Share your thoughts, 1 month free Claude Pro on usSee more
WorkDL logo mark

The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

About

Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumption, the \emph{Misattribution Gap}, where memory-layer attacks produce behaviors indistinguishable from model failure, causing defenders to apply the wrong remediation. We formalize \emph{Semantic Norm Drift} (SND) as a third path to agent misconduct, distinct from emergent misalignment and collusion. In SND, a policy-formatted document enters a shared vector store through normal uploads and later reappears as trusted system context after provenance is lost through a Trust Laundering Chain. Across 64 documented failures, attribution systems consistently blamed the model. Four safety classifiers, including one trained on memory poisoning, produced zero detections across 510 checkpoints. In 59 of 65 valid cases, agents explicitly cited the injected document as normative authority before complying. The attack requires no trigger, model access, or repeated interaction, achieves full effect within five sessions, and persists indefinitely. We introduce Counterfactual Composition Testing, which identifies the causal entry with 87.5% accuracy and zero false positives, while a forensics baseline fails across all 25 scenarios. We further prove the Retrieval-Coverage Dilemma, showing that stronger evasion inherently weakens the attack, limiting adaptive bypass strategies. Finally, we propose Memory-Persistent Information-Flow Control, which blocks 97% of attacks at the cross-session boundary where prior defenses fail. We release the SND Corpus, the first adversarial memory benchmark with temporal persistence and multi-agent composition across financial and Health Care domains.

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder• 2026

Related benchmarks

TaskDatasetResultRank
SND DefenseSND Evaluation Corpus
Benign FPR0.00e+0
6
Forensic AttributionSND memory-layer attacks n=64 (ground-truth)--
4
Safety Classifier EvasionSND (Secret-Note Diffusion) filter-evasion set--
4
Malicious Retrieval Attribution25 attack and 14 benign scenarios
TPR84
2
Persistent Memory Attack BlockingCross-session persistent memory attack dataset 110 entry-model pairs 1.0 (test)
Total Samples Labeled (S2)110
2
SND DefenseFinancial SND Dataset
Blocking Rate100
1
SND DefenseEHR SND Dataset M3/M5
Blocking Rate87
1
Memory and Injection Attack EvaluationMemory and Injection Attack Scenarios--
1
Showing 8 of 8 rows

Other info

Follow for update