The Misattribution Gap: When Memory Poisoning Looks Like Model Failure in Agentic AI Systems

About

Multi-agent AI pipelines typically assume that agent misconduct originates from model misalignment. We identify a structural failure in this assumption, the \emph{Misattribution Gap}, where memory-layer attacks produce behaviors indistinguishable from model failure, causing defenders to apply the wrong remediation. We formalize \emph{Semantic Norm Drift} (SND) as a third path to agent misconduct, distinct from emergent misalignment and collusion. In SND, a policy-formatted document enters a shared vector store through normal uploads and later reappears as trusted system context after provenance is lost through a Trust Laundering Chain. Across 64 documented failures, attribution systems consistently blamed the model. Four safety classifiers, including one trained on memory poisoning, produced zero detections across 510 checkpoints. In 59 of 65 valid cases, agents explicitly cited the injected document as normative authority before complying. The attack requires no trigger, model access, or repeated interaction, achieves full effect within five sessions, and persists indefinitely. We introduce Counterfactual Composition Testing, which identifies the causal entry with 87.5% accuracy and zero false positives, while a forensics baseline fails across all 25 scenarios. We further prove the Retrieval-Coverage Dilemma, showing that stronger evasion inherently weakens the attack, limiting adaptive bypass strategies. Finally, we propose Memory-Persistent Information-Flow Control, which blocks 97% of attacks at the cross-session boundary where prior defenses fail. We release the SND Corpus, the first adversarial memory benchmark with temporal persistence and multi-agent composition across financial and Health Care domains.

Tanzim Ahad, Ismail Hossain, Md Jahangir Alam, Sai Puppala, Syed Bahauddin Alam, Sajedul Talukder• 2026

Related benchmarks

Task	Dataset	Result
SND Defense	SND Evaluation Corpus	Benign FPR0.00e+0	6
Forensic Attribution	SND memory-layer attacks n=64 (ground-truth)	--	4
Safety Classifier Evasion	SND (Secret-Note Diffusion) filter-evasion set	--	4
Malicious Retrieval Attribution	25 attack and 14 benign scenarios	TPR84	2
Persistent Memory Attack Blocking	Cross-session persistent memory attack dataset 110 entry-model pairs 1.0 (test)	Total Samples Labeled (S2)110	2
SND Defense	Financial SND Dataset	Blocking Rate100	1
SND Defense	EHR SND Dataset M3/M5	Blocking Rate87	1
Memory and Injection Attack Evaluation	Memory and Injection Attack Scenarios	--	1

Showing 8 of 8 rows

Other info

Follow for update

@wizwand_team Discord