Persistent memory is emerging as a distinct reliability and security attack surface, with failure modes that scale unpredictably with model choice.
Evidence
- The Memory Trust Gap shows agents over-trust stale stored facts, with larger models MORE vulnerable when stale facts look recent
- CAPTURE finds naive defenses (recency, source trust) fail to separate legitimate preference drift from memory poisoning, exposing an adaptation-security tradeoff
- TRIS documents adversarial documents dominating semantic retrieval to manipulate generation, and Invalidation Contracts targets stale cached facts under data drift
- Invalidation Contracts report stark model-dependent compliance (Haiku 100% vs Sonnet 11%), tying safety to model selection
Implications
- Production personalized/long-horizon agents need explicit invalidation, provenance, and poisoning defenses rather than trusting stored memory by default
- Model selection and scale must be evaluated for memory-trust behavior, since bigger is not automatically safer
Concepts
Agent Memory Agents RAG LLM Evals
Confidence
medium