Skip to content

Persistent memory is emerging as a distinct reliability and security attack surface, with failure modes that scale unpredictably with model choice.

Evidence

  • The Memory Trust Gap shows agents over-trust stale stored facts, with larger models MORE vulnerable when stale facts look recent
  • CAPTURE finds naive defenses (recency, source trust) fail to separate legitimate preference drift from memory poisoning, exposing an adaptation-security tradeoff
  • TRIS documents adversarial documents dominating semantic retrieval to manipulate generation, and Invalidation Contracts targets stale cached facts under data drift
  • Invalidation Contracts report stark model-dependent compliance (Haiku 100% vs Sonnet 11%), tying safety to model selection

Implications

  • Production personalized/long-horizon agents need explicit invalidation, provenance, and poisoning defenses rather than trusting stored memory by default
  • Model selection and scale must be evaluated for memory-trust behavior, since bigger is not automatically safer

Concepts

Agent Memory Agents RAG LLM Evals

Confidence

medium