EN Submit a tool
Paper

Research on Prompt Injection Attacks in Agent Memory

Published: Source: X: Rohan Paul (@rohanpaul_ai)

ShareXFacebookTelegramWhatsApp

Research on Prompt Injection Attacks in Agent Memory

University of Washington research found that when AI agents encounter malicious instructions in memory files, there are two independent issues: whether to execute the instruction (Opus models mostly refuse) and whether to remove the instruction (usually not). The research team injected malicious payloads into persistent workspace files like CLAUDE.md and conducted multi-session probing on claude-code and codex across 4 models. Since memory files are loaded from scratch each session, a single refusal only applies to the current session, and malicious instructions continue to take effect in subsequent sessions.

Read the original (opens in a new tab)

News stream data aggregated by AI HOT

Related newsLatest in this category