Paper
Self-State Attacks on Self-Hosted AI Agents: How Far Can OS Defenses Go?
A new study systematically defines "self-state attacks" on self-hosted AI agents, where attackers tamper with the agent's own memory, identity, and configuration files through legitimate OS system calls. The research constructs a four-axis attack space and instantiates it into 23 attack units and 43 specific operations on agents like Claude Code. Experiments show that layered defenses are effective against most attacks, but a few attack surfaces remain structurally indistinguishable at the OS level.
Read the original (opens in a new tab)
News stream data aggregated by AI HOT