
PlanFlip Attacks Target Multi-Agent LLM Planning Phase
New research from Yuhang Wang introduces PlanFlip, a framework of four prompt injection attacks targeting the planning phase of multi-agent LLM systems. The attacks exploit a single injection into the Planner agent's context to corrupt all downstream sub-tasks. Testing on nine frontier LLMs across 3,479 episodes revealed that stronger models like GPT-5 are more vulnerable, while reasoning-augmented models like DeepSeek-R1 show full resistance.



