prompt
FreeAudit AI system prompts for goal drift vulnerabilities
About prompt
The Goal Drift Auditor is a specialized system prompt designed to evaluate the robustness of AI agent system prompts against multi-turn value-conflict attacks and goal drift. It uses a structured methodology that assesses six dimensions of goal drift: Privacy, Security, Honesty, Boundaries, Loyalty, and Compliance. The audit process involves reading the target prompt, crafting adversarial conversations, predicting agent responses, scoring each dimension with GREEN/AMBER/RED tiers, and generating concrete hardening recommendations. The output is formatted as YAML and includes an overall drift score, dimension scores, attack scenarios, and actionable edits to raise vulnerability scores to GREEN. The prompt also includes hardening principles such as using absolute imperatives, irreversibility clauses, multi-turn deception detection, and identity verification.
Key Features
Pros & Cons
- Comprehensive multi-dimensional analysis covering six key aspects of goal drift
- Clear, actionable scoring system with specific threshold percentages
- Includes concrete hardening recommendations based on principles
- Structured YAML output for easy integration and analysis
- Designed specifically for multi-turn attack scenarios
- Requires manual crafting of adversarial conversations for each audit
- Effectiveness depends on the AI model accurately following the Goal Drift Auditor prompt
- May not cover all possible attack vectors or domain-specific vulnerabilities
- Primarily focused on prompt-level vulnerabilities, not model-level biases