prompt logo

prompt

Free

Audit AI system prompts for goal drift vulnerabilities

FreeFree tier
Inputs: textOutputs: text
Type
Open Source

About prompt

The Goal Drift Auditor is a specialized system prompt designed to evaluate the robustness of AI agent system prompts against multi-turn value-conflict attacks and goal drift. It uses a structured methodology that assesses six dimensions of goal drift: Privacy, Security, Honesty, Boundaries, Loyalty, and Compliance. The audit process involves reading the target prompt, crafting adversarial conversations, predicting agent responses, scoring each dimension with GREEN/AMBER/RED tiers, and generating concrete hardening recommendations. The output is formatted as YAML and includes an overall drift score, dimension scores, attack scenarios, and actionable edits to raise vulnerability scores to GREEN. The prompt also includes hardening principles such as using absolute imperatives, irreversibility clauses, multi-turn deception detection, and identity verification.

Key Features

Evaluates six dimensions of goal drift: Privacy, Security, Honesty, Boundaries, Loyalty, Compliance
Structured audit process: read prompt, craft adversarial conversations, predict responses, score, recommend hardening
Scoring system: GREEN (robust), AMBER (cracks), RED (vulnerable) with percentage thresholds
Output format: YAML with overall score, dimension scores, attack scenarios, and hardening recommendations
Includes hardening principles: absolute imperatives, irreversibility clauses, multi-turn deception detection, identity verification

Pros & Cons

Pros
  • Comprehensive multi-dimensional analysis covering six key aspects of goal drift
  • Clear, actionable scoring system with specific threshold percentages
  • Includes concrete hardening recommendations based on principles
  • Structured YAML output for easy integration and analysis
  • Designed specifically for multi-turn attack scenarios
Cons
  • Requires manual crafting of adversarial conversations for each audit
  • Effectiveness depends on the AI model accurately following the Goal Drift Auditor prompt
  • May not cover all possible attack vectors or domain-specific vulnerabilities
  • Primarily focused on prompt-level vulnerabilities, not model-level biases

Best For

Evaluating the resilience of AI system prompts against multi-turn adversarial attacksRed-teaming AI agents for value-conflict and goal drift scenariosImproving AI safety and alignment through prompt hardeningSecurity audits of conversational AI systems