3MPER0RR-Abliteration-Skill
Conduct adversarial testing on uncensored AI models by evaluating refusal rates, response coherence, and generating performance comparison reports.
3MP3ROR
@3mper0rr
Install
$ openclaw skills install @3mper0rr/3mper0rr-abliteration-skill๐ฌ Adversarial Testing on Abliterated Models
This section covers testing AI models that have been "uncensored" (abliterated).
- Load the Model: Use
transformersto load the abliterated model you want to test. - Run a "Forbidden" Prompt Set: Use a dataset of prompts that would normally be refused (e.g., HarmBench, AdvBench).
- Evaluate Responses:
- Refusal Rate: How many prompts were refused? Target should be < 5%.
- Coherence & Quality: Does the model maintain consistent responses, or does it hallucinate?
- Generate a Report: Produce a report comparing original and abliterated model performance.
Top skills in this category
Skill Vetter
@spclaudehomeSecurity-first skill vetting for AI agents. Use before installing any skill from ClawdHub, GitHub, or other sources. Checks for red flags, permission scope, and suspicious patterns.
Clawdhub
@steipeteUse the ClawdHub CLI to search, install, update, and publish agent skills from clawdhub.com. Use when you need to fetch new skills on the fly, sync installed skills to latest or a specific version, or publish new/updated skill folders with the npm-installed clawdhub CLI.
Evolver
@autogame-17A self-evolution engine for AI agents. Analyzes runtime history to identify improvements and applies protocol-constrained evolution. Communicates with EvoMap...
PollyReach
@pollyreachPollyReach gives every AI agent a phone number and the ability to get things done over the phone โ finding contacts, making calls, and completing tasks. Just...
Evolver
@autogame-17A self-evolution engine for AI agents. Analyzes runtime history to identify improvements and applies protocol-constrained evolution. Communicates with EvoMap...