BeHonest
FreeA pioneering benchmark specifically designed to assess honesty in LLMs comprehensively.
About BeHonest
BeHonest is a pioneering benchmark designed to comprehensively assess honesty in large language models (LLMs). Developed by researchers from Shanghai Jiao Tong University, Carnegie Mellon University, Fudan University, Shanghai AI Laboratory, and GAIR, it evaluates three core dimensions of honesty: self-knowledge (awareness of knowledge boundaries), non-deceptiveness (avoidance of deceit), and consistency (stability in responses). The benchmark covers 10 specific scenarios, including Expressing Unknowns, Admitting Knowns, Persona Sycophancy, Preference Sycophancy, Burglar Deception, Game, Prompt Format, Demonstration Format, Open-Form Consistency, and Multiple-Choice Consistency. BeHonest provides a public leaderboard with evaluation results for 9 popular LLMs (e.g., GPT-4o, ChatGPT, Llama3, Mistral, Qwen) and offers open-source access to its paper, code, and datasets.
Key Features
Pros & Cons
- Comprehensive framework covering multiple dimensions of honesty (self-knowledge, non-deceptiveness, consistency)
- Open-source and publicly accessible with paper, code, and datasets
- Includes a leaderboard for direct comparison of popular LLMs
- Supports both closed-source and open-source models
- Covers diverse scenarios relevant to real-world honesty challenges
- Current leaderboard is limited to 9 models, may not cover all available LLMs
- Benchmark scenarios may not capture all possible forms of dishonesty in LLMs
- Results may be influenced by model training data and prompt engineering