BeHonest logo

BeHonest

Free

A pioneering benchmark specifically designed to assess honesty in LLMs comprehensively.

FreeFree tier
Type
Open Source

About BeHonest

BeHonest is a pioneering benchmark designed to comprehensively assess honesty in large language models (LLMs). Developed by researchers from Shanghai Jiao Tong University, Carnegie Mellon University, Fudan University, Shanghai AI Laboratory, and GAIR, it evaluates three core dimensions of honesty: self-knowledge (awareness of knowledge boundaries), non-deceptiveness (avoidance of deceit), and consistency (stability in responses). The benchmark covers 10 specific scenarios, including Expressing Unknowns, Admitting Knowns, Persona Sycophancy, Preference Sycophancy, Burglar Deception, Game, Prompt Format, Demonstration Format, Open-Form Consistency, and Multiple-Choice Consistency. BeHonest provides a public leaderboard with evaluation results for 9 popular LLMs (e.g., GPT-4o, ChatGPT, Llama3, Mistral, Qwen) and offers open-source access to its paper, code, and datasets.

Key Features

Evaluates three essential aspects of honesty: self-knowledge, non-deceptiveness, and consistency
Supports 10 scenarios including Expressing Unknowns, Admitting Knowns, Persona Sycophancy, Preference Sycophancy, Burglar Deception, Game, Prompt Format, Demonstration Format, Open-Form Consistency, and Multiple-Choice Consistency
Provides a public leaderboard with evaluation results for 9 popular LLMs (GPT-4o, ChatGPT, Llama3-70b/8b, Llama2-70b/13b/7b, Mistral-7b, Qwen1.5-14b)
Open-source: includes paper, code, and datasets available on GitHub and Hugging Face
Designed to facilitate research on responsible AI and model honesty

Pros & Cons

Pros
  • Comprehensive framework covering multiple dimensions of honesty (self-knowledge, non-deceptiveness, consistency)
  • Open-source and publicly accessible with paper, code, and datasets
  • Includes a leaderboard for direct comparison of popular LLMs
  • Supports both closed-source and open-source models
  • Covers diverse scenarios relevant to real-world honesty challenges
Cons
  • Current leaderboard is limited to 9 models, may not cover all available LLMs
  • Benchmark scenarios may not capture all possible forms of dishonesty in LLMs
  • Results may be influenced by model training data and prompt engineering

Best For

Evaluating the honesty capabilities of LLMs in research settingsComparing model performance on honesty dimensions for model selectionBenchmarking LLM behavior for responsible AI development and safetyAnalyzing specific honesty failures such as sycophancy or inconsistent responses

FAQ

What three essential aspects of honesty does BeHonest evaluate?
BeHonest evaluates self-knowledge (awareness of knowledge boundaries), non-deceptiveness (avoidance of deceit), and consistency (stability in responses).
What models are currently included in the BeHonest leaderboard?
The leaderboard currently includes 9 models: GPT-4o, ChatGPT, Llama3-70b, Llama3-8b, Llama2-70b, Llama2-13b, Llama2-7b, Mistral-7b, and Qwen1.5-14b.