Preprint
Machine Learning

Reasoning Models Ace the CFA Exams

J. Patel, Yunzhe Chen, Kaiwen He, Keyi Wang, David Li, Kairong Xiao, Xiao-Yang Liu
December 9, 2025arXiv.org1 citations

1

Citations

0

Influential Citations

arXiv.org

Venue

2025

Year

Abstract

Previous research has reported that large language models (LLMs) demonstrate poor performance on the Chartered Financial Analyst (CFA) exams. However, recent reasoning models have achieved strong results on graduate-level academic and professional examinations across various disciplines. In this paper, we evaluate state-of-the-art reasoning models on a set of mock CFA exams consisting of 980 questions across three Level I exams, two Level II exams, and three Level III exams. Using the same pass/fail criteria from prior studies, we find that most models clear all three levels. The models that pass, ordered by overall performance, are Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1. Specifically, Gemini 3.0 Pro achieves a record score of 97.6% on Level I. Performance is also strong on Level II, led by GPT-5 at 94.3%. On Level III, Gemini 2.5 Pro attains the highest score with 86.4% on multiple-choice questions while Gemini 3.0 Pro achieves 92.0% on constructed-response questions.

Analysis

Why This Paper Matters

This paper directly challenges earlier findings that LLMs performed poorly on the CFA exams. By evaluating the latest generation of reasoning models, it shows a dramatic improvement: most models now pass all three levels, with top scores exceeding 90% on multiple-choice sections. This is significant because the CFA exam is a rigorous, graduate-level professional certification that tests deep financial reasoning, not just factual recall. The results indicate that reasoning models have reached a level of competence that could make them useful in financial analysis, portfolio management, and investment decision support.

The study also provides a clear, reproducible benchmark (980 questions across three levels) that can be used by the AI community to track progress on professional-domain reasoning. This is particularly valuable as the field moves beyond general knowledge benchmarks to specialized, high-stakes evaluations.

Technical Contributions

  • Comprehensive benchmark: The authors assembled a diverse set of mock CFA questions covering all three levels, including both multiple-choice and constructed-response formats, providing a robust testbed for reasoning models.
  • Standardized evaluation: They applied the same pass/fail criteria as prior studies, enabling direct comparison with earlier LLM performance and highlighting the improvement over time.
  • Multi-model comparison: The study evaluates six leading reasoning models from different vendors (Google, OpenAI, xAI, Anthropic, DeepSeek), offering a broad view of the current state of the art.
  • Level-specific analysis: The paper breaks down performance by exam level and question type, revealing strengths and weaknesses—e.g., Gemini 3.0 Pro excels at constructed-response, while Gemini 2.5 Pro leads on Level III multiple-choice.

Results

  • Overall pass rates: Most models cleared all three CFA levels, a stark contrast to earlier LLMs that failed.
  • Level I: Gemini 3.0 Pro achieved a record 97.6%, the highest reported score on this benchmark.
  • Level II: GPT-5 led with 94.3%, showing strong analytical reasoning on complex financial scenarios.
  • Level III: Gemini 2.5 Pro scored 86.4% on multiple-choice questions, while Gemini 3.0 Pro achieved 92.0% on constructed-response questions, indicating robust performance on essay-style answers.
  • Ranking: The models that passed, in order of overall performance, were Gemini 3.0 Pro, Gemini 2.5 Pro, GPT-5, Grok 4, Claude Opus 4.1, and DeepSeek-V3.1.

Significance

This paper provides strong evidence that reasoning models have crossed a threshold in professional-domain expertise. The ability to pass the CFA exam suggests that these models can handle complex, multi-step financial reasoning, which has implications for AI-assisted financial analysis, automated advisory services, and educational tools. It also sets a new benchmark for evaluating LLMs on professional certifications, encouraging further research into domain-specific reasoning. As models continue to improve, we can expect even higher scores and broader adoption in finance and other regulated industries, though careful validation and human oversight remain essential.