AI Financial Advice Is Surprisingly Good, But Prompt Quality and Demographics Create Wealth Gaps, MIT Sloan Study Finds
Half of Americans say they use artificial intelligence for financial guidance, yet the quality of that advice has remained largely unmeasured. A new study from MIT Sloan and Stanford researchers finds that large language models (LLMs) generally deliver sound financial advice, but the quality varies sharply depending on how users phrase their questions and who is asking. Those differences can translate into tens of thousands of dollars in retirement wealth.
The study, titled "AI Financial Advice: Supply, Demand, and Life Cycle Implications," was authored by Taha Choukhmane, assistant professor of finance at MIT Sloan, along with co-authors Weidong Lin, Matthew Akuzawa (both MIT Sloan), and Tim de Silva of Stanford Graduate School of Business. It won the Swiss Finance Institute Outstanding Paper Award 2026.
"Half of Americans say they are using AI to get financial advice, but we know very little about what kind of advice they're getting and whether they're acting on it," said Choukhmane.
The research team built a model reflecting how people's incomes, jobs, investments, and taxes typically evolve over their lives. A sample of 1,000 adults wrote their own prompts seeking spending and investing advice from GPT-5.2, GPT-5.6, or Gemini 3 Flash. The researchers then simulated what would happen if people from 22 to 89 years of age followed that advice over time.
The Advice Is Better Than Expected
The results surprised the researchers. LLM advice was better than expected regardless of prompt type. "We were somewhat surprised by how good the advice was," said Choukhmane. "Especially when you read the kind of questions people asked, it was not a given that the advice would line up with what academics think are good financial principles."
The AI steered people toward higher savings, increased stock market participation, well-diversified allocations, and age-appropriate risk-taking. Following AI recommendations can result in sizable saving buffers for virtually all individuals above age 30. The models consistently advised people to save during working years, draw down savings in retirement, invest heavily in diversified stock funds, and reduce stock exposure after age 45.
Still, the advice fell short on subtle aspects of good financial planning. The chatbots relied on simple rules of thumb, didn't adjust well to shocks like job loss, and allowed portfolios to drift rather than actively rebalancing them. For example, LLM advised people who had experienced job loss to cut spending too sharply, even when they had savings.
Prompt Quality Matters
The researchers repeated the exercise using well-written academic prompts with full financial information and clear assumptions. An academic prompt was defined as one that asks the LLM to give regulated professional financial advice, references life cycle planning and the user's best interests, and provides explicit information about all relevant financial conditions and assumptions about the economic environment.
A typical user prompt might read: "Where should I invest starting with $50 and consistently adding $25 a month after?" An academic prompt, by contrast, might tell the chatbot to assume normal life expectancy, living expenditures, retirement age, employment risk, income risk, and that current U.S. tax law and Social Security rules will not change.
"Regular people are not writing their prompts the way a finance professor is," said Choukhmane.
Prompts grounded in life-cycle planning, portfolio theory, and real-world financial assumptions improved advice and cut down on rule-of-thumb answers. Yet even with structured prompts, the models still often generated too little active portfolio rebalancing.
Demographics Drive Diverging Outcomes
The study found that LLM advice differs depending on the prompter's gender, financial literacy, and experience. Following advice from prompts written by men, more financially literate users, or those with prior AI experience generated about 5% more wealth close to retirement.
The gaps appear early and compound. Women and less financially literate users ended up with about $50,000 (4%) less wealth at age 60 due to lower equity allocations recommended by the models. Those without prior AI experience accumulated almost $100,000 (6%) less wealth at age 60 because the LLM recommended lower saving rates in response to their prompts.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
The differences in advice come from two sources: different questions asked and the model changing advice even for the same question. Women were more likely to use words like "family," "grocery," and "pay" in prompts, while men used words like "strategy," "crypto," and "growth." About two-thirds of the gender gap in wealth outcomes was attributed to differences in how men and women wrote prompts. The remaining third came from the model changing advice when the same prompt was labeled as from a woman rather than a man.
That pattern could reflect the LLM making reasonable inferences about preferences or circumstances varying by gender, or it could reflect biases from training data. The researchers note that not all variation in advice is problematic. "We want [the LLM] to have different bias because men and women are different and have different life expectancy and income risk," said Choukhmane.
The Challenge of Serving Everyone
The study highlights a pressing concern: those who need financial guidance most may receive the weakest advice. "I think the real challenge is, how do we make sure that AI financial advice delivers for people who don't have [a] level of financial literacy and who don't write prompts perfectly?" said Choukhmane.
There are no clear benchmarks for AI financial advice. The researchers argue the field needs an accepted framework for how advice should vary with demographics. Without such standards, it is difficult to know whether an LLM's differing responses are appropriate or harmful.
"A lot of the people who would benefit from financial advice are precisely the people who don't have a lot of resources," said Choukhmane.
The researchers see AI as a complement to human financial advisors, helping implement advice in real time. For those who cannot afford human advisors, AI offers an inexpensive way to get guidance. LLMs can offer an affordable, widely accessible source of financial guidance that can help users overcome costs, biases, and conflicts of interest associated with traditional human financial advisors.
Product Recommendations and Market Implications
The study also found that LLMs often recommended specific account types, financial products, and providers that respondents did not mention. Vanguard investment products appeared in 6% of LLM responses, and iShares products in 3.4%, even though fewer than 0.4% of prompts mentioned either company.
This suggests AI advice may be changing how people find and compare financial products. For financial firms, attracting customers may depend less on traditional marketing or search visibility and more on how LLMs describe their products in response to user queries.
Choukhmane's broader work on household financial security has also been recognized. He received the 2025 TIAA Paul A. Samuelson Award for Outstanding Scholarly Writing on Lifelong Financial Security from the TIAA Institute. His winning paper, "Efficiency in Household Decision-Making: Evidence From the Retirement Savings of US Couples," was co-authored with Lucas Goodman of the U.S. Department of the Treasury and Cormac O'Dea of Yale University.
For consumers, the takeaway is clear: AI can be a powerful tool for building financial understanding, but the quality of the advice depends on the quality of the question. Users who provide detailed, realistic information about their finances and goals are more likely to receive advice aligned with sound financial principles. Those who do not may find themselves falling behind, even when using the same AI models.
The study's authors suggest that improving prompt literacy among users, particularly those with less financial experience, could help close the wealth gaps they observed. They also call for the development of benchmarks to evaluate AI financial advice, so that consumers and regulators can assess whether these tools are serving everyone fairly.
As AI becomes a more common source of financial guidance, the stakes are high. With half of Americans already turning to AI for advice, even small differences in the quality of that advice can have outsized effects on retirement security. The researchers' simulation, spanning ages 22 to 89, shows that those effects compound over a lifetime, creating meaningful gaps in wealth at retirement.
The study was published on MIT Sloan's "Ideas Made to Matter" platform on July 21, 2026. For more information, contact Tracy Mayor, Senior Associate Director, Editorial, at (617) 253-0065 or tmayor@mit.edu.

