Research

AI beats law professors at answering student questions, study finds

A Stanford Law School study found that law professors preferred AI-generated answers to student questions over responses written by fellow instructors. In blind evaluations of nearly 3,000 comparisons, AI won 75% of head-to-head matchups. The results challenge assumptions about AI's role in legal education.

Neura News

Neura News

Neura Market Editorial

June 3, 20264 min read

Originally reported by law.stanford.edu

AI beats law professors at answering student questions, study finds

Law professors showed a strong preference for AI-generated answers to student questions over responses written by their own colleagues, according to a study led by Stanford Law School. The findings could reshape how legal education is delivered.

The study, titled "Law Professors Prefer AI Over Peer Answers," was led by Professor Julian Nyarko and involved 16 law professors from U.S. law schools. It tested whether large language models could serve as effective tutors for contract law courses. In a blind evaluation of nearly 3,000 anonymized comparisons, professors rated AI responses significantly higher than answers written by other professors, with AI winning 75% of head-to-head matchups.

Rigorous testing in a judgment-rich field

"This study challenges important assumptions about AI's role in legal education," said Nyarko, who directs Stanford Law School's Legal Innovation through Frontier Technology Lab, or liftlab. He co-authored the paper with colleagues from Yale, NYU, the University of Chicago, and other leading institutions. "We focused on law precisely because it requires judgment, nuanced reasoning, and the ability to navigate ambiguity, not just factual recall."

The study stands out because previous AI evaluations focused mainly on subjects with clear right-or-wrong answers. Legal reasoning, by contrast, demands careful analysis of competing arguments and defensible conclusions.

"We were frankly surprised by the magnitude of the results," Nyarko added. "These weren't just simple questions with obvious answers. Many of them required synthesizing complex material, applying it to new situations, and explaining legal concepts in ways that would help students develop their own analytical skills."

Participants created 40 representative contract law questions that students might ask after class or during office hours. They wrote their own answers and then evaluated responses without knowing whether they came from AI or other participating professors. The AI systems performed comparably to the best human instructor in the study.

Perhaps most striking: professors flagged AI responses as pedagogically harmful only 3.5% of the time, compared to 12% for peer-written answers.

AI meets the professional standard

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

"In most fields where AI gets tested, there's a right answer. In law, there often isn't," said Sarath Sanga, co-author and professor at Yale Law School. "Two opposing arguments can both be good. What we wanted to know is whether AI can meet the latent professional standard that lawyers use to evaluate each other's arguments. In this case, the answer was yes."

The research team took extensive precautions to ensure the study's validity. They calibrated AI responses to match the length and structure of human answers, used multiple evaluation methods, and had professors assess whether responses might mislead or confuse students.

Alejandro Salinas, first author of the study and a researcher at Nyarko's liftlab, emphasized the educational implications: "Our study shifts attention to what AI tutoring can contribute to learning in judgment-rich fields like law. We find that, when evaluated by legal educators, AI tutors can offer high-quality, on-demand support that complements classroom instruction, and may broaden access to expert guidance."

The study also examined specific AI models, including commercial tutoring systems and Google's NotebookLM, finding varying levels of performance. However, even when context limitations affected AI responses, professors still frequently preferred them to human-written alternatives.

Caution and opportunity in legal education

The findings arrive as law schools nationwide grapple with integrating AI tools into legal education while maintaining rigorous academic standards. Some institutions have embraced AI experimentation, while others remain cautious about potential risks including hallucinations, overreliance, and the erosion of critical thinking skills.

"Our study evaluates the quality of answers given by AI tools. But how to implement these tools to most effectively improve student learning is still an open question. So we're not advocating for wholesale adoption of AI tutors," Nyarko cautioned. "But our data suggests that blanket skepticism may be equally unwarranted. The conversation should shift from whether AI can give accurate, high quality responses to how we can deploy it responsibly to the benefit of our students."

Liftlab, based at Stanford Law School, is among the first academic efforts in legal AI to unite research, prototyping, and real-time collaboration with industry. Its mission is to increase access to high quality legal services in the private sector by leveraging AI and other frontier technologies.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read