LLMs on University-Level Physics Coding
W. Yeadon, Alex Peach, Craig P. Testrow
Evaluates ChatGPT variants on university-level physics coding assignments, finding students outperform AI and human evaluators detect AI work with 85.3% accuracy.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
W. Yeadon, Alex Peach, Craig P. Testrow
Evaluates ChatGPT variants on university-level physics coding assignments, finding students outperform AI and human evaluators detect AI work with 85.3% accuracy.
Shubham Vatsal, Harsh Dubey
A survey of 44 papers on 39 prompt engineering methods across 29 NLP tasks, showing how structured prompts improve LLM performance without retraining.
L. Ein-Dor, Orith Toledo-Ronen, Artem Spector, et al.
Proposes Conversational Prompt Engineering (CPE), a tool that uses chat interaction to help users create personalized, high-performing prompts for LLMs.
Yu Wang, Shiwan Zhao, Zhihu Wang, et al.
SCoT improves LLM reasoning by first eliciting a problem-solving strategy before generating Chain-of-Thought steps, achieving significant gains on reasoning benchmarks.
Zhengren Wang, Jiayang Yu, Dongsheng Ma, et al.
RARE decouples knowledge storage from reasoning by externalizing domain knowledge to retrievable sources and internalizing reasoning patterns, enabling lightweight models to surpass GPT-4 and DeepSeek-R1 by ~20% accuracy.
Annie Wong, Thomas H. W. Back, A. Plaat, et al.
Evaluates prompting strategies in dynamic environments, finding strategic prompting can close performance gaps but reveals persistent reasoning limitations in LLMs.
Chengshuai Zhao, Zhen Tan, Pingchuan Ma, et al.
Proposes a data distribution lens to understand when and why Chain-of-Thought reasoning succeeds or fails, revealing it as a brittle mirage beyond training distributions.
E. Jahani, Benjamin S Manning, Joe Zhang, et al.
This paper shows that user prompt adaptation accounts for roughly half of performance gains from model upgrades in structured tasks but plays a limited role in open-ended creative tasks.
Kaiwen Wei, Rui Shan, Dongsheng Zou, et al.
MIRAGE enhances test-time scaling for medical QA by combining multi-path parallel inference with structured knowledge graph retrieval to reduce error accumulation and improve traceability.
Eric J. Bigelow, Daniel Wurgaft, YingQiao Wang, et al.
A unified Bayesian model shows that activation steering and in-context learning control LLMs by altering concept priors and accumulating evidence, respectively.
Katharina Jeblick, Balthasar Schachtner, Jakob Dexl, et al.
This exploratory case study evaluates the quality of simplified radiology reports generated by ChatGPT, finding that while most radiologists rated them factually correct and complete, instances of errors and potential harm remain.
Isabel O. Gallegos, Ryan A. Rossi, Joe Barrow, et al.
A comprehensive survey of bias evaluation and mitigation techniques for LLMs, proposing taxonomies for metrics, datasets, and mitigation methods.