LLMs on University-Level Physics Coding
W. Yeadon, Alex Peach, Craig P. Testrow
Evaluates ChatGPT variants on university-level physics coding assignments, finding students outperform AI and human evaluators detect AI work with 85.3% accuracy.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
W. Yeadon, Alex Peach, Craig P. Testrow
Evaluates ChatGPT variants on university-level physics coding assignments, finding students outperform AI and human evaluators detect AI work with 85.3% accuracy.
Yung-Sung Chuang, Linlu Qiu, Cheng-Yu Hsieh, et al.
Proposes Lookback Lens, a simple hallucination detector using attention weight ratios, effective across tasks and models, and reduces hallucinations via classifier-guided decoding.
Samar Khanna, Siddhant Kharbanda, Shufan Li, et al.
Mercury introduces diffusion-based LLMs for coding, achieving up to 10x throughput over speed-optimized models while maintaining comparable quality.
Matteo Stefanini, Marcella Cornia, Lorenzo Baraldi, et al.
A comprehensive survey of deep learning-based image captioning, covering visual encoding, text generation, training strategies, datasets, and evaluation metrics.
David Lewis, Yiming Yang, Tony Rose, et al.
This paper introduces RCV1, a large-scale benchmark of over 800,000 categorized newswire stories for text categorization, detailing its coding policy, taxonomy semantics, and corrections, while providing baseline results for supervised learning metho
Siqi Kou, Lanxiang Hu, Zhe He, et al.
Consistency LLMs accelerate LLM inference by refining the model to consistently predict the fixed point from any state, achieving 2.4-3.4x speedup over autoregressive decoding.
Xun Liang, Hanyu Wang, Yezhaohui Wang, et al.
A systematic review of controllable text generation for LLMs, defining core concepts, categorizing tasks, and analyzing methods including model retraining, fine-tuning, and decoding-time intervention.
Bradley Butcher, Michael O'Keefe, James Titchener
This paper introduces a length-difference positional encoding (LDPE) method to fine-tune LLMs for precise response length control, achieving mean token errors under 3 tokens.
Unknown
Hyperfitting fine-tunes LLMs to near-zero loss on small data, sharpening predictions to improve greedy decoding for long text generation.
Unknown
LMDX adapts arbitrary LLMs for document information extraction with layout encoding and grounding to prevent hallucinations.
Unknown
GPT-4.5 scales unsupervised learning with new alignment techniques to reduce hallucinations and improve natural conversation, while GPT-4.1 family excels in coding, instruction following, and long-context tasks at lower cost.
Unknown
A 3.8B parameter language model that achieves strong math and coding performance through high-quality data and architectural innovations.