LLMs for Data Annotation
Zhen Tan, Dawei Li, Song Wang, et al.
This survey uniquely focuses on LLMs for data annotation, covering generation, assessment, and utilization of LLM-generated annotations.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Zhen Tan, Dawei Li, Song Wang, et al.
This survey uniquely focuses on LLMs for data annotation, covering generation, assessment, and utilization of LLM-generated annotations.
Ranjay Krishna, Yuke Zhu, Oliver Groth, et al.
The Visual Genome dataset provides dense annotations of objects, attributes, and relationships in over 108K images to enable cognitive tasks like image description and question answering.
Saif M. Mohammad, Peter D. Turney
This paper demonstrates how crowdsourcing can efficiently create a large, high-quality word-emotion and polarity lexicon, addressing challenges like malicious annotation and sense-level disambiguation.
Caleb Ziems, William A. Held, Omar Ahmed Shaikh, et al.
This paper provides a roadmap for using zero-shot LLMs as tools in computational social science, showing they can augment human annotation and bootstrapping creative generation tasks.
Teng Wang, Zhangyi Jiang, Zhenqi He, et al.
Proposes Hierarchical Reward Model (HRM) and Hierarchical Node Compression (HNC) to improve multi-step reasoning evaluation in LLMs while reducing annotation costs.
Yifei Zhou, Sergey Levine, J. Weston, et al.
Self-Challenging framework enables LLM agents to self-generate high-quality training tasks via Code-as-Task, achieving over two-fold improvement on tool-use benchmarks without human annotation.
Yung-Sung Chuang, Benjamin Cohen-Wang, Shannon Zejiang Shen, et al.
SelfCite introduces a self-supervised method using context ablation to generate reward signals for improving citation quality in LLMs without human annotations.
Unknown
Proposes an iterative self-training method using only synthetic data to improve LLM-as-a-Judge without human annotations.
Saurabh Dash, Yiyang Nan, John Dang, et al.
Aya Vision introduces open-weight multilingual VLMs for 23 languages using synthetic annotation and cross-modal model merging to retain text skills.
Peiyi Wang, Lei Li, Zhihong Shao, et al.
Math Shepherd automatically scores each reasoning step in math solutions using MCTS-inspired rollouts, enabling step-by-step PPO without human annotations.
Unknown
MAmmoTH2 introduces a scalable pipeline to harvest 10 million natural instruction-response pairs from web corpora, boosting LLM reasoning without manual annotation.
Unknown
Self-Instruct bootstraps instruction-following in language models using their own generations, reducing annotation needs.