Scalable Oversight for Superhuman AI via Recursive Self-Critiquing
Unknown
This paper proposes a recursive self-critiquing framework for scalable oversight of superhuman AI, showing promising results in experiments.
A comprehensive index of artificial intelligence and machine-learning research with AI-generated summaries, citation metrics, and direct links to papers and code.
Unknown
This paper proposes a recursive self-critiquing framework for scalable oversight of superhuman AI, showing promising results in experiments.
Unknown
This paper proposes collaborative multi-agent debate for scalable oversight, improving error detection in LLM responses.
Unknown
Introduces a principled benchmark for empirically evaluating scalable oversight protocols, enabling competitive comparison of methods for supervising AI systems.
Unknown
Introduces a principled benchmark for empirically studying and competitively evaluating scalable oversight protocols.
Unknown
This paper formalizes scalable oversight in hierarchical reinforcement learning, studying how to scale human feedback to complex tasks.
Unknown
This paper outlines a research paradigm for scalable oversight of large language models, based on Cotra's sandwiching approach, to improve model reliability.
Unknown
This paper explores scalable oversight protocols, specifically evaluating whether weak LLMs can effectively judge strong LLMs, and expresses optimism about debate as a scalable oversight method.
Unknown
This paper investigates scaling laws for scalable oversight, where weaker AI systems monitor stronger ones, to understand how oversight effectiveness scales with model capabilities.
Kiranjit Atwal
This paper provides a brief overview of AI applications in clinical nutrition and dietetics, highlighting opportunities and risks, and emphasizes the need for standardized protocols and human oversight.
Riikka Koulu
This paper critically examines human oversight in EU AI policy for legal decision-making, arguing it risks becoming an empty procedural safeguard without addressing inherent human limitations.
Manuel Cossio
This paper provides a comprehensive taxonomy of LLM hallucinations, arguing their theoretical inevitability and emphasizing the need for robust detection, mitigation, and human oversight.