Preprint
AI Safety & Alignment

Beyond Preferences in AI Alignment: T. Zhi-Xuan et al.

January 1, 2025

0

Citations

0

Influential Citations

Venue

2025

Year

Abstract

… The dominant practice of AI alignment assumes (1) that … we term a preferentist approach to AI alignment. In this paper, we … a reframing of the targets of AI alignment: Instead of alignment …

Analysis

Why This Paper Matters

This paper challenges a foundational assumption in AI alignment: that aligning AI systems to human preferences is the ultimate goal. The 'preferentist' approach, as the authors call it, underpins many current alignment techniques, from RLHF to preference-based reward modeling. By questioning this assumption, the paper opens up a critical discussion about whether preferences are the right target for alignment, especially as AI systems become more capable and their impacts more profound.

The significance lies in its potential to reshape the alignment research agenda. If preferences are not the sole or even primary target, then alignment methods may need to incorporate other normative dimensions, such as human rights, fairness, or long-term societal well-being. This could lead to more holistic alignment strategies that are less susceptible to the pitfalls of preference aggregation, such as manipulation or gaming of preference models.

Technical Contributions

The paper's main contribution is conceptual rather than technical. It provides a clear articulation of the preferentist assumption and its limitations. It then proposes a reframing that broadens the targets of alignment beyond preferences. This includes:

  • A critique of preference-based alignment's reliance on revealed or stated preferences, which may be inconsistent, adaptive, or influenced by the AI itself.
  • A suggestion that alignment should consider 'normative' targets, such as ethical principles or societal values, which may be more stable and principled.
  • A framework that could accommodate multiple alignment targets, potentially leading to multi-objective alignment approaches.

Results

As a theoretical paper, there are no empirical results or metrics. The 'results' are the conceptual arguments and the proposed reframing. The paper's impact will be measured by its influence on subsequent research and discourse in AI alignment.

Significance

The broader impact of this paper is to encourage the AI community to think more deeply about what we are aligning AI to. By moving beyond preferences, alignment could become more robust against issues like preference change, moral uncertainty, and the potential for AI to shape human preferences. This could lead to AI systems that are not only more aligned with what we say we want, but also with what we truly value, even as those values evolve. The paper is a timely contribution to the growing field of AI safety and alignment, offering a philosophical foundation for future technical work.