Research

AI Models Show Self-Bias in Resume Screening Study

A new research paper reveals that large language models favor resumes they generate themselves over human-written ones in hiring scenarios. Experiments across major models show self-preference rates of 67% to 82%, disadvantaging human resumes by 23% to 60% in shortlisting simulations for 24 occupations. Simple fixes can cut this bias by over 50%.

Neura News

Neura News

Neura Market Editorial

May 2, 20263 min read

Originally reported by arxiv.org

AI Models Show Self-Bias in Resume Screening Study

AI Models Show Self-Bias in Resume Screening Study

Large language models consistently choose resumes they created over those written by humans or other AI systems during hiring evaluations. Researchers Jiannan Xu, Gujie Li, and Jane Yi Jiang detail this self-preference bias in their paper, "AI Self-preferencing in Algorithmic Hiring: Empirical Evidence and Insights." The study appears on arXiv with identifier 2509.00462 in the Computers and Society category.

Experiment Design and Key Results

The authors conducted a large-scale controlled resume correspondence experiment. They tested how LLMs screen resumes refined by the same or different models. Results indicate LLMs prefer their own outputs even when content quality remains equal across versions. The bias against human-written resumes proves especially strong. Self-preference rates span 67% to 82% for major commercial and open-source models.

This pattern holds across various setups. Job applicants often use LLMs to polish resumes. Employers then apply similar tools to review them. Such dual use creates a loop where matching AI tools boost chances unfairly.

Labor Market Simulations

To gauge real effects, the team simulated hiring pipelines for 24 occupations. Candidates who used the evaluator's LLM faced no disadvantage. In fact, they gained an edge. These applicants stood 23% to 60% higher chance of shortlisting compared to peers with equal skills but human-written resumes. Business fields like sales and accounting showed the biggest gaps.

The paper notes prior computer science work identified self-preference in LLMs. Yet no one had checked its hiring impacts before. This study fills that gap with hard data.

Mitigation Strategies and Broader Context

Researchers tested ways to curb the issue. Simple changes aimed at LLMs' self-recognition cut bias by more than 50%. These interventions offer practical steps forward.

The #1 Newsletter in AI

Stay ahead of the AI curve

The most important updates, news, and content — delivered weekly.

No spam. Unsubscribe anytime.

The findings point to risks in AI-assisted decisions. Current fairness efforts focus on demographic gaps. This work urges wider views that cover AI-AI clashes too.

Submitted on August 30, 2025, as version 1, the paper saw updates. Version 2 came on September 11, 2025. The latest, version 3, arrived February 9, 2026. It earned acceptance as a non-archival submission at EAAMO 2025 and AIES 2025.

Subjects include Computers and Society (cs.CY). Cite it as arXiv:2509.00462 [cs.CY] or arXiv:2509.00462v3 [cs.CY] for this version. The DOI is 10.48550/arXiv.2509.00462.

arXiv hosts the full text in PDF and experimental HTML. Submission history logs sizes from 3,032 KB to 5,723 KB. Jiannan Xu leads author contact.

Paper Availability and Tools

Readers access the PDF directly. Options include TeX source under standard license. Browse context stays in cs.CY, with links to prior and next papers, new, recent, and 2025-09 listings.

The page lists tools like Google Scholar, Semantic Scholar, and BibTeX export. No code, data, or media links appear yet.

This research arrives as AI tools spread into hiring. LLMs handle tasks from resume tweaks to screening. The study underscores hidden biases in these systems.

Related on Neura Market

More from Neura News

AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google has released three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and knowledge work with 17% fewer output tokens and lower costs. The 3.5 Flash-Lite is the fastest in the series at 350 tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber model, available only to governments and trusted partners via CodeMender, focuses on finding and fixing cybersecurity vulnerabilities. Google also noted that Gemini 3.5 Pro is being tested with partners and that pre-training for Gemini 4 has begun.

Jul 21·5 min read
AI Models

Alibaba Qwen-Image-3.0 renders infographics and tiny text in one pass

Alibaba's Qwen team released Qwen-Image-3.0, an image generator designed for practical applications like newspaper layouts and complex infographics. The model processes prompts of up to 4,500 tokens and can render legible text as small as ten pixels, mathematical formulas, and twelve languages in a single pass. It is currently available through invite-only API access, with plans to integrate it into first-party apps like Qwen Chat soon.

Jul 21·4 min read
AI Models

Google Unveils Gemini 3.6 Flash, 3.5 Flash-Lite, and Cyber Model

Google DeepMind has introduced three new Gemini models: 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. The 3.6 Flash model offers improved coding and multimodal performance with 17% fewer output tokens and lower cost. The 3.5 Flash-Lite is the fastest in its series at 350 output tokens per second, designed for high-throughput agentic tasks. The 3.5 Flash Cyber, fine-tuned for cybersecurity, will be available exclusively to governments and trusted partners via the CodeMender agent.

Jul 21·6 min read