A new study published in the academic journal Judgment and Decision Making found that readers could not distinguish between AI-generated and human-written short stories better than chance. When authorship was unknown, participants rated the AI stories higher on perceived quality and immersion. But when they were told a machine wrote them, those ratings dropped.
The research, co-authored by Sydney Sears and Deena Skolnick Weisberg, ran three experiments with more than 2,500 total participants. The findings suggest that people's bias against AI may be stronger than their ability to actually spot it.
The First Experiment: 1,682 Readers, Six Stories
The first experiment involved 1,682 participants. Each person read one of six short stories, each about 1,000 words long. Three stories came from well-known literary magazines and short story collections. The other three were generated using ChatGPT 4.0, with prompts based on the theme, style, and narrative perspective of the human originals.
Half of the participants were told the story was written by a human. The other half were told it came from ChatGPT. The authorship information given was accurate for only half the participants in each group.
The results were striking. AI-generated stories received a mean quality score of 1.54 on a scale from -3 to +3. Human-written stories received a mean quality score of 0.97. On immersion, AI stories scored 1.42, while human stories scored 1.00.
Participants also gave higher scores when they were told a human was the author, regardless of the actual origin. That bias flipped among AI-skeptical participants, who rated stories lower when told they were AI-written.
Direct Comparison Didn't Help
The two additional experiments had 905 total participants. In these, each person read both a human-written and an AI-generated story and had to identify which was which. Even with direct comparison, participants performed no better than chance.
Self-reported experience with AI systems correlated positively with the ability to correctly identify story origins. Self-reported experience with fiction did not help participants tell the stories apart.
Participants with a positive attitude toward AI gave higher ratings overall. They gave even higher scores when told the story came from ChatGPT.
Stay ahead of the AI curve
The most important updates, news, and content — delivered weekly.
No spam. Unsubscribe anytime.
Why AI Stories Might Score Higher
The study's authors suggest that higher ratings don't necessarily mean better writing. AI-generated texts tend to be smoother, easier to read, and more emotionally upbeat than human-written texts. High-quality literary fiction is often intentionally hard to access and pushes readers to work for meaning.
People tend to prefer material that is easier to process. A story can be high quality but not very engaging, and vice versa. The short story format likely works in AI's favor, since telling a coherent story in 1,000 words is different from doing so across hundreds of pages.
All data and materials from the study are freely available on the Open Science Framework.
Earlier Research Points the Same Way
A prior study from October 2025 by Stony Brook University and Columbia Law School examined reader expertise. With simple prompts, professional readers clearly preferred human-written texts. But when models were trained on individual authors' styles, experts preferred AI-generated texts eight times more often for style imitation and twice as often for writing quality.
An earlier study on AI-generated poems found a similar bias to the one in this study. The pattern is consistent: people perceive AI-generated creative works as at least on par with human work, yet they don't believe AI is capable of that.
The Bottom Line
The study suggests that AI can generate creative works people perceive as at least on par with human work. The bias against AI appears to be a matter of belief, not detection. Readers can't tell the difference, but they still judge the machine harshly when they know it's there.

