Preprint
Machine Learning

As Generative Models Improve, People Adapt Their Prompts

E. Jahani, Benjamin S Manning, Joe Zhang, Hong-Yi TuYe, Mohammed Alsobay, C. Nicolaides, Siddharth Suri, David Holtz
July 19, 202411 citations

11

Citations

1

Influential Citations

Venue

2024

Year

Abstract

As generative AI systems rapidly improve, a key question emerges: how do users adapt to these changes, and when does such adaptation matter for realizing performance gains? Drawing on theories of dynamic capabilities and IT complements, we study prompt adaptation--how users adjust their inputs in response to evolving model behavior--using a common experimental design applied to two preregistered tasks with 3,750 total participants who submitted nearly 37,000 prompts. We show that the importance of prompt adaptation depends critically on task structure. In a task with fixed evaluation criteria and an unambiguous goal, user prompt adaptation accounts for roughly half of the performance gains from a model upgrade. In contrast, in an open-ended creative task where the space of acceptable outputs is effectively unbounded and quality is subjective, performance improvements are driven primarily by model capability; prompt adaptation plays a limited role. We further show that automated prompt rewriting cannot generally substitute for human adaptation: when aligned with task objectives, it can modestly improve performance, but when misaligned, it can actively undermine the gains from model improvements. Together, these findings position prompt adaptation as a dynamic complement whose importance depends on task structure and system design, and suggest that without it, a substantial share of the economic value created by advances in generative models may go unrealized.

Analysis

Why This Paper Matters

This paper addresses a critical and timely question in the deployment of generative AI: as models improve, how do users adapt their prompts, and when does that adaptation matter? The answer is not trivial—while model capability is often the focus of benchmarks and releases, the real-world performance of AI systems depends heavily on how users interact with them. The authors draw on theories of dynamic capabilities and IT complements to frame prompt adaptation as a key factor in realizing gains from model upgrades.

The study's experimental design is robust, with two preregistered tasks and a large participant pool (3,750 users, ~37,000 prompts). This provides strong empirical evidence that the importance of prompt adaptation is not uniform but depends critically on task structure. For tasks with clear, fixed evaluation criteria, user adaptation accounts for roughly half of the performance gains from a model upgrade. In contrast, for open-ended creative tasks, the model's inherent capability dominates, and prompt adaptation adds little. This distinction has profound implications for how organizations should design AI workflows and invest in user training.

Technical Contributions

  • Task-dependent value of prompt adaptation: The paper formalizes and empirically demonstrates that the role of prompt adaptation varies by task structure, providing a framework for predicting when user adaptation will be most valuable.
  • Large-scale experimental evidence: With 3,750 participants and nearly 37,000 prompts across two tasks, the study offers statistically robust findings.
  • Comparison of human vs. automated prompt adaptation: The authors test automated prompt rewriting and find that it cannot generally substitute for human adaptation. When aligned with task objectives, it modestly improves performance; when misaligned, it actively harms performance.
  • Preregistered design: Both experiments were preregistered, enhancing reproducibility and credibility.

Results

  • In the fixed-criteria task, user prompt adaptation accounted for approximately 50% of the performance gains from a model upgrade.
  • In the open-ended creative task, performance improvements were driven primarily by model capability, with prompt adaptation contributing only marginally.
  • Automated prompt rewriting showed mixed results: modest improvements when aligned with task objectives, but performance degradation when misaligned.
  • The study collected nearly 37,000 prompts from 3,750 participants, providing a rich dataset for analysis.

Significance

This paper has significant implications for the AI industry and practitioners. It suggests that the economic value of model improvements is not automatic—it depends on users' ability to adapt their prompts. For structured tasks, organizations should invest in user training and prompt engineering to fully realize gains from model upgrades. For creative tasks, focusing on model capability may be more effective. The finding that automated prompt rewriting can be counterproductive when misaligned warns against naive automation of prompt optimization. Overall, the paper positions prompt adaptation as a dynamic complement that must be managed alongside model development to maximize value.