Preprint
Computer Vision

Enhancing university level English proficiency with generative AI: Empirical insights into automated feedback and learning outcomes

Sumie Tsz Sum Chan(English Language Teaching Unit, The Chinese University of Hong Kong, Hong Kong, CHINA), Noble Po Kan Lo(Department of Educational Research, Lancaster University, Lancaster, UNITED KINGDOM), Alan Man Him Wong(English Language Teaching Unit, The Chinese University of Hong Kong, Hong Kong, CHINA)
November 13, 2024Contemporary Educational Technology47 citations

47

Citations

4

Influential Citations

Contemporary Educational Technology

Venue

2024

Year

Abstract

This paper investigates the effects of large language model (LLM) based feedback on the essay writing proficiency of university students in Hong Kong. It focuses on exploring the potential improvements that generative artificial intelligence (AI) can bring to student essay revisions, its effect on student engagement with writing tasks, and the emotions students experience while undergoing the process of revising written work. Utilizing a randomized controlled trial, it draws comparisons between the experiences and performance of 918 language students at a Hong Kong university, some of whom received generated feedback (GPT-3.5-turbo LLM) and some of whom did not. The impact of AI-generated feedback is assessed not only through quantifiable metrics, entailing statistical analysis of the impact of AI feedback on essay grading, but also through subjective indices, student surveys that captured motivational levels and emotional states, as well as thematic analysis of interviews with participating students. The incorporation of AI-generated feedback into the revision process demonstrated significant improvements in the caliber of students’ essays. The quantitative data suggests notable effect sizes of statistical significance, while qualitative feedback from students highlights increases in engagement and motivation as well as a mixed emotional experience during revision among those who received AI feedback.

Analysis

Why This Paper Matters

This paper addresses a critical gap in the application of generative AI to education: rigorous, large-scale empirical evidence of its effectiveness. While many studies speculate on AI's potential, this work delivers a randomized controlled trial with 918 students, providing concrete data on how LLM-based feedback impacts writing proficiency, engagement, and emotional responses. For AI practitioners, it validates that models like GPT-3.5-turbo can serve as effective automated tutors, not just content generators.

The focus on both quantitative outcomes (essay grades) and qualitative experiences (motivation, emotions) offers a holistic view of AI's role in learning. This is particularly relevant as educational institutions grapple with integrating AI tools without undermining student development. The mixed emotional findings also caution against over-optimism, highlighting the need for thoughtful deployment.

Technical Contributions

  • Large-scale RCT design: The study uses a rigorous experimental setup with 918 participants, ensuring statistical power and causal inference.
  • Multi-modal assessment: Combines statistical analysis of essay scores with student surveys (motivation, emotional states) and thematic analysis of interviews, providing a comprehensive evaluation.
  • Use of GPT-3.5-turbo: Demonstrates the practical application of a widely available LLM for automated feedback, making the approach accessible to many institutions.
  • Focus on revision process: Unlike many studies that assess one-shot AI use, this work examines iterative revision, a key learning activity.

Results

The paper reports statistically significant improvements in essay quality for students receiving AI feedback, with notable effect sizes (exact values not provided in abstract). Quantitative data shows clear grading improvements. Qualitatively, students reported higher engagement and motivation, but also a mixed emotional experience—some found AI feedback helpful, others felt anxious or overwhelmed. This nuanced result is important for designing supportive AI systems.

Significance

For the AI field, this paper provides a template for evaluating generative AI in real-world educational settings. It moves beyond hype to evidence, showing that LLMs can enhance learning outcomes when used appropriately. The mixed emotional findings underscore the need for human-in-the-loop systems that balance automation with empathy. As AI becomes ubiquitous in education, studies like this guide practitioners in building tools that are both effective and emotionally considerate.