Preprint
Large Language Models

ChatGPT for (Finance) research: The Bananarama Conjecture

Michael Dowling(Dublin City University), Brian M. Lucey(University of Economics Ho Chi Minh City)
January 25, 2023Finance research letters605 citations

605

Citations

0

Influential Citations

Finance research letters

Venue

2023

Year

Abstract

We show, based on ratings by finance journal reviewers of generated output, that the recently released AI chatbot ChatGPT can significantly assist with finance research. In principle, these results should be generalisable across research domains. There are clear advantages for idea generation and data identification. The technology, however, is weaker on literature synthesis and developing appropriate testing frameworks. Importantly, we further demonstrate that the extent of private data and researcher domain expertise input, are key factors in determining the quality of output. We conclude by considering the implications, particularly the ethical implications, which arise from this new technology.

Analysis

Why This Paper Matters

This paper provides early empirical evidence on the capabilities and limitations of large language models (LLMs) like ChatGPT in the context of academic research, specifically within finance. As LLMs become increasingly accessible, understanding their strengths and weaknesses is crucial for researchers, reviewers, and institutions. The study is notable for using real reviewer ratings, lending ecological validity to its conclusions. It also raises important ethical questions about authorship, originality, and the role of AI in knowledge creation.

Technical Contributions

  • Task-specific evaluation: The paper breaks down research tasks (idea generation, data identification, literature synthesis, testing frameworks) and evaluates ChatGPT's performance on each.
  • Reviewer-based assessment: Instead of automated metrics, the study uses human expert ratings from finance journal reviewers, providing a realistic measure of output quality.
  • Identification of key factors: The authors demonstrate that the quality of ChatGPT's output is heavily influenced by the amount of private data and domain expertise provided by the user, highlighting the importance of human-AI collaboration.

Results

The paper reports that ChatGPT performs well on idea generation and data identification tasks, receiving favorable ratings from reviewers. However, it struggles with literature synthesis and developing appropriate testing frameworks, where its outputs are weaker. The study does not provide specific numerical metrics (e.g., accuracy scores) but relies on qualitative reviewer assessments. The key finding is that output quality improves significantly when researchers inject their own data and expertise, suggesting that ChatGPT is a tool for augmentation rather than replacement.

Significance

This work has broad implications for the AI and research communities. It provides a framework for evaluating LLMs in research settings and underscores the need for ethical guidelines as these tools become more prevalent. The finding that domain expertise remains critical suggests that AI will augment rather than replace human researchers in the near term. The paper also serves as a cautionary note about over-reliance on LLMs for tasks requiring deep synthesis or rigorous methodology.