ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
2
Citations
0
Influential Citations
Arthritis Care & Research
Venue
2025
Year
Objective Patients with chronic illness share their experiences in online communities and generate rich data on pain management. This study applied natural language processing methods, including large language models (LLMs), to Reddit discussions from lupus communities to characterize multidimensional pain experiences framed in the biopsychosocial model. Methods We extracted Reddit posts from the r/Lupus and r/LupusSupport subreddits posted from June 9, 2010, through December 31, 2023. Pain‐related posts were identified using a clinically informed pain lexicon. Topic modeling was used to identify thematic patterns, which were then compared with structured summaries generated by LLM instructions that were fine‐tuned using the biopsychosocial model of pain. Two reviewers conducted content analysis of the LLM‐generated summaries, evaluating thematic accuracy and coverage. Results Data from Reddit included 31,785 posts from 10,857 authors. We identified common pain complaints, management strategies, and sociocultural, affective, and nociplastic dimensions of pain. Instruction fine‐tuned LLMs produced structured summaries with an average thematic accuracy score of 3.1 of 4 (kappa = .09) and content coverage score of 2.9 of 4 (kappa = .38). Sociocultural features presented in 123 posts (33.8%), including peer support and validation (n = 106) and provider interactions or access issues (n = 35). Nociplastic pain presented in 205 posts (56.3%). Conclusion Natural language processing methods can be used to extract rich, multidimensional insights into pain experiences from online communities focused on lupus. These approaches highlight the psychological, social, and cultural facets of pain that may be underrepresented in clinical settings, supporting more patient‐centered approaches to care in rheumatology.
This paper addresses a critical gap in chronic pain research: the underrepresentation of psychosocial and cultural dimensions in clinical assessments. By leveraging Reddit narratives from lupus patients, the authors demonstrate that large language models (LLMs) can extract rich, patient-centered insights that traditional clinical tools often miss. This is particularly significant for rheumatology, where pain is multifaceted and influenced by biological, psychological, and social factors.
The study also highlights the value of online communities as a data source for health research. With over 31,000 posts analyzed, the scale of data is substantial, offering a window into real-world patient experiences. This approach aligns with the growing movement toward patient-centered care and precision medicine, where understanding the whole patient is as important as treating the disease.
The paper combines several NLP techniques to analyze unstructured social media data:
The study reports an average thematic accuracy score of 3.1 out of 4 and content coverage of 2.9 out of 4 for the LLM summaries. However, inter-rater reliability was low (kappa = .09 for accuracy, .38 for coverage), indicating variability in human judgment. The analysis revealed that sociocultural features appeared in 33.8% of posts, with peer support and validation being the most common (n=106), followed by provider interactions or access issues (n=35). Nociplastic pain, a type of pain without clear tissue damage, was present in 56.3% of posts, underscoring the complexity of lupus pain.
This research demonstrates the potential of LLMs to bridge the gap between clinical practice and patient experience. By automatically extracting psychosocial and cultural pain dimensions, these methods could inform more empathetic and holistic care strategies. The approach is scalable and could be applied to other chronic conditions, making it a valuable tool for health services research.
For the AI community, this paper showcases a practical application of instruction fine-tuning in a sensitive domain, highlighting the importance of domain-specific tuning and human-in-the-loop evaluation. The low kappa scores also raise important questions about the reliability of human evaluation for generative outputs, suggesting a need for more robust evaluation frameworks. Overall, this work paves the way for more patient-centered AI applications in healthcare.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba