ImageNet classification with deep convolutional neural networks
Alex Krizhevsky, Ilya Sutskever et al.
1.8k
Citations
125
Influential Citations
PLoS ONE
Venue
2013
Year
We analyzed 700 million words, phrases, and topic instances collected from the Facebook messages of 75,000 volunteers, who also took standard personality tests, and found striking variations in language with personality, gender, and age. In our open-vocabulary technique, the data itself drives a comprehensive exploration of language that distinguishes people, finding connections that are not captured with traditional closed-vocabulary word-category analyses. Our analyses shed new light on psychosocial processes yielding results that are face valid (e.g., subjects living in high elevations talk about the mountains), tie in with other research (e.g., neurotic people disproportionately use the phrase 'sick of' and the word 'depressed'), suggest new hypotheses (e.g., an active life implies emotional stability), and give detailed insights (males use the possessive 'my' when mentioning their 'wife' or 'girlfriend' more often than females use 'my' with 'husband' or 'boyfriend'). To date, this represents the largest study, by an order of magnitude, of language and personality.
This 2013 study by Schwartz et al. is a landmark in computational social science, demonstrating that the language people use on social media can reliably predict their personality, gender, and age. By analyzing an unprecedented 700 million words from 75,000 Facebook users, the authors moved beyond small-scale lab studies to a large-scale, ecologically valid analysis of natural language. The open-vocabulary approach—letting the data reveal linguistic patterns rather than imposing predefined categories—was a methodological breakthrough that has since become standard in psycholinguistic and NLP research.
The paper's findings have practical implications for personalized marketing, mental health screening, and user modeling. For AI practitioners, it showed that simple word and phrase frequencies can capture complex psychological constructs, paving the way for modern personality-aware AI systems. The study also highlighted the importance of big data in understanding human behavior, setting a precedent for later work on sentiment, emotion, and demographic inference from text.
This paper fundamentally changed how researchers study personality and language, shifting from small-scale experiments to large-scale observational studies. It demonstrated that machine learning and natural language processing can extract meaningful psychological signals from noisy social media text. The open-vocabulary methodology has been widely adopted in computational social science, marketing analytics, and AI-driven mental health tools. For AI practitioners, the paper underscores the value of data-driven feature engineering and the importance of large, diverse datasets for training robust models. It also raised ethical considerations about privacy and inference of sensitive traits from public text, a topic that remains highly relevant today.
Alex Krizhevsky, Ilya Sutskever et al.
Ashish Vaswani, Noam Shazeer et al.
Douglas M. Bates, Martin Mächler et al.
Diederik P. Kingma, Jimmy Ba