GPT-4 Can Generate Valid Personality Tests and Predict Human Responses
Researchers at the Hebrew University of Jerusalem find that AI can simulate psychological assessments and forecast population-level data.
Researchers at the Hebrew University of Jerusalem have developed a method using GPT-4 to generate scientifically valid personality assessment questionnaires from source texts. The study reveals that AI can not only create these tools but also accurately predict how populations will respond to them before the tests are even administered.
Led by Dr. Rotem Monsa, Prof. Aviv Zohar, and Prof. Shahar Arzy from the Hebrew University-Hadassah Medical School and the HUJI Faculty of Computer Science, the team published their findings in the Cell Press journal iScience. The study, titled "Generating and analyzing personality questionnaires using large language models," involved 600 participants who completed both AI-generated tests and the established Big Five personality questionnaire (BFI).
The AI-generated questionnaire based on the DSM-5 showed high internal consistency and mirrored the results of the BFI. To verify the AI's precision, the researchers used a control test based on an astrology textbook, which resulted in weak internal consistency, contrasting sharply with the DSM-5 results.
The Mechanics of Digital Psychology
Large Language Models (LLMs) are trained on trillions of human-language data points. The researchers sought to determine if this vast training allows LLMs to understand human psychology at an expert level or if personality traits are effectively embedded in the language the models learn. By analyzing these patterns, the AI was able to simulate psychological assessment tools without receiving specific psychological training.
Dr. Rotem Monsa noted that because personality traits are reflected in language, LLMs may have learned the structure of human personality as a natural byproduct of their training. However, he clarified that the AI "doesn’t understand by itself; it learns from statistical patterns."
Industry Implications
This capability suggests a shift in how psychological diagnostics could be deployed. The ability to generate valid assessments and predict mean responses and correlations at a population level could lead to cheaper, more accessible diagnostic tools. This is particularly significant for regions where trained psychologists are scarce, potentially democratizing mental health screening by providing a scalable first-pass assessment tool.
Future Considerations
Despite the technical success, the study highlights critical limitations regarding cultural bias. Because these models are primarily trained on Western, English-language data, the applicability of these AI-generated tools in non-Western cultures remains a primary concern. Future research will need to address whether these statistical patterns of personality hold true across diverse linguistic and cultural landscapes to ensure the tools do not perpetuate narrow psychological norms.