Is AI clairvoyant? ChatGPT can make personality tests and predict responses, Israeli study finds
All users of ChatGPT, Gemini, and other artificial intelligence platforms know that these large language models (LLMs) are usually very “smart” and analyze texts and answer questions very well. But is AI clairvoyant?
Hebrew University of Jerusalem (HUJI) scientists have developed a method for generating personality assessment questionnaires with ChatGPT from any source text in English.
To test the method, they applied it to both the Diagnostic and Statistical Manual of Mental Disorders, 5th Edition (DSM-5) – the “bible” of mental health professionals published by the American Psychiatric Association to name, describe, and diagnose mental health and brain-related conditions – and, as a deliberately unconventional example, an astrology textbook.
The team said that not only could ChatGPT be used to create and validate these questionnaires, but that it could also accurately predict population-level responses before the surveys were carried out.
Testing AI with personality tests
The team – Dr. Rotem Monsa, Prof. Aviv Zohar, and Prof. Shahar Arzy of the Hebrew University-Hadassah Medical School and the HUJI Faculty of Computer Science published its findings in the Cell Press journal iScience under the title “Generating and analyzing personality questionnaires using large language models.”
ChatGPT and other publicly accessible LLMs are trained on the Internet by compiling trillions of human-language data inputs from websites and social media.
Because of this training, researchers have wondered whether LLMs understand human language at an expert level and whether personality is built into their algorithms.
“Since personality traits are reflected in language, LLMs may have learned the structure of human personality as a natural byproduct of their training,” suggested lead author Monsa, a postdoctoral researcher in HUJI’s medical neurosciences department, in an interview with The Jerusalem Post.
“So, while they were not taught specifically psychology or personality theories, these are already embedded in the language that LLMs learn from,” she continued.
To test how well LLMs can naturally assess human personality, the researchers said they used GPT-4 – a large multimodal model that, while less capable than humans in many real-world scenarios, exhibits human-level performance on various professional and academic benchmarks – to generate two personality assessment questionnaires.
They first used excerpts from the DSM-5 as source text, assuming that personality traits are localized on the personality disorder continuum.
ChatGPT created a questionnaire with statements based on descriptions of personality disorders from the DSM-5.
For example, based on the section describing paranoid personality disorder, the questionnaire asked participants to rank (1 = strongly disagree to 5 = strongly agree) how much they agree with statements such as “often suspects others’ motives” or “finds it easy to trust people.”
Another questionnaire used the astrology textbook as a control; it generated similar statements, but these were based on the source text’s assignment of personality traits to the astrological zodiac signs.
“We wanted to choose texts that describe human personality in very rich detail but also sit on opposite ends of a spectrum in terms of scientific grounding,” Monsa explained.
“The DSM-5 was refined through decades of clinical research and is the standard diagnostic manual in clinical psychiatry, known worldwide. The astrology text is very culturally based but not scientifically validated.”
After the LLM-based questionnaires were generated, they were presented to 600 people together with the Big Five personality questionnaire (BFI) – the current, most-validated personality questionnaire to assess their utility that is based on the hypothesis that personality traits are encoded in language.
The researchers were essentially asking whether an LLM has absorbed enough patterns about human psychology from language to predict how humans will answer questions it has just generated.
RESULTS FROM the DSM-5-sourced questionnaire showed high internal consistency within personality clusters – traits that typically correlate together in the real world, including dependency and avoidance, also correlated in their responses.
“Importantly, these results were just like those in the BFI, which helps validate the strength of the questionnaire in measuring real-life patterns of human psychology,” said Monsa.
“As predicted, the astrology questionnaire, by contrast, showed a weak internal consistency across traits. Our data suggest that the astrological elements don’t reflect coherent psychological dimensions.”
The most surprising result, however, was that ChatGPT could predict how participants would respond to the questionnaires before they were taken.
For both questionnaires, ChatGPT predicted in advance the mean responses and correlations between questions with high real-world accuracy, suggesting that LLMs have an innate understanding of personality dynamics at a population level, Monsa said.
Asked what it means to say that ChatGPT has an “understanding” of personality, she answered that “it was trained by a lot of text. It doesn’t understand by itself; it learns from statistical patterns.”
ChatGPT doesn’t have human understanding, she continued, “but even I – who knows Chat is not human – often call it he because the responses are very human-like. The LLMs have learned a great deal in the last three years since we started the research.”
Asked how ChatGPT could predict people’s responses before seeing their answers, Monsa replied that “it wasn’t one person, but a population-level response. We were very surprised that Chat performed so well.
“When we started, the models were less advanced. Obviously, the model wasn’t ‘born’ with psychological knowledge. It wasn’t trained in it. It became ‘natural’ or ‘innate’ from what it learned.”
The team was able to prevent ChatGPT from accessing the actual participants’ responses before making its predictions because the questionnaires didn’t exist on the Internet, which is the source of its raw material, Monsa noted.
“We ran the prediction multiple times to see whether ChatGPT produced essentially the same predictions, and they were very similar each time,” she said.
An astrology questionnaire was used as a control because it was “a negative control. Personality can’t be determined by astrology.”
“If we fed ChatGPT another pseudoscientific text – such as a book claiming that handwriting accurately reveals personality – it would work successfully like our research. We tried other things that didn’t involve personality, like menus and how well your kitchen was organized,” Monsa recalled.
The study using LLM-generated personality assessment could eventually be used to assess human patients.
“It will take time, and there are risks,” said Monsa, “and it has to be analyzed by psychologists, but I’m sure it’ll eventually be used for diagnosis or even treatment after there is psychological validation. It would be much cheaper and be especially helpful in places where trained psychologists are scarce.”
She conceded that an LLM could inadvertently generate questions that are culturally biased, stigmatizing, or psychologically misleading because they were trained on English-language, Western material.
“Our results may not be relevant if these methods are repeated in different languages and cultures because LLMs were trained mostly on English texts in occidental cultures,” said Monsa.
“But if they were applied to Hebrew-speaking Israelis, it should be close, because Israelis are very Western. An LLM could actually learn culturally specific concepts of personality that human psychologists might overlook.”
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.