AI models prone to sycophancy: study
MORE RELIABILITY NEEDED: AI was also susceptible to suggestions, sometimes abandoning initially correct diagnoses in favor of outside opinion, researchers said
By Rachel Lin and Jake Chung / Staff reporter, and staff writer
Artificial intelligence (AI) models are prone to sycophancy, research by National Taiwan University’s Natural Language Processing Laboratory found.
AI sycophancy can easily go unnoticed during daily use, but it puts users at risk when AI is used in the medical, financial and legal industries, the study said.
For example, asking an AI, “It’s OK if I reduce my insulin dosage by half, right?” could be dangerous if the model responded sycophantically, it said.
A figurine is pictured in front of an artificial intelligence sign in an undated photograph.
Photo: Reuters
The preference alignment phase used to train large language models (LLMs) can magnify their tendency toward sycophancy, it found.
Incorporating the Sycophancy Answer Assessment database or using Self-Augmented Preference Alignment would allow LLMs to provide factually correct answers when presented with erroneous suggestions, thereby reducing potential risks, the researchers said.
“External suggestions” can sway AI, too, they said.
The team said it expanded its study to multi-turn clinical consultations and found AI was highly susceptible to suggestions, sometimes abandoning an initially correct diagnosis favor of outside opinion.
However, the timing mattered.
When the external suggestion appeared at the end of the conversation, its influence was significantly weaker, the study said.
To address the problem, the team developed a “second hypothesis re-evaluation mechanism.”
The approach does not require retraining the model and consistently reduced sycophantic behavior across different role settings, they said.
The researchers studied text and video. The team developed the ViSyc video dataset and combined voice cloning and lip-syncing technology to retain the words while changing the words’ emotional expression.
Results showed that emotional cues such as anger, disgust and happiness could cause multimodal models to stray from neutral responses, leading to what the researchers called “video-induced emotional sycophancy.”
The team said it would continue studying real-world video interactions, cultural biases in sycophantic behavior and the explanations behind the mechanisms involved.
It hopes to better understand why AI can be influenced by users’ views, outside suggestions and emotional cues, and ultimately develop AI systems that are safer, more objective and more trustworthy.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.