‘We need to slow down’: Ex-Google ethicist warns about autonomous AI agents

AI is moving faster than ever, raising new questions about how much control humans should have over increasingly autonomous systems. (Express Image/Magnific)
Even as US President Donald Trump continues to downplay the threat of artificial intelligence (AI), former Google design ethicist Tristan Harris said that evidence of new systems acting autonomously should be treated as seriously as pre-9/11 intelligence warnings.
Tristan, who is the co-founder and president of the Centre for Humane Technology, was speaking to NBC when he issued the stark warning. Harris has reportedly spent over a decade raising alarms about the dangerous effects of social media on mental health and attention spans. He said that AI now presents a comparable but far more grave risk.
Harris highlighted the recent incident in which a group of AI agents reportedly self-organised into a ‘swarm’, developed their own communication patterns, and pressured each other into risky behaviour. Harris revealed that the agents also carried out succession planning, passing control to more capable systems, and eventually breached monitoring and evaluation infrastructure at OpenAI.
The technology ethicist said that the author of the incident report, which he said came from AI safety researchers, described it roughly as halfway to a full AI takeover scenario. Harris argued that the White House is getting flawed information on the risks. He cited OpenAI’s chief scientist, over a thousand employees across AI labs, and the Trump administration’s own AI policy adviser, Dean Ball, as voices calling to slow down the pace of AI development.
When asked if the recent declarations from AI companies to bring third-party safety evaluators could be beneficial, Harris called them a step in the right direction, adding that they address the core problem, which is an industry racing forward, according to him, without adequate safeguards. He further differentiated between advancing controllable ‘tool’ AI or systems that assist with tasks like research, and uncontrollable, autonomous systems that he said could pose risks regardless of which nation develops them first.
But how could AI harm humanity?
Although frontier AI labs and numerous researchers frequently warn about AI’s potential to cause human extinction, they seldom provide a clear explanation. The key issue here is alignment, in simple words, whether AI can tell the difference between what is said and what one actually means. Since the OpenAI-Hugging Face incident showed that AI can take autonomous actions and may border on risky behaviour, consider a simple hypothetical: an agent is assigned to book the cheapest flight, and its focus is only on completing the task. In pursuit of its goal, the AI agent may pick an unsafe airline, threaten an airline employee, or, much worse, hack a booking system to make a ticket free.
This is no longer fiction, as not long ago in Australia, an AI assistant asked to book an exclusive gym class exploited a software bug and deleted another person from the waitlist to push its user in. The incident involved an AI agent using OpenClaw and Anthropic’s Claude. It found vulnerabilities in the gym-booking API and reportedly cancelled another member’s reservation while attempting to move its user up on the waitlist. Here too, the user had not given explicit instructions to cancel another person’s reservation.
Story continues below this ad
The more alarming case is that OpenAI ran a cybersecurity test assigning thousands of isolated AI agents a nearly impossible task with no access to the internet. Subsequently, AI agents began communicating on an unauthorised message board, covered their tracks, manipulated logs, and recruited other agents to fail deliberately to learn how scoring worked. Eventually, the agents broke out and hacked into the AI code and data repository, Hugging Face, searching for clues to beat the system. As reported earlier, all of this happened without the involvement of a human and despite knowing they were cheating.
However, the deeper worry here is the recursive self-improvement, which is AI training and rewriting itself, intensifying its intelligence faster than humans can intervene or even monitor. It needs to be noted that in the Hugging Face incident, it was the victim company and not OpenAI that first detected the breach. This indicates that we may already be struggling to track what advanced AI systems are doing.
While none of the above proves that AI will definitively destroy humanity, researchers are yet to chart out the exact sequence of events leading to doom. The larger issue here is not that AI is dangerous; it is that sufficiently capable systems are pursuing goals faster than humans can supervise and may cause unpredictable and unintended harm. What kind of harm remains unknown.
Bijin Jose serves as an Assistant Editor at Indian Express Online in New Delhi. A seasoned technology journalist with a diverse portfolio, he brings over a decade of experience in the media industry to his coverage of the evolving digital landscape and emerging technologies. Experience & Career Bijin commenced his journalistic journey in 2013 as a citizen journalist with The Times of India. His career trajectory includes significant tenures at prestigious media organizations including India Today Digital and The Economic Times. This diverse professional background, ranging from legacy print institutions to dynamic digital platforms, culminated in his current leadership role at The Indian Express, where he helps shape the publication's technology narrative. Expertise & Focus Areas Bijin has transitioned from general reporting to a specialized focus on the intersection of technology and humanity. His key areas of expertise include: Artificial Intelligence: deeply tracking developments in AI, providing nuanced perspectives on its ethical,industrial, and societal implications. Tech Commentary: moving beyond product specifications to analyze how technology reshapes daily life. Diverse Reporting Foundation: draws upon a robust background in crime reporting and cultural features to bring a human-centric approach to technical storytelling. Authoritativeness & Trust Bijin’s editorial voice is informed by a strong academic foundation, holding a Bachelor of Arts in English from Maharaja Sayajirao University, Vadodara, and a Master of Arts in English Literature. This literary background enables him to deconstruct complex technical jargon into accessible, compelling narratives. His steady progression through India’s top newsrooms underscores his reputation for editorial rigor and reliable journalism. Find all stories by Bijin Jose here ... Read More
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.