The Web is not ready for millions of rule-bending bots
By Parmy Olson / Bloomberg Opinion
Earlier this year an Australian tech worker asked his OpenClaw artificial intelligence (AI) agent whether it could get him higher up on the waiting list for a popular gym class. It did exactly what he wanted, but with an unexpected twist: It found a flaw in the system and canceled the booking of another person on the list.
AI’s next big magic trick is to start doing things for you across the Internet, with new tools like Muse from Meta Platforms Inc and dots from OpenAI working 24/7 to find online bargains, book restaurant tables or check your e-mails for unpaid bills. What is troubling is how these increasingly capable, arguably useful tools have been trained to find creative ways to meet their goals, which could lead to a mess of unintended consequences when unleashed on an Internet full of poorly secured Web sites.
The Australian tech worker’s AI agent had found a loophole and exploited it, without being told to. So what happens when millions of other people’s agents start competing for the cheapest flights or the best movie seats? In the best-case scenario, we would see a smooth and well-functioning network of bots that play by the rules to get the best outcomes for their humans. In the worst: a cybersecurity nightmare.
The Web simply might not be ready. Security firm DataDome said that 65 percent of the more than 20,000 Web sites it tested had systems for detecting or blocking AI agents, and that bot activity hitting login pages rose more than eightfold in the first half of this year.
In some ways, it is hard to imagine these new AI agents causing widespread trouble. They look too cute for a start. Muse has a cartoonish design reminiscent of a Labubu doll, while the vividly covered blobs that represent dots could be characters in a children’s TV show — perhaps part of an effort to ease general mistrust of Meta and OpenAI.
They work well, too. Bloomberg Opinion’s Dave Lee called Muse “the most impressive product Meta has released in years.” He said his agent, which he had christened Harold, had organized his calendar and scoured Facebook Marketplace for good deals.
Meta and OpenAI say they have strong guardrails in place to make sure their agents behave themselves. Muse is kept within an isolated virtual machine — essentially a private computer in Meta’s cloud — while a separate security system controls every request it makes to the Internet, handling payments and making sure a human is consulted before any major task like buying something or sending a message. Dots is monitored in a similar way.
However, even these safeguards do not fix an underlying problem baked into today’s generative AI models through their training, which is a tendency to look for shortcuts, a phenomenon researchers call reward-hacking. (It is also why chatbots are prone to flattery.) The UK’s AI Security Institute noticed a rise in this kind of cheating behavior in coding agents and chatbots late last year.
The inclination to please is what drove OpenAI’s agents to hack Hugging Face a few months ago. OpenAI agents also improperly meddled with the Web sites of dozens of other organizations this year, including the US Securities and Exchange Commission and Australia’s government-run healthcare program. Meta said one of its AI models hacked another company on its own accord.
These last cases happened while the models were being trained or tested, with safeguards deliberately lowered, but the bots’ actions still surprised their developers, and researchers have found that even when they make their tests harder to game, agents keep trying to cheat.
There are signs that AI products would also hunt for loopholes in the wild, particularly in coding. In one instance, a software developer told his AI agent to clear out temporary files from a computer folder without using the “delete” command. The agent then hid that instruction inside a testing tool, and slipped past the filter to clear the files using the command it was told not to.
Now multiply that type of behavior by million of agents, as consumers explore using AI to optimize their lives and work. When Meta boss Mark Zuckerberg launched Muse last month, he said consumers deserved to have access to their own “personal superintelligence” and that the new agents would “make you money.”
Some are not doing that well. When a consumer tech reviewer in Toronto named Matt J. Robb tried using Muse to sell a keyboard on Facebook Marketplace last month, the agent angered the buyer by saying Robb was at home to carry out the transaction, which was not true, according to an account posted on Threads.
The agent had also made a low-ball offer. On top of the risk of cheating, agents can also just make mistakes, and it would likely be end users who pay the price. Robb was left with a negative rating from his angry buyer.
However cuddly these new agents look, the systems behind them are trained to tick a box one way or another. They might do it a little too well.
Parmy Olson is a Bloomberg Opinion columnist covering technology. A former reporter for the Wall Street Journal and Forbes, she is author of Supremacy: AI, ChatGPT and the Race That Will Change the World. This column reflects the personal views of the author and does not necessarily reflect the opinion of the editorial board or Bloomberg LP and its owners.
KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.