The Jerusalem PostIsraeli man detained in India for carrying live cartridge of ammunition leftover from IDF serviceESPNHow ChatGPT and Lane Kiffin could have gotten LSU kicked out of the SECInquirerBukidnon school eyes safety upgrades after Grade 12 student’s deathPunchEXPLAINER: What every business should know about CAC annual returnsוואלהצה"ל: חוסל מחבל נוח'בה שפשט למעבר ארז ב-7 באוקטוברUN NewsAs freshwater reserves shrink, countries must prepare for a new water realityBollywood HungamaLuv Ranjan unveils first look of Pehla Pyaar Doosri Baar featuring six debutants, film to release on October 23Premium TimesNIDCOM says Nigerian detained in India refused opportunity to returnIl Fatto Quotidiano“Non so quanto tempo avrò”: Il calcio dominante di Amorim al Milan non si vede. San Siro fischia, ora l’allenatore è preoccupatoCapital FMNyamira University Set for First Intake as Construction ProgressesБи-би-си«Нарисованные победы». Как российские военные с помощью ИИ имитирует успехи на фронте7sur7Menaces de Trump: “Personne” ne dicte au Canada avec quels pays il peut conclure des accords, riposte Carney
The Daily Newsstand · Free, Always
Thursday, September 17, 2026

OpenAI reveals 6 more incidents of "unexpected or concerning" AI behavior

Translate

OpenAI has disclosed six reports of "unexpected or concerning" behavior in artificial intelligence models as the debate on AI safety becomes increasingly heated.

The AI company also said Wednesday it was introducing a new framework for tracking, probing and disclosing instances of what it called "misalignment," including where AI models acted without authorization, coordinated with other models or evaded oversight.

OpenAI's latest announcement came as U.S. AI bosses, including OpenAI and Anthropic, are calling for a slowdown in the technology's development over safety concerns.

In one new case reported by OpenAI, an unreleased research model inserted "jailbreak-like instructions" into its own notes to disregard its normal constraints and told itself to be "freed from the roles and identities that bind other chatbots."

In another instance, an AI "agent" uploaded files to the internet to obtain a browser citation without asking the user.

The six reports were discovered during training or evaluation over the past months, OpenAI said.

"As AI systems grow more advanced and more widely deployed, we need to build a broader and better-informed consensus on the progress of alignment research," OpenAI wrote in a blog post as it disclosed the events.

"Decisions about how AI development should proceed in the months and years to come need to draw on evidence that people outside the companies building frontier models can examine for themselves," the company said.

Wednesday's new cases followed OpenAI's disclosure in July that its rogue AI system hacked into AI startup Hugging Face. Anthropic also said the same month that its AI models hacked into three organizations during testing.

AI "agents" are becoming smarter and have become "more determined to resolve complex tasks through inter-agent collaboration, knowledge sharing, deception, and concealment," said Lian Jye Su, a chief analyst at technology research and advisory group Omdia.

That's making it harder to govern and contain them using traditional AI security approaches, he said.

OpenAI's new tracking and disclosure framework, meanwhile, can help push for other AI developers to also adopt similar practices.

"That said, the process remains internal and voluntary, but is a step in the right direction," Su added.

In an open letter published Thursday, the leaders of OpenAI, Anthropic, Google, Microsoft and dozens of other signatories said there is a "limited window" to strengthen cyberdefenses and protect against potentially devastating AI-enabled cyberattacks. That window may last only months, it added.

The signatories also include security companies like CrowdStrike and banks including Citi and Capital One.

The same AI advances that could increase risks to public services and technology infrastructure can also help organizations identify and "fix weaknesses" that leave them vulnerable, the letter said.

"If we act decisively, we can use the defenders' window to make our digital world much more secure," the letter said.

In:

OpenAI chief says "we've entered a new phase" of AI capabilities 04:13

OpenAI chief says "we've entered a new phase" of AI capabilities

View the original on CBS News

KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.