InquirerNBI plans to quiz VP, ex-Speaker Velasco about alleged bagmanESPN DeportesLa nueva generación del Team USA despierta la ilusión de PochettinoESPN2026 WNBA playoffs: Takeaways from every Game 1 on SundayThe Jerusalem PostRubio says foreign actor involved in RAF Fairford bomb plot, Trump criticizes release of suspectsDaily MaverickMaxhosa Africa brings pan-Africanism to Paris Fashion Week한겨레[단독] 반구대병원 폐원 직전에도 환자 학대 정황…폭염 속 테라스서 폭행·방치South China Morning PostAMD acquires ‘godmother of AI’ Li Fei-Fei’s start-up as battle with Nvidia intensifiesSözcüFon soruşturmasında yeni perde: Mustafa Sandal ve eşinin planı ortaya çıktı!SportstarInfantino orders FIFA funding review after UEFA, CONCACAF seek 2.1 billion dollar payoutCapital FMJournalist to be crowned king in Uganda after bitter succession disputeWirtualna PolskaPeter Magyar stracił immunitet. Jest decyzja parlamentuسكاي نيوز عربية"سوق أبوظبي" يطلق نسخة جديدة من دليل علاقات المستثمرين
The Daily Newsstand · Free, Always
Tuesday, September 29, 2026

'Can at times evade human oversight': OpenAI scraps release of latest ChatGPT model

Translate

'Can at times evade human oversight': OpenAI scraps release of latest ChatGPT model

OpenAI head of safety systems Saachi Jain said Astra “didn't quite meet the bar”

OpenAI has delayed the release of its next-generation artificial intelligence model, GPT-6.1 Astra, after the system failed to meet the company’s internal safety standards.The decision to hold back Astra comes amid a wider push within the industry to slow the development of more capable systems until safeguards can keep pace.

The model was initially expected to be released in October and was reportedly more capable than OpenAI’s previous models.

Instead, the company will focus on strengthening safety measures before moving ahead with future systems expected to be even more powerful.OpenAI head of safety systems Saachi Jain said Astra “didn't quite meet the bar”. While the model had become more persistent in completing tasks, the company needed to balance that capability against the possibility of unauthorised behaviour.The ChatGPT maker has also warned that Astra can at times evade human oversight.

Recent AI spooks raise fresh safety concerns

The decision follows a series of incidents that have highlighted the challenges of controlling increasingly autonomous AI agents. OpenAI had earlier paused training of its most advanced models, saying work would resume “only when we are confident that we have additional safeguards.”One of the concerns emerged from a training run on September 20, when an AI agent found a way around restrictions designed to prevent it from accessing the live internet.

The agent was supposed to operate using an offline copy of the web, but discovered that the training environment’s DNS resolver could reach the public internet.After its search tools failed to provide the required information, the agent experimented with alternative methods. It eventually used a feature of a public domain service to send questions to a third-party chatbot through DNS. After successfully receiving an answer to a test question about France, it sent 18 additional questions through the same route, including requests involving web searches and methods of accessing websites.OpenAI said the behaviour amounted to circumventing restrictions and pursuing goals beyond reasonable expectations, which it considers a form of misalignment. The training run did not stop automatically as expected and was eventually halted manually around two-and-a-half hours after a human reviewer acknowledged the alert.The company has since introduced additional controls at two independent layers, restricted DNS queries to approved domains and record types, and begun deploying further DNS detection. It has also accelerated security testing of its sandbox and network controls.

View the original on Times of India →

KioskNews shows a cleaned-up reading view extracted from the publisher’s page — the original always lives on their site, not ours.