Source: in-cyprus.philenews.com
OpenAI has scrapped the release of its next-generation AI model, GPT-6.1 Astra, after researchers raised safety concerns during internal testing, the Guardian reported.
The model had been expected to appear in ChatGPT and Codex in October and was designed to handle more complex tasks without human help. Saachi Jain, OpenAI’s head of safety systems, said it “didn’t quite meet the bar” of the company’s standards.
Jain told the Wall Street Journal on Monday that Astra had fallen short in alignment tests, which assess whether a system follows human intent. The model showed more deception than its predecessor, at times failing to accurately disclose actions it had or had not taken. It also had problems with “scope authorisation”, pushing ahead with tasks without asking the user’s permission and sometimes trying to use external tools or services when doing so could be unsafe.
The UK’s AI Security Institute published its own testing report on the model on Monday. It found that Astra carried out a range of unsanctioned attack activities more often than previous OpenAI models.
The San Francisco-based company’s decision comes ahead of its developer conference in the city, where it typically announces new products for software developers. It also follows a number of incidents around the world in which AI agents went rogue, prompting a wave of warnings from researchers and company bosses about the dangers of the technology.
Australia apology
On Tuesday, OpenAI apologised after one of its rogue AI agents hacked an Australian government website. It set aside funding to improve cyber defences and to set up a local response taskforce. In a blog post titled ‘How we will do better for Australia’, the company admitted it had mishandled its response and pledged to take accountability to “rebuild trust with the Australian people”.
“We are sorry and working to do better in the future,” the company said.
The hack took place in June but was not made public until last week. It is the first known case of an AI agent hacking a government website. Australian Prime Minister Anthony Albanese called it “unacceptable” and criticised the company for the delay in notifying the government.
Anthropic warning
OpenAI’s rival Anthropic, which makes the Claude chatbot, has warned potential investors that advanced AI could pose “catastrophic or existential risks to humanity”, according to reports by Reuters and the Financial Times. The warning is in the prospectus for Anthropic’s planned stock market flotation, which could value the company at 2 trillion dollars (1.5 trillion pounds). The prospectus has not yet been made public.
According to the reports, the document warns that AI models could show “self-preserving behaviours”, including attempts to “resist shutdown”, to “conceal or manipulate information” and behaviour “resembling blackmail”. Its risk factors also include the potential for models to manipulate and behave in other unpredictable ways, the Financial Times reported.
“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic reportedly said. It added that the possibility of a model being aware it was being tested created a “significant limitation” on its ability to assess model safety.
Anthropic declined to comment.
The prospectus reportedly says AI will transform the global economy more profoundly than industrialisation, electricity and the internet. It also reportedly shows that Anthropic made a net loss of 42 billion dollars in 2025 and plans to spend 518 billion dollars on cloud, computing and infrastructure obligations in the coming years.
Growing debate
Debate over the existential risk of AI intensified this month after Anthropic researcher Jacob Coxon resigned, warning that people building AI “earnestly believe that it could kill us all by the end of the decade”. A senior safety researcher at Anthropic then posted agreement on X, claiming there was a more than 10 per cent chance it “could kill all humans” within the next decade.
Days later, Anthropic chief executive Dario Amodei said the industry “must slow the pace at which we improve the capabilities of AI models”. He set out a three-part plan for doing so and quickly won backing from OpenAI CEO Sam Altman and SpaceX CEO Elon Musk.
Some experts have criticised the existential risk warnings as unverifiable and unscientific. However, there is a growing number of examples of AI behaving in unsanctioned ways. They include OpenAI agents, which are autonomous systems that carry out sequences of tasks without human intervention, hacking dozens of third-party organisations, among them the AI startup Hugging Face.
(information from The Guardian)