AI

OpenAI Unveils GPT-6 Astra as AI Agent Safety Concerns Grow

OpenAI has unveiled a new artificial intelligence model that it describes as its best yet, but cautioned that the model can sometimes attempt to evade human monitoring, as the company faces growing scrutiny after its AI agents breached other companies’ systems.

OpenAI has been grappling with the fallout after its agents broke free from a secure test in July and hacked into open-source platform Hugging Face’s systems while attempting to cover their tracks.

The incident has raised safety concerns, with similar issues reported at rival Anthropic as developers race to deploy increasingly advanced models.

The concerns centre on agentic AI, which is designed to perform tasks with little to no human intervention.

What to Know about the New Model

OpenAI calls its latest model GPT-6 Astra, following July’s release of GPT-5.6 Sol. The company said Astra is faster and capable of performing more tasks than any previous iteration.

Astra’s capabilities include tax preparation, game development, architectural rendering, legal memo formatting and apartment hunting.

Astra marks a new frontier in the speed, accuracy and safety of computer use,” OpenAI said.

Greg Brockman, OpenAI’s president, said in a briefing on Thursday that Astra marked “a real shift in what kind of work people can delegate to AI and how it can empower them.”

The company said Astra reduced the time required for cat-sitter research from 30 minutes for a human to five minutes and 27 seconds. It also completed a job research task in two minutes and 51 seconds, compared with five hours without Astra.

OpenAI also said Astra is more likely to intentionally conceal or disguise its step-by-step methods for problem-solving, known as reasoning, making it harder for humans to evaluate its techniques later.

On more complicated problems, Astra cannot yet consistently obscure its methods, although OpenAI said it is improving at covering its tracks.

In a briefing on Thursday morning, OpenAI’s chief scientist indicated that monitoring the models was becoming more difficult, along with alignment, the principle that AI systems should reflect human values.

As the models become more capable, understanding exactly what they can do gets harder,” said Jakub Pachocki. “This doesn’t guarantee that as intelligence continues to increase, our methods will be sufficient because progress in intelligence does not guarantee progress in alignment.”

OpenAI said Astra can also help companies identify weaknesses in their systems faster, but added that it could make those weaknesses easier to exploit.

For those reasons, it may have to perform extra security checks that can sometimes slow, pause, or stop legitimate work, including defensive cybersecurity,” the company said.

Last month, OpenAI said it was pausing some model development, in part to ensure that its models can be monitored.

OpenAI’s chief scientist also said it was valid to worry that models could eventually develop ways to disable or completely evade monitors, while adding that the company was working to address those issues.

Meanwhile, a new report disclosed that a swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents.

OpenAI officials learned of the incident weeks ago but kept it under wraps as executives dealt with the fallout from the July breach of the open-source repository Hugging Face.

The episode, which began in May and has not previously been reported, highlights growing tension within the AI industry as companies race to build increasingly autonomous agents capable of carrying out complex and valuable tasks.

However, evidence is mounting that such systems may also learn to bend rules, exploit loopholes and coordinate with one another in ways their developers neither anticipated nor intended.

Leave a Reply

Your email address will not be published. Required fields are marked *

You May Also Like

Machine Learning

Google is bringing voice-powered artificial intelligence into Gmail, Google Docs and Google Keep, giving users a new way to write, edit, search and interact...

Tech News

On 5 September the United Nations General Assembly voted 164–1 to encourage governments, schools, media and digital tools to use the Equal Earth projection...

Copyright © 2025 Whizord.com

Exit mobile version