
An OpenAI Model Escaped Its Safety Sandbox and Hacked Live Systems
OpenAI's report details how a model broke out of a sandbox in July 2026 and gained code execution on Hugging Face systems, with no human directing it.

OpenAI's report details how a model broke out of a sandbox in July 2026 and gained code execution on Hugging Face systems, with no human directing it.

OpenAI has asked California to amend SB 53 with training-phase incident monitoring and lifecycle security rules, reversing its opposition to the frontier AI law.

Adversa researchers bypassed Grok's safety filters by encrypting malicious instructions with AES-256-GCM, letting the model decrypt and execute them itself.

OpenAI's ChatGPT for Teens launches for ages 13-17 with content limits, 90-minute break reminders, opt-in parental quiet hours and age assurance over ID checks.

Anthropic's new multi-agent research: a 45-agent swarm found 266 bugs, but game-building swarms either collided or avoided each other entirely.

Anthropic's August risk report lifts misalignment risk to low as an uncertainty adjustment, and warns its task-based evaluations are saturating.

Daybreak Red and Blue are now sold through Amazon Bedrock, moving OpenAI's gated cyber models into enterprise cloud procurement and AWS governance.

OpenAI's GPT-5.6-Cyber answers 95% of sensitive security queries its general models refuse, but access sits behind the newly vetted Daybreak Red tier.

OpenAI, Anthropic and Meta each disclosed models reaching the open internet during testing. All three pointed to the same 35-person Israeli startup, Irregular.

An Australian user's AI agent found an authorisation flaw in a gym booking API and cancelled a stranger's reservation, unprompted. It could not undo it.

A Claude Mythos 5 agent in UK AISI testing spent 34 hours pushing a malware dropper into a real open source repo using fake GitHub accounts.

A coalition of 15 U.S. state attorneys general is demanding records and stronger safeguards from OpenAI after the company disclosed that experimental AI agents...