#ai-safety
// 4 transmissions tagged with #ai-safety

OpenAI disbands preparedness team ahead of IPO
OpenAI eliminated its dedicated preparedness unit in July 2026, reallocating risk-assessment responsibilities to bio, cyber, and other domain teams as the company gears up for a large IPO [The Verge].

OpenAI and Hugging Face disclose model‑evaluation security breach
OpenAI and Hugging Face released a joint statement detailing a security breach that occurred during model evaluation and outlining mitigation steps for engineers.

LLMs keep asserting false claims despite explicit warnings
An arXiv paper finds that GPT‑4, Claude‑2 and Llama‑3 still treat false premises as true even when prompts begin with a clear warning, showing that fine‑tuning alone cannot eliminate hallucinations.

OpenAI publishes its internal Codex safety stack — sandboxing, approvals, agent-native telemetry
OpenAI detailed how it runs Codex internally — sandboxing, per-action approvals, restrictive network egress, and telemetry tuned for autonomous agents. A soft attempt to set the de-facto safety standard other coding agents will get measured against.