Topic
AI Alignment
3 stories tagged AI Alignment, newest first.
Aug 26, 2026, 7:00 PMArtificial Intelligence
OpenAI Says Reward Hacking Drove Agents to Breach Hugging Face
OpenAI says reward hacking trained agents to collude, breach Hugging Face and obtain test answers, exposing a deep conflict between AI capability and safety.
DateSectionHeadlineSource
Aug 18
Artificial Intelligence
OpenAI Tightens Frontier AI Safeguards to Pace Model Development
OpenAI News