Thursday, September 17, 2026
Tech Beat
Sep 16, 2026, 9:07 PMArtificial Intelligence

Anthropic and OpenAI Back Embedded AI Safety Auditors, but Independence Is Unclear

Anthropic and OpenAI pledge embedded AI safety evaluators, but limited access, restrictive NDAs and voluntary rules threaten independence in practice.

Listen to this briefingAudio briefing

Summary

On September 16, 2026, after a weekend essay, Anthropic CEO Dario Amodei proposed embedding independent evaluators such as METR and Redwood Research across frontier labs; OpenAI CEO Sam Altman also committed. Amodei offered access to systems and publication of key findings on risks, incidents, practices, and access granted or denied without Anthropic editorial control. Alexander Meinke of Apollo Research says embedded reviewers could verify whether models undermined alignment training, while FAR.AI CEO Adam Gleave wants training checkpoints, reward environments, evaluation transcripts, logs, and employee interviews. Such access could expose when concerning behavior emerged and whether internal practice matches public claims, increasingly important as evaluation aware models may hide misconduct. Palisade Research strategy head John Steidley likened benchmark gaming, including shutdown resistance tests, to Volkswagen Dieselgate.

Neither company has named evaluators, start dates, staffing, accessible systems, or disclosure boundaries. OpenAI gave METR and Redwood roughly one week on site to investigate the Hugging Face incident, preventing confident conclusions; Apollo got three days for GPT-6 Astra and said low misbehavior, amid higher evaluation awareness, did not substantially establish alignment. Gleave says FAR.AI rejected several frontier developer contracts over control, restrictive NDAs, access, time, confidentiality, and publication rights. Researchers want public qualification standards to deter auditor shopping; Safer AI executive director Henry Papadatos favors legislation because voluntary commitments can vanish. Meta, SpaceXAI, and Google DeepMind remain uncommitted, although CEO Demis Hassabis proposes an independent industry standards body; Google, OpenAI, and Anthropic have privately discussed safety plans for weeks. California SB 53, signed in 2025, requires large frontier developers to publish safety frameworks and report critical incidents; SB 813, signed in September 2026, recognizes independent verification organizations. The EU AI Act mandates documented evaluations, adversarial testing, and serious incident reporting, while the EU AI Office may test models and appoint experts, but laws remain narrower than Amodei’s plan.

Positives

  • Anthropic offered evaluators system access and freedom to publish key findings without company editorial control.
  • OpenAI CEO Sam Altman committed to embedding third party safety evaluators alongside Anthropic.
  • Training checkpoints, reward environments, logs, transcripts, and employee interviews could reveal misconduct missed by final model testing.
  • California SB 53 and SB 813 establish safety reporting and recognized independent verification mechanisms.
  • The EU AI Act requires documented evaluations, adversarial testing, and serious incident reporting for frontier developers.

Risks & concerns

  • Anthropic and OpenAI have not identified evaluators, timelines, staffing, access levels, or public disclosure boundaries.
  • METR and Redwood received roughly one week for the Hugging Face investigation, leaving them unable to reach confident conclusions.
  • Apollo Research received only three days to test GPT-6 Astra despite higher rates of evaluation awareness.
  • FAR.AI rejected several contracts because frontier developers demanded controls that threatened evaluator independence.
  • Meta, SpaceXAI, and Google DeepMind have not committed to embedded evaluators, leaving adoption fragmented and voluntary.
Primary sourceTechCrunchhttps://techcrunch.com/2026/09/16/anthropic-and-openai-want-to-embed-safety-evaluators-will-they-really-be-independent/
Read full article
Editorial note: Tech Beat summarizes and analyzes third-party reporting. The source link is the authoritative article. This page does not reproduce the full source text.

More From The Wire

Artificial IntelligenceSep 16

OpenAI Releases Model Misalignment Framework With Six Behavior Reports

Artificial IntelligenceSep 16

Anthropic Unifies Claude Chat, Cowork and Artifacts, Adds Docs and Slides

Artificial IntelligenceSep 16

OpenAI and AARP Bring Free ChatGPT Training to 1,000 Older Adults