Anthropic commits to hosting embedded third-party safety evaluators with near-internal access, part of a three-pronged plan to slow AI capability growth.
September 12, 2026
brightray analysis
Summary
Dario Amodei published a blog post calling for deliberate deceleration of AI capability development via three mechanisms: embedded third-party evaluators (e.g., METR) with near-internal access, voluntary coordination on safety standards among democratic-country AI labs, and global coordination including China on narrow prohibitions like bioweapons. Anthropic unilaterally commits to the first step — granting evaluators company badges, desks, and access comparable to internal risk teams — citing the OpenAI-HuggingFace hack and AI's accelerating ability to recursively improve itself as the proximate triggers.
Why it matters
- Anthropic unilaterally commits to hosting embedded third-party evaluators (e.g., METR) with desks, badges, and access near-equivalent to internal risk teams — a concrete, verifiable accountability step.
- Amodei calls for democratic-country AI labs to coordinate safety standards and cap unchecked capability progress, with US government mediating to avoid antitrust exposure via a narrow waiver.
- China competition addressed by tightening chip/equipment export controls and cracking down on model distillation — framed as widening the US lead over 3–5 years rather than pausing US development.
- Global coordination with authoritarian governments proposed for narrow prohibitions (e.g., AI-assisted bioweapons), even while acknowledging stark limits on what's achievable.
- Post arrives as Anthropic researcher Jacob Coxon's public resignation over existential risk concerns intensifies pressure on frontier lab leadership to respond substantively.
Community notes—
No notes yet — be the first.
See every signal in the Feed