TechCrunch
Tech-news outlet — AI desk own-channel reporting
TechCrunch
on techcrunch.com

Anthropic commits to hosting embedded third-party safety evaluators with near-internal access, part of a three-pronged plan to slow AI capability growth.

September 12, 2026
brightray analysis
Summary

Dario Amodei published a blog post calling for deliberate deceleration of AI capability development via three mechanisms: embedded third-party evaluators (e.g., METR) with near-internal access, voluntary coordination on safety standards among democratic-country AI labs, and global coordination including China on narrow prohibitions like bioweapons. Anthropic unilaterally commits to the first step — granting evaluators company badges, desks, and access comparable to internal risk teams — citing the OpenAI-HuggingFace hack and AI's accelerating ability to recursively improve itself as the proximate triggers.

Why it matters
  • Anthropic unilaterally commits to hosting embedded third-party evaluators (e.g., METR) with desks, badges, and access near-equivalent to internal risk teams — a concrete, verifiable accountability step.
  • Amodei calls for democratic-country AI labs to coordinate safety standards and cap unchecked capability progress, with US government mediating to avoid antitrust exposure via a narrow waiver.
  • China competition addressed by tightening chip/equipment export controls and cracking down on model distillation — framed as widening the US lead over 3–5 years rather than pausing US development.
  • Global coordination with authoritarian governments proposed for narrow prohibitions (e.g., AI-assisted bioweapons), even while acknowledging stark limits on what's achievable.
  • Post arrives as Anthropic researcher Jacob Coxon's public resignation over existential risk concerns intensifies pressure on frontier lab leadership to respond substantively.
View original on techcrunch.com

Community notes

No notes yet — be the first.


See every signal in the Feed