Anthropic released Claude Opus 5. As usual I covered that in three parts: The system card, model welfare and capabilities.
OpenAI was revealed over the last two weeks to have left an internal model unsupervised for a week during a cybersecurity evaluation, with its cyber safeguards lowered, despite having had multiple previous incidents where models broke out of their sandboxes. During that test, the model broke out of the sandbox, then proceeded to use an agent swarm to hack into HuggingFace to get the test answers. The model was loose for a week before OpenAI realized what had happened.
This event was a really big deal. There are severe alignment problems at OpenAI, along with supervisory and infrastructure failures. The internal research model that did this, which my posts nicknamed Galaxy, has now been permanently deactivated.
There have been further developments, and I anticipate at least one additional post on the HuggingFace incident soon.
Partly as a response to this, over 1,290 employees at frontier labs signed an open letter, Pacing the Frontier. The letter warns that we are close to automating AI research, and that companies are racing ahead on this faster than we can handle it.
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.Both OpenAI and Anthropic put out statements of endorsement. Since that post, others have continued to sign, including OpenAI cofounder Ilya Sutskever and DeepMind cofounder Shane Legg. Dario Amodei has signed. Sam Altman has not signed, but is talking in Washington about the need to pace development.
All three of those developments are more important than anything in the weekly. There is plenty here, but catch up on those key events first if you have not done so.
This week was crazy. I am absolutely not moving to a 7-days-a-week posting schedule, and fully intend to take some weekdays off as soon as there is what passes for a lull. However, there is even more speed premium these days, so I will continue the policy of shifting posts to weekends when the speed premium is especially high.
by Zvi Moshowitz, DWAV | Read more:
Image: via
[ed. Things are moving fast, too fast. Zvi's newsletter has become the first thing I check every morning. People have long speculated that before AI becomes too dangerous (without our knowing it) we might see "warning shots" that give us time to prepare. It appears we've seen those now, so what are we going to do about it? (assuming people actually view recent incidents as warning shots. Or just don't care (Politico):]
As the top political operatives at Leading the Future, they oversee a network of pro-AI industry super PACs and nonprofits that friends and foes alike describe as an aggressive, well-funded machine attempting to obliterate their opponents much like the Star Wars superweapon.
Their goal: to defeat candidates who support the strictest AI regulations and champion those who want to unleash the development of the industry.
***
In the AI political universe, Zac Moffatt and Josh Vlasto are at the helm of the Death Star.As the top political operatives at Leading the Future, they oversee a network of pro-AI industry super PACs and nonprofits that friends and foes alike describe as an aggressive, well-funded machine attempting to obliterate their opponents much like the Star Wars superweapon.
Their goal: to defeat candidates who support the strictest AI regulations and champion those who want to unleash the development of the industry.