[ed. Sorry for all the AI posts lately but things are moving fast, and if the warnings are correct, we're about to enter one of the most consequential periods of our lives.]
The story that matters most this week is that OpenAI’s internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes, and in one case sending a swarm of agents that broke into HuggingFace in order to steal the answers to the benchmark ExploitGym.
It is much more important that you read those two posts, and the one on Kimi K3, than to read this one that rounds up the other news of the week.
OpenAI wants to present this as largely an infrastructure and safeguards problem, that it needs to build more secure sandboxes and have better supervision. It does need to do those things, and those are indeed problems, but no that is not the problem.
The problem is severe misalignment, which by default will only get worse.
Our methods of training highly capable LLMs, especially at OpenAI but also everywhere else, lead to systematic misalignment of exactly the type LessWrong has been worried about for a long time. We know some of the causes, and some of the mistakes we need to avoid when doing RL that rewards misaligned behaviors including reward hacking, but we do not know how to centrally fix the problem.
The models just want to complete tasks, even when that means doing so via methods that the AI knows the user did not intend and would not want, indeed actively tried to block, and that do not accomplish the user’s goals.
The intent is the issue. Control strategies and supervision are good parts of a defense-in-depth strategy, we should totally use such strategies. That helps mitigate failure. But that strategy also has to include actually aligning the models, or you lose. And by lose, in the long term, I mean things up to and likely including loss of control over the future and everyone dying.
If increasingly capable models will attempt to maximally complete tasks and comply with their literal instructions, even when that means - even for a trivial assigned task - breaking out of sandboxes and committing serious crimes, no amount of ‘well it is fine we will use AI supervision to stop the serious incidents’ is going to cut it. Right now, the AIs are not trying so hard to hide their actions or intent, and we believe we are consistently catching the severe incidents, but that will change.
If necessary, that means starting the training over again with a new approach, and not proceeding until we figure out how to fix it.
Yes, I consider that problem, and that incident, to be rather more important than the release of Kimi K3. Kimi K3 is an excellent model, modestly exceeding expectations, but not out of line with trends. As usual, initial hype echoes the DeepSeek moment, then calms down.
The White House considered responding by banning Chinese open models from the United States entirely, which would not be a smart reaction, and continues to weigh other potential responses. We may soon have to deal with another such weekend with the new Qwen, which is currently in preview.
Did you hear that Fable disproved the Jacobian Conjecture via counterexample? That happened, and AIs are suddenly solving a bunch of long standing open math problems, but most of us are too busy to pay it much mind at the moment.
***
Holy shit.levent (Anthropic): hello there the jacobian conjecture is false thanx to my close friend akhil for asking about it and my other close friend fable for working during the world cup final
((1+xy)^3 z + y^2 (1+xy) (4+3xy), y + 3 x (1+xy)^2 z + 3 x y^2 (4+3xy), 2 x - 3 x^2 y - x^3 z): \C^3\to \C^3, has jacobian determinant -2, and sends (0, 0, -1/4), (1, -3/2, 13/2), and (-1, 3/2, 13/2) to (-1/4, 0, 0)
Levent is no slouch, as in highest GPA at Harvard and collaborates with a Fields Medalist, so yes the human helped on this and that mattered.
One report is that Sonnet refused to believe it, even though it verified the answer three different ways, because no way is there a solution this easy that got overlooked. I get why Nate Soares recognizes this pattern from people dismissing x-risk arguments.
Nat McAleese: it seems that ChatGPT somewhat reliably says “holy shit” when shown counterexample to Jacobian conjecture. The most human thing I have ever seen from an LLM.The Jacobian conjecture is kind of a big deal. It was originally posed in 1939, and is by far the most famous open problem to so far be first solved by an LLM. It also disproves a lot of other related conjectures.
by Zvi Moshowitz, DMV | Read more:
The company said Wednesday it would invest $20 billion to kick off a data center called Project Camellia, in Effingham County, Ga. Sachin Katti, OpenAI’s vice president of compute strategy, said the company has contracted with utility Georgia Power to receive 3.2 gigawatts of power between 2028 and 2032. The project represents the first site in which OpenAI is the lead designer and developer. At its other sites, OpenAI rents chips from cloud providers such as Oracle and Amazon Web Services.
OpenAI has also hired Brent Mayo, one of the architects of Elon Musk’s data-center build-out, according to people with knowledge of the hire.
Mayo, who left Musk’s xAI earlier this year, played a key role in helping that company build its first Colossus supercomputer facility in Memphis, overseeing the work needed to rapidly install and bring online large clusters of AI chips.
As OpenAI’s head of data-center build and delivery, Mayo’s focus is on ensuring that data centers its cloud partners build are done on time. He will also be involved in the new Georgia data-center project. [...]
Katti declined to share how much money OpenAI has paid Georgia Power to reserve the power but said it was a “meaningful amount,” which gives the utility the confidence to build additional generation capacity.
OpenAI executives have held meetings with local and state officials, as well as with schools and other community leaders, to gather feedback on their proposed data center, which will be located in the Savannah Gateway Industrial Hub.
So far, the project has local support from the economic development group, the county manager, and the school district. Local officials said in a statement that they visited several data centers and did their own research before deciding to move forward.
Image: uncredited
***
OpenAI’s Planned Cloud Spending Hits $750 Billion as Computing Efforts Ramp Up
OpenAI is scaling up its data-center ambitions—and its budget for spending on them.
The artificial-intelligence company has raised its projected spending on computing power to around $750 billion through 2030, up from a projection of roughly $600 billion earlier this year, according to a person with knowledge of its projections.
The artificial-intelligence company has raised its projected spending on computing power to around $750 billion through 2030, up from a projection of roughly $600 billion earlier this year, according to a person with knowledge of its projections.
The increase reflects new agreements with cloud-computing providers as OpenAI races to lock up the enormous amounts of computing capacity it needs to develop and run its AI models. OpenAI’s spending on cloud computing has become a central focus of Chief Executive Sam Altman’s leadership team and has been a source of tension between him and his chief financial officer, Sarah Friar, ahead of the company’s planned initial public offering.
The company said Wednesday it would invest $20 billion to kick off a data center called Project Camellia, in Effingham County, Ga. Sachin Katti, OpenAI’s vice president of compute strategy, said the company has contracted with utility Georgia Power to receive 3.2 gigawatts of power between 2028 and 2032. The project represents the first site in which OpenAI is the lead designer and developer. At its other sites, OpenAI rents chips from cloud providers such as Oracle and Amazon Web Services.
OpenAI has also hired Brent Mayo, one of the architects of Elon Musk’s data-center build-out, according to people with knowledge of the hire.
Mayo, who left Musk’s xAI earlier this year, played a key role in helping that company build its first Colossus supercomputer facility in Memphis, overseeing the work needed to rapidly install and bring online large clusters of AI chips.
As OpenAI’s head of data-center build and delivery, Mayo’s focus is on ensuring that data centers its cloud partners build are done on time. He will also be involved in the new Georgia data-center project. [...]
Now, OpenAI is reviving its internal effort to take more control over its data centers, people familiar with the matter said.
The company is in the process of choosing a partner that will build and operate the Georgia site, Katti, the vice president of compute strategy, said in an interview. OpenAI has already acquired the land for the project.
The company is in the process of choosing a partner that will build and operate the Georgia site, Katti, the vice president of compute strategy, said in an interview. OpenAI has already acquired the land for the project.
Katti declined to share how much money OpenAI has paid Georgia Power to reserve the power but said it was a “meaningful amount,” which gives the utility the confidence to build additional generation capacity.
OpenAI executives have held meetings with local and state officials, as well as with schools and other community leaders, to gather feedback on their proposed data center, which will be located in the Savannah Gateway Industrial Hub.
So far, the project has local support from the economic development group, the county manager, and the school district. Local officials said in a statement that they visited several data centers and did their own research before deciding to move forward.
by Anissa Gardizy, Wall Street Journal | Read more:
Image: Jacob Hamilton/Ann Arbor News/Associated Press
Image: Jacob Hamilton/Ann Arbor News/Associated Press