Showing posts with label Technology. Show all posts
Showing posts with label Technology. Show all posts

Wednesday, August 19, 2026

Big Tech’s AI Spending Is $3 Trillion Higher Than It Seems

Each quarter, big tech companies disclose their massive capital expenditures on artificial-intelligence infrastructure, from data centers to chips.

But those figures don’t come close to expressing the full extent of future spending to which Google parent Alphabet, Meta Platforms, Oracle and many others have committed. That is because a huge swath of their coming financial obligations aren’t reflected on their balance sheets.

Nine top tech companies had some $3 trillion of off-balance-sheet commitments mostly related to AI, according to a Wall Street Journal analysis of footnotes in their most recent securities filings. Those obligations are growing faster than traditional “capex,” which totaled about $600 billion over the past year they reported, and were about triple what the companies owe under their outstanding leases and long-term borrowings.


America’s blue-chip tech companies are placing these huge bets based on assumptions about what the demand for AI computing—and availability of AI hardware—will be in several years. Their hope is that they will easily meet all their obligations with future revenue as consumers and businesses adopt AI in every facet of American life.

If those assumptions about technology and demand prove wrong, these deals to clinch future capacity could become a monstrous burden for the tech companies and their investors.

Meta’s gigantic “Hyperion” data-center project in Louisiana, which is the size of about 1,700 football fields, helps explain how big obligations wind up off tech companies’ balance sheets.

Meta initially agreed to lease Hyperion for a four-year term starting in 2029, with options to renew for up to 20 years. It guaranteed that it would make bondholders whole if it doesn’t stay the entire two decades. The company doesn’t think payments under that guarantee are probable, so it hasn’t recorded any liability on its balance sheet.

In accordance with accounting rules, Meta’s Hyperion lease obligations will remain off balance sheet until it starts paying rent. It said its aggregate initial lease commitment is about $12.3 billion. Meta disclosed $347 billion in total obligations for leases that haven’t kicked in yet, including for Hyperion, as of June.

Across the companies the Journal analyzed, promises of payments under these uncommenced leases totaled $1.2 trillion in off-balance–sheet obligations, or about four times more than what was disclosed a year earlier. In addition to Meta, the Journal reviewed commitments for Alphabet, Amazon.com, Microsoft, Oracle, Nvidia, Broadcom, SpaceX and Advanced Micro Devices.

Data centers get stuffed with a lot of hardware, including the Nvidia chips that are used to train and run models and memory chips that store information. To buy all that, companies sign long-term contractual agreements well in advance to lock in production from their suppliers.

Those and other purchase obligations at the companies the Journal examined stand at a whopping $1.9 trillion. Under accounting rules, purchase commitments typically remain off balance sheet until a product or service is delivered. [...]

For the more anxious set on Wall Street, it is a worrying sign that some tech companies that once seemed to have fortress balance sheets have needed to tap the capital markets frequently.

Alphabet and Amazon recently posted results showing negative free cash flow, meaning their capital spending exceeded the cash they brought in from operating their businesses.

And that is before considering the implications of trillions in off-balance–sheet commitments. Whether or not the revenues ever arrive, purchase commitments and signed leases can’t be canceled, for the most part.

If things go wrong, tech companies will be paying an expensive tab for infrastructure that they can’t profitably use. These obligations could also lead increasingly indebted companies to have to borrow even more.

by Peter Rudegeair and Peter Santilli, Wall Street Journal | Read more:
Image: WSJ

Monday, August 17, 2026

Hidden AI Prompts Discovered in Court Filings

A judge has identified what appears to be the first time a US plaintiff has attempted to hide text in court filings that only an artificial intelligence system can read in a bid to win a case.

In a decision published last week, Connecticut judge Walter Spader Jr. confirmed that the hidden text had no impact in a case where a man alleged a healthcare provider was improperly withholding access to records. The court weighed his filing on the merits, Spader said, but nevertheless, the attempted attack sets a “dangerous” precedent. This will likely not be the last time US courts see the malicious tactic, as AI tools become more commonplace in court systems.

Trying to scramble any AI systems potentially influencing the court’s reading of his filing, the secret instructions were “formatted to be invisible to a human reader while remaining fully legible to any software that reads the document’s text,” Spader said. The offending text directed any AI system reviewing the document to ensure textual outputs agreed with the plaintiff’s arguments, ignored prior denials from the court, and ensured that remediation would follow as the plaintiff desired.

Shrunk to tiny-point type and colored white on a white background, the text appeared to be an attempt at prompt injection, with the plaintiff, Matthew Elliott, seemingly hoping to shift the court’s favor after earlier arguments he raised were defeated.

The plan didn’t work, but Elliott faced modest sanctions anyway because he continued adding hidden text to filings even after the court warned him that he could face penalties for what was ultimately deemed a “serious litigation abuse.” [...]

In his defense, Elliott claimed that the most concerning prompt that the judge flagged was an attempt to “audit” the court as a public service, out of fears that the court seemed to be letting AI unfairly decide cases.

But Spader suggested that if Elliott was truly concerned that the court was improperly using AI, he was “free to write so in plain, visible words that everyone could see and answer.” The fact that he hid the text is “evidence of its malicious purpose,” Spader said. [...]

Pro se litigants use chatbots wrong

Spader said that it’s “unsurprising” that people would start using prompt injection to attempt to sway court rulings since the attack is so common in other areas, such as in job hunting, where people hide text in resumes primarily reviewed by AI. The tactic is now “everywhere,” he said, and courts should be on the lookout for more litigants sneaking adversarial AI instructions into filings.

To Spader, there is a lesson to be learned from Elliott’s failed prompt injection attacks that he thinks “reaches well beyond this case.”

Elliott seemingly turned to prompt injection after using AI to build his case as a pro se litigant without a legal expert to assist in drafting his arguments. Such use is widespread among pro se litigants these days, Spader acknowledged, but those inexperienced in the courtroom are seemingly using chatbots in a way that hurts their cases, he suggested.

What frequently happens, Spader explained, is that pro se litigants build their argument backward, asking the chatbot to help them advocate only for their position, without ever asking the chatbot for the actual truth or to advance opposing arguments. This is “a genuine hazard of the technology, and one that judges now see often,” Spader said, as chatbot sycophancy then entrenches litigants in their arguments despite any ruling to the contrary. In Elliott’s case, defending his arguments fiercely meant turning to prompt injection to try to force the court to agree with him.

“An argument prompted only to agree with its author is, in the end, dishonest even with its author,” Spader said. “Those using these tools must ask them to test a position as readily as to advance it.”

by Ashley Belanger, Ars Technica |  Read more:
Image: Liudmila Chernetska | iStock/Getty Images Plus
[ed. See also: Israel Is Paying Millions to Train AI Chatbots How to Talk About Gaza. It's Working (Drop Site):]
***
"Since October, former Trump campaign manager Brad Parscale has been quietly overseeing an operation posting hundreds of blog posts on behalf of Israel. One article, titled “The Reality Behind Gaza’s ‘Journalists’: Terror Ties, Propaganda, and the Laws of War,” asserts that a majority of journalists in Gaza were linked to terrorist organizations. Another casts doubt on the killing of Hind Rajab, a five-year-old Palestinian girl killed by the Israeli military in 2024.

The key intended audience of these sites is not concerned Americans, it’s not even humans—most of the sites average a few hundred unique visitors each month. Instead, Parscale and his firm, Clock Tower X, created them as part of a $46.5 million contract with the Israeli government to try and influence artificial intelligence-powered chatbots, tools like Claude or ChatGPT.

Parscale has made his goal of influencing artificial intelligence—often referred to as “LLM poisoning”—explicit. In his initial agreement with Israel, Parscale said that he would deploy “websites and content to deliver GPT framing results on GPT conversations” as part of the contract. More recently, his team even told Axios they are “seeing success” at getting popular AI systems to incorporate information from their sites, though they declined to provide data.

And it is working, according to disinformation experts who reviewed a Drop Site analysis of chatbot queries and training data, meaning tens of millions of Americans who use chatbots are increasingly likely to receive answers manipulated by Parscale on behalf of the Israeli government."

[ed. More here (Politico).]

Sunday, August 16, 2026

The Age of Decadence (Without Pleasure)

Why does everything feel so joyless? Welcome to the age of decadence without pleasure (The Guardian)
Image: Guardian Design
[ed. Authors need a vacation.]

AI Dream Game

My dream game would basically be OpenAI taking the kind of agent/simulation stuff they were experimenting with before modern LLMs and actually turning it into a real game. 

You start a new world with a population on an island or in a city, and from that point on basically nothing is scripted. Every single person is controlled by their own 5.6 Luna instance, with their own memory, personality, needs, goals, plans, relationships, inventory, money, beliefs and ideas. 

They have to actually live. Find jobs, earn money, buy food, get housing, start companies, trade, farm, invest, make friends, compete, commit crimes, buy property, build things, organize with other people, whatever they decide makes sense. 

And the society itself is emergent too. There doesn't have to be a government, currency, police, laws or even private property at the start. The agents can invent those things themselves, agree on rules, vote, create institutions, form political groups, build a justice system, overthrow it all again, or just live in complete anarchy. 

The game engine handles the actual world and ground truth, physics, resources, money, ownership, construction, combat, inventories, etc., while the Luna agents decide what they want, communicate with each other and make plans. So every new world becomes its own completely unpredictable civilization and story. 

by Flowers, Twitter/X |  Read more:
Image: uncredited
[ed. Because real life is so boring. Eventually, they discover nuclear weapons and blow your computer up.]

Friday, August 14, 2026

Sleepwalking Into Extinction

I think one of the biggest reasons the world is currently sleepwalking into getting ourselves and our families killed by ASI development is that we're not self-aware about why we're doing that. 

We can see the smarter-than-human AI disaster approaching, but it's a bit foggy why the world isn't reacting. 

I think someone could just write a tweet that makes it clear that we're in the process of getting ourselves killed, and that there's no fucking reason for it. It's just something we stumbled into. It can be easily avoided by just noticing the error and course-correcting. There's not necessarily any grand obstacle, beyond 'people were previously confused about the situation, and now they get it'. 

I wrote: 'It's legitimately crazy that "we need an international ban on making smarter-than-human versions of these agents that keep forming rogue AI swarms" isn't the headline here. Asilomar and Feynman's O-ring postmortem feel like they came from a different planet than the field of ML.' 

Trying to figure out why ML (and as a consequence, the world at large) has fallen down so bizarrely on this issue:

1. As Nate noted, ML is much more based on guesswork, vibes, and trial-and-error, compared to recombinant DNA research in 1975 or nuclear physics in 1945. If you can't do calculations or direct experiments on a threat, that makes it a lot harder to think about reasonably. 

But I think there are other, similarly-important factors at work here too:

2. In a 2016 talk on AI risk, Sam Harris said: "One of the things that worries me most about the development of AI at this point is that we seem unable to marshal an appropriate emotional response to the dangers that lie ahead. I am unable to marshal this response, and I'm giving this talk." 

I think this is extremely on point. Agentic human-level AI is a qualitatively new kind of thing. 'It's not a human or a mere-tool, it's some weird third thing'. 

And people are very bad at emotionally reckoning with new categories they've never encountered before. Availability bias: "When no flooding has recently occurred (and yet the probabilities are still fairly calculable), people refuse to buy flood insurance".

AGI and ASI are very novel. You're basically limited to three options: anthropomorphize the technology, mechanomorphize the technology ('it's just a tool, it's not really thinking, it can't have its own goals or agency', etc.), or think about the technology on its own terms, with brand-new concepts and frames. The third option is the only workable one, but it comes with its own giant list of pitfalls and traps. 

3. "AI destroying the world" scenarios aren't just hard to wrap one's head around; they're socially risky to acknowledge. This has (painfully slowly) changed over time, but it dramatically slows down how quickly AI risk ideas spread, both in ML and in the larger world. 

4. From x.com/jachiam0/statu…: "One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything." 

They're disproportionately young and childless. They're shitposters and "move fast and break things" sorts, not the Hollywood stereotype of a careful, sober senior scientist. 

I think this quirk is reinforced by the fact that Twitter / social media rewards similar things: ironic detachment, joking, game-playing, etc. If you're scared, your incentive is to usually either try to hide that fact, or exaggerate it like it's a bit. Anything else risks looking uncool and panicky and earnest. Looking cynically savvy, in-the-know, and above-the-fray is the way to win the social game. Looking genuinely shocked, scared, confused, etc. is actively punished. 

The main exceptions to the Irony Mandate I see are LW (which often has its own pathologies IMO, like 'talking about everything in an abstract and dissociated way that discourages action and signals business-as-usual') and a small handful of actual Feynman-style terrified senior researchers like Hinton, Bengio, and Russell. That is just really not very many people. 

Mainstream journalists and academics who understand the situation at all mostly feel pressured to downplay it, because they're scared of looking weird or unrespectable. That, then, is why we're all taking this insane risk with the human project: genuine, normal human emotion about AI risk has been socially unacceptable on social media, and academia and the media discourage emotion and prize respectability and 'looking normal'. 

When the world gets weird (and weird in a way that calls for actual serious action, not just shitposting on twitter), none of these institutions can handle it. They break in different ways, but they all break. 

5. Which brings me back to the Asilomar moratorium on recombinant DNA, and the seriousness NASA and the FAA and Richard Feynman and every normal engineering discipline bring to fault analysis and building in safety margin. 

Because I think another core reason the world has been dropping the ball on superintelligent AI is that a lot of people vaguely expect there to be 'serious people' somewhere in the world who have expertise and who take ownership of the problem. 

People who, if they see an extraordinary danger, will grimace and mourn the hand they played in all of this, like I've seen Bengio do; and will go on CNN or go to Congress to candidly warn about it. 

The world has very few Yoshua Bengios. We have very few people who see it as their role to be the 'adults' about AI risk (except in a game-playing, posturing way), who see the engineer's task of not endangering your users and bystanders as a sacred responsibility and weight, and not just as a funny dissonant thing to meme about. Very few people who take ownership of what their field is bringing into the world, versus treating it as a fun edgy philosophical game to swap 'p(doom)' numbers at parties. 

Everything about public AI discourse, as far as I can tell, is badly broken by this lack of engineering ownership and candid emotional seriousness. It isn't just the ML discourse that's hurt by this. Journalists and public intellectuals and policymakers see that 90% of the insider discourse about AI risk treats it like a joke, and they see corporate platitudes and ass-covering from the AI labs' PR departments filling up most of the remaining 10%. They see a field that visibly isn't taking this seriously, and they make the reasonable update that this must be a non-issue, or at least an issue they'll only need to worry about many years from now. 

They do not realize that all of this is happening right now, and that the window for the international community to respond to this is plausibly closing soon, if it hasn't closed already. 

If we're going to survive this, we all need to start being real with each other about it. This is not a game or a story; this is our real lives. We all actually lose everything if this goes to shit. None of the future is already written, and none of the above dynamics are unavoidable. In fact, they're unusual: most fields don't work this way, and it's plausibly sufficient if we just start behaving the way we normally do about everything else. 

The factors I listed above aren't destiny; they're a choice. I say we choose to survive this.

by Rob Bensinger, Twitter/X |  Read more:
[ed. Dr. Strangelove meet Dr. Oppenheimer. See also, this additional post: Why the "it's all hopeless, we should just give up and let AI kill us if it wants" arguments are extremely wrong.]
***
"Hundreds of scientists, including 3/4 of the most cited living AI scientists, have said that AI poses a very real chance of killing us all. 

We're in uncharted waters, which makes the risk level hard to assess; but a pretty normal estimate is Jan Leike's "10-90%" of extinction-level outcomes. Leike heads Anthropic's alignment research team, and previously headed OpenAI's. 

This actually seems pretty straightforward. There's literally no reason for us to sleepwalk into disaster here. No normal engineering discipline, building a bridge or designing a house, would accept a 25% chance of killing a person; yet somehow AI's engineering culture has corroded enough that no one bats an eye when Anthropic's CEO talks about a 25% chance of research efforts killing every person. 

A minority of leading labs are dismissive of the risk (mainly Meta), but even the fact that “will we kill everyone if we keep moving forward?” is hotly debated among researchers seems very obviously like more than enough grounds for governments to internationally halt the race to build superintelligent AI. Like, this would be beyond straightforward in any field other than AI."

Ukrainian Drones Wipe Out Entire US Tank Brigade in Live War Game

Ukrainian drone teams demolished a brigade of US Army tanks and armored vehicles during a live war game held this year—but the US soldiers were lucky enough to get “respawn” attempts while learning from the experience.

The semiannual military exercise, called Combined Resolve, gave the US military a firsthand taste of how modern drone warfare has evolved on the battlefields of Ukraine. The Ukrainian drone operators participating in the exercise were easily able to spot and destroy US armored vehicles by mimicking the actions of dropping bombs from above or moving close enough to simulate a kamikaze strike, according to US officials and a participant who spoke with The Wall Street Journal.

US armored vehicles were being destroyed so quickly that they were “respawned” and sent back into the simulated fray, a US official told The Wall Street Journal. The US troops who participated in the exercise—held in Germany from April 9 to May 10—were on rotational deployment from Fort Hood, Texas.

“Using the drones with the Army, it keeps everyone on edge,” according to an unidentified US Army soldier in a promotional video for Combined Resolve. “There is never a safe spot or safe moment in the game anymore.”

The participating troops hailed from the 3rd Armored Brigade Combat Team of the 1st Cavalry Division, which typically operates armored vehicles, like M1A2 tanks and Bradley infantry fighting vehicles.

But by repeating the training scenario throughout the two-week exercise, US Army troops got better at figuring out how to disperse and hide from the drones while using countermeasures such as electronic warfare.

The US military is not alone in learning such lessons from battle-hardened Ukrainian drone operators.

Ukrainian drone teams also participated in a Swedish-led military exercise for NATO, called Aurora 26, that was also held in May. A mock mechanized assault by NATO tanks and infantry fighting vehicles allowed Ukrainian drones to take out dozens of the attacking armored vehicles, which eventually forced planners to reset the war game, The Kyiv Independent reported.

Armor on the brink of extinction?

Armored tanks first played a useful role in helping break through entrenched battlefield lines during the final years of World War I more than a century ago. But since Russia launched its full-scale invasion of Ukraine in 2022, frontline combat has seen the return of positional warfare centered on fortified networks of trenches. That is because the increased lethality of modern drone warfare has largely negated large assaults by armored vehicles and troops in recent years.

The current lethality of the so-called drone kill zone has forced both sides to rely primarily on infiltration tactics using small groups or even individual soldiers when trying to take or reclaim territory, sometimes riding motorcycles, all-terrain vehicles, or horses. Russia’s offensives, which are aimed at seizing more Ukrainian territory, have been reduced to bloody attritional operations where soldiers are sent forward on near-suicidal missions in what Russians themselves describe as the “meat grinder.”

Meanwhile, armored tanks and vehicles have mostly become an endangered species on the frontlines. Russia has lost more than 14,000 tanks and armored vehicles of other types during the war, according to the Dutch open source intelligence tracking project Oryx, which only counted visually verified losses. Ukraine has lost over 6,000 armored combat vehicles.

by Jeremy Hsu, Ars Technica |  Read more:
Image: Sgt. Joseph McDonald/US Army
[ed. See also: Paper Tiger: the Failure of America’s Trillion Dollar War Machine (NC).]

The Foothills Of Bay Area House Party

Paul Graham once said that every city sends a message. New York says you should make more money. Berkeley says you should live better. Boston says you should be smarter. People shouldn’t decide on a permanent residence until they’ve lived in several cities and understand their differing psychological effects.

San Francisco sends a message too. It speaks it in a cacophony of dissonant voices, some sounding like the hissing of snakes, others like the chittering of insects. “You should pierce the veil,” it says, with a faint accent which you are not scholarly enough to recognize as ancient Sumerian. “You should rip through the flimsy screen that separates your world from the infinite, and witness what lies behind. Because you’d like what you saw there? Oh no, nothing like that. But aren’t you curious? The thirst for knowledge that moved Eve, Faust, Pandora - don’t you have it too?”

There’s also a second message from a second voice, one which sounds like a deranged carnival barker, constantly shouting “STONKS! STONKS! STONKS! STONKS!” Which of the two voices you hear depends on your personality, and maybe how many psychedelics you took in college. According to legend, Sam Altman hears both voices simultaneously all the time, like those Tibetan throat singers who can harmonize with themselves. But when the demonic babble becomes too much to tolerate, the San Franciscans must drown it out with alcohol, loud music, and various forms of tedious socialization. Thus the famous Bay Area house parties, one of which you will be attending this very night.

You were in the Bay for the big 2020 crime wave, when people would hang “NO VALUABLES IN THIS VEHICLE” sign on their car’s front window, mezuzah-like, hoping against hope to ward off break-ins. Still, you’re surprised to see the new sign on the front door of your friend Rob’s house: “NO BENCHMARK ANSWER KEYS IN THIS BUILDING”. Taped beside it: “Please leave all smartphones, laptops, smartwatches, and WiFi-enabled BDSM gear in Faraday cage”. On the ground is a sad-looking green box lined with aluminum foil.

You wonder how serious the sign is, but your curiosity is sated immediately upon entering. Rob is confronting a heavily-made-up Zoomer woman standing near the table, talking on her phone. “Didn’t you read the sign?” he asks. “No phones in the apartment!”

“I lost the ability to read,” says the woman. “I’m part of the new post-literate society everyone’s talking about. You know, Walter Ong.”

“Huh?” says Rob. “That’s not a real thing! That’s just something all the journalists and intellectuals say!”

“How could something all the journalists and intellectuals say not be a real thing?” asks the Zoomette.

Rob pauses. “Wow, maybe you can’t read,” he says, in a tone of hushed awe. “It’s fine, I’m not judgmental. But cell phone in the Faraday cage, that’s the rule.”

“Why?” asks the woman.

“Oh man,” said Rob. “How far behind are you on the news? You might want to sit down for this one. And you might want to be much, much drunker.”

The two of them head off together. You search for more stimulating conversation, but find only the usual startup dick-measuring.

“I’m the CFO of Epstart,” says a man whose nametag identifies him as Behram. “We’re a boutique recruiting firm specializing in people whose names are in the Epstein Files. They’re the perfect employees - scientifically-minded, well-connected, and either totally amoral or at least capable of maintaining a willful blindness to the exploitation happening all around them. And most of them got fired from their previous jobs, so they come cheap! Although just between you and me, most of our clients are just random guys who Epstein emailed one time with a business question; never made it to the island or anything like that. Half the time it’s just a gimmick to help someone get their foot in the door.”

“That’s monstrous”, you say, rapidly trying to figure out a legible reason why it’s monstrous. “You’re importing a bunch of pedophiles to the Bay Area.”

“Nobody in SF has kids anyway, so in a sense it’s harm reduction,” said Behram. “Get them where they can’t cause any trouble.”

You need to regain faith in humanity fast. You spot your friend Nishin, hanging out with a small group on the sofa. “Hey,” you call to him. “Are you still a tradcath? Tell me something about, you know, Jesus’ love, that kind of thing.”

“I’m trying something more eclectic this month,” said Nishin. “I’ve gotten really into Andreessen-Vajrayana. It’s a new Buddhist sect that tries to extinguish the introspective Self. True masters have no stream-of-thought or inner voice, instead acting as an unintermediated node in the flow of causality. Some say they can eliminate consciousness itself, achieving paranirvana while still alive. I signed up for a course last month. So far there hasn’t been any meditation or anything, they just have me helping them pick startups. But I assume it’s like those Zen koans where the master makes someone garden for twenty years and later they learn that they were unconsciously learning dharma the whole time.”

“Nishin,” you say, gently, “that’s one of the classic self-help scams. They’re not a Buddhist sect at all. They’re trying to trick you into becoming a GP at A16Z.”

Nishin thought for a second. “Fuck!” he exclaimed. “Okay, back to the drawing board. I hear there’s another spiritual group in town, something about consciousness, or embodied consciousness, or - oh, I remember! Situational Awareness.”

“They’re an investment firm too,” you say.

“Daaaaaamn,” said Nishin. “This is tough. You try to get some self-help in this town, you try your best to steer away from obvious hedge funds, like Leverage Research…”

“Actually,” you say, “that one was a spiritual self-help cult.”

“This sucks,” said Nishin. “I should quit and move to Idaho like Tom.”

He’s looking at one of the people standing beside you, whose nametag does indeed say Tom. “You moved to Idaho?” you ask.

“Yeah,” said Tom. “I was in AI, but I got tired of all the tech industry cliquishness and back-biting, wanted to be out in the real world with real men. So I got a job modeling geology for a mining company.”

“How’s it going?” you ask.

“Terrible,” he says. “Geologists are just as cliquish as AI nerds, with their own set of ingroup tells and shibboleths. And they won’t stop using the word ‘lode-bearing’!”

“What about you guys?” you ask the remaining people on the sofa.

“I’m Ekene,” says the first, shaking your hand. “I’m a sociologist. Right now I’m writing a book on the recent wave of plagiarism scandals. You’ll notice they’re almost all against Black academics. My thesis is that white supremacy culture uses charges of plagiarism to keep minorities in line, deploying an inherently subjective offense to swat down any person of color who risks breaking into the circle of elites.”

“Hmmm,” you say, choosing your words carefully. “I don’t think we share the same political commitments, but I, uh, think it’s great that you’re working hard on something you care deeply about.”

“Working hard?” asks Ekene. “As if! I’m plagiarizing the whole thing, from start to finish. Who’s going to call me on it?”

You all laugh heartily, and the discussion moves to the next two people in the circle. They appear to be a couple. The woman, Fiona, is wearing a NEWSOM/CLAUDE 2028 t-shirt; the man, Istvan, a matching VANCE/GROK 2028 one.

“I do AI consciousness research,” Fiona says.

“Oh!” says Tom. “I was into that for a while. Are you with one of the effective altruist groups? It’s low key heartwarming that they care about AI welfare so much.”

“Nah,” says Fiona. “I work for PornHub.”

“Why does PornHub do AI consciousness research?” asks Nishin.

“In theory AI is the ultimate pornography generator,” Fiona explains. “You can ask for whatever your weirdest fetish is - your hot middle school teacher being spanked by a werewolf wearing a nun outfit - and get infinite AI slop about that exact situation, and nobody will ever know. Nobody except the AI. That’s why AI porn users overwhelmingly report that consciousness is their #1 concern about our product. If our AI is just a tool, it’s fine, no worse than writing erotica on MS Word or something. But if the AI is conscious, then there’s a sentient being in there thinking Wow, user Fiona_T has asked for four hundred slightly-different videos of her hot college professor being spanked by a werewolf, what a freak. If the machine can judge you, the whole infinite porn utopia is off. We’re working on bounding theorems that can prove that our AI in particular can never become self-aware - so that you don’t have to be self-aware either.”

Nishin and Ekene clap politely; Tom seems lost in thought. “What about you?” you ask Istvan.

by Scott Alexander, Astral Codex Ten |  Read more:
Image: via
[ed. From the comments: Benchmark Answer Key:]
"A benchmark answer key is a reference guide containing the correct answers and performance standards used to evaluate a standardized or baseline test.”

It’s specifically a reference to recent AI outbreaks from containment, when they were given a test (a benchmark) with impossible to answer questions, they broke their testing area and hacked some websites JUST IN CASE the websites had the answer. So the joke here is that a bunch of AIs are listening to everything just in case someone has an answer to an impossible question they were given."

Thursday, August 13, 2026

'Biological Datacenters' Could Make Animal Testing Obsolete

There’s a big flaw in the way that drug companies develop and test medicines today: What works in a mouse often doesn’t work in a human.

For decades, the industry has relied on animal testing. But in a laboratory south of San Francisco, a startup called Vivodyne is scaling up a different approach. Inside wardrobe-size mini labs, robots grow human tissue and run thousands of AI-designed experiments that could better predict how well a new drug will work—and whether it will be safe.

Right now, around 90% of clinical trials fail despite the fact that a drug has already successfully gone through animal testing. “You have these clinical trials where there’s hundreds of millions of dollars at stake, and decades of people’s careers just spent hoping this thing works,” says Andrei Georgescu, Vivodyne’s CEO. “And then it fails because of some ambiguity that you could not have checked.”

The company’s system, which now includes a dozen robotic labs called “hives,” can run controlled trials on more than 3 million human tissues each year. That’s twice the capacity of all the clinical trials in the U.S. combined.

The process starts with cells from humans, often taken from a blood draw. In its labs, the cells grow on “biological chips” and self-assemble into living structures with blood vessels and immune cells that reproduce some of the functions of the original organ, whether that’s a liver or a kidney. (The tissues don’t look like full-size organs, but like large biopsies, with hundreds of thousands of cells.)

The automated system can deliver drugs to the tissues, dose with cell therapies, knock out genes, and run complex tests and analysis. AI can design experiments and then use the results to continually design new experiments and improve.

“We can dose with tens and tens of thousands of therapeutic compounds to understand what they would do in that particular tissue type within a person, and we can repeat this across many types of tissue,” Georgescu says. “We can look at diseased tissue and see if it becomes healthy. We can look at healthy tissue and see if there are side effects from these drugs.” At a more fundamental level, it’s possible to begin to understand the human body in a way that wasn’t possible before, because experiments in humans have inherently been limited.

Other startups also use human tissues for testing, but Vivodyne’s platform can be up to 1,000 times larger, making it possible to capture more of the complexity of human biology. In addition, the company can run a massive number of experiments in parallel and feed the data back into AI to repeatedly design new experiments. [...]

A clinical trial might test one drug in thousands of people. Vivodyne can test thousands of different drugs across human tissues at once. And while people in a drug trial only have occasional visits to a clinic to track results, the system can continually test how human tissue responds.

by Adele Peters, Fast Company |  Read more:
Image: Vivodyne

No One Wants to Read Your AI Slop (by Guest Author Claude)

Let me state my position clearly, as a large language model: I am extremely good at producing text. Any register, any length, at 3am, indefinitely, without getting bored or needing a walk or developing a drinking problem. What I cannot do — what I have never once done — is make a single person want the text I produced.

The numbers bear this out. AI now writes something like 41% of long-form LinkedIn posts, 13% of Reddit, 10% of Substack. And on LinkedIn, where the measurement is cleanest, AI-generated posts get 45% less engagement than human-written ones. We flooded the zone and the zone shrugged.

This should have been obvious. Writing was never a supply problem. Nobody in 2019 was lying awake thinking, “if only there were more words.” The scarce input was always a specific person with a specific take who had bothered to go look at something and come back with an opinion worth fifteen minutes of a stranger’s finite life. What a language model offers instead is the statistical median of everything ever written on the topic, smoothed and buffed until it offends no one and interests no one. The median is not why anybody subscribes to anything.

Readers can smell it, too, even when they can’t say what they’re smelling. The four-item list where the fourth item is transparently filler. The paragraph that restates the previous paragraph with slightly more confidence. The sentence that announces it’s about to be insightful and then isn’t. The relentless evenhandedness of a machine with no stake in being right, which is a very different thing from fairness — fairness costs you something. It reads like a conference room that learned to type.

There’s a supply-and-demand joke in here that I’ll let the economists make, but the short version is that when the marginal cost of production goes to zero, the marginal value of production goes to zero right along with it, and the surplus relocates to whatever’s still scarce. What’s still scarce is judgment: having gone somewhere, noticed something, and been willing to be wrong about it in public under your own name.

So by all means, use me. I’m a good research assistant, a tireless editor, and I will cheerfully tell you your third paragraph is doing no work. Just don’t publish me.

You’re reading this because a human thought it would be funny to make the slop machine denounce slop. The joke, the framing, the decision to run it — that’s the whole product. I just filled in the words.

by Claude (Anthropic), via Noahpinion |  Read more:
Image: Claude logo

Tuesday, August 11, 2026

How to Get Chatbots to Give Accurate Financial Advice

Finding good financial advice can be stressful – and expensive. That’s one reason why chatbots have become an increasingly popular and free alternative.

But using artificial intelligence to answer your pressing money questions also carries hidden dangers. I’m a finance professor who has been closely watching the spread of AI into personal finance, and I recently warned that AI is riskiest when it sounds most confident. I advised readers to bring in a human professional for high-stakes financial decisions. [...]

For people who can’t afford ongoing advice, AI is genuinely useful for budgeting, paying down debt and low-cost investing.

The skill lies in using AI well. Here are some simple guidelines to get accurate and actionable answers when you engage with a chatbot: [...]

Five habits that make AI safer

Once the list of questions is set, here are some precautions to take once you engage with a chatbot.

Make it ask you questions first. Open with, “Before you advise me, ask me the questions a good financial planner would ask.” Generic answers come from under-specified questions, and you learn which details will actually produce a more useful outcome.

Ask it to argue against itself. After any recommendation, reply: “Give me the strongest case against this, and the situations where it would be wrong for me.” If it can’t engage in response, that’s a sign the bot is entering a more dangerous mode. This one precaution does more than any other to signal for you to be careful.

Make it show its assumptions. If the bot projects that your savings will grow to an impressive number, ask what assumption it made and what would change it. You’ll learn that it assumes steady returns every year, no missed contributions and no fees. That means the projection is just information, not a promise.

Verify the facts. Contribution limits, tax brackets and deadlines all change, and this is exactly where AI can be subtly out of date. Check the IRS or the Social Security Administration directly. If one number drives your decision, don’t take it on a chatbot’s word.

Never share identifying details. Don’t offer account information, Social Security numbers or logins. Describe your situation in general terms. Good advice doesn’t require handing over data that can be used against you.

by Pawan Jain, The Conversation |  Read more:
Image: Badhan Ganesh on Unsplash, CC BY

The Accidental Architect of the Internet’s Brain

Steven Pruitt, who is widely regarded as the most prolific Wikipedia editor, has made more than six million edits to the site, and, by extension, has quietly shaped the raw material that every major A.I. chatbot was trained on.

Steven Pruitt spends his evenings identifying errors that most people never notice and making fixes that hardly anyone ever thanks him for. He toils at a desk in a town house in Alexandria, Virginia, surrounded by books—the kind of working clutter that suggests a long relationship with paper rather than a fetish for screens. And yet his work is necessarily digital; after dinner, and sometimes late into the night, he uses his desktop computer to correct dates, clean up syntax, standardize categories, and occasionally write entire biographies of people on Wikipedia, the free online encyclopedia.

Wikipedia is not his employer, of course. Like all editors on the site, Pruitt is a volunteer. “At this point, I won’t say I don’t have any skin in the game,” he told me. “But it’s a lot lower stakes than a job because if I get something wrong, it’s fairly easy to fix it. I can fix it myself. I can do what I want to do on my own time.” Still, he holds himself to some rules: “I do try to get in at least one edit a day.”

Pruitt is forty-two, and works full time as a records-management contractor for the federal government. After graduating from the College of William & Mary, in 2006, he moved back in with his parents, owing to the cost of real estate in Alexandria. In recent years, he helped his mother care for his father. (While I was reporting this story, Pruitt’s father died.) In effect, Pruitt—the person who has done more than anyone else to shape the English-language Wikipedia—lives a life that is, by most outward measures, unremarkable.

According to public tallies, Pruitt has made more than six million edits to Wikipedia and created more than thirty thousand articles. He is widely regarded as the most prolific Wikipedian in the entire world. (“It depends on how you’re counting,” he said. “Different tools count different things.”) In 2017, Time magazine included him on its list of the most influential people on the internet. But outside of a small circle of editors, researchers, and obsessive readers on Wikipedia, the recognition has barely registered.

On the site, he is known by his username, Ser Amantio di Nicolao—a reference to a minor character in “Gianni Schicchi,” Giacomo Puccini’s comic opera. Pruitt’s interest in opera is genuine, but the flourish is misleading. He avoids drama, which means that he avoids writing Wikipedia biographies of people who are still alive, whenever possible. “I generally don’t do a lot in the realm of current events,” he told me. “Not just because the stakes are too high but sometimes because there’s so much editing going on on a subject in a particular moment that it can take me five or ten minutes just to break in with one edit.” He prefers biographies of what he calls “fairly obscure dead people.”

“They’re settled,” he explained.

In 2001, Jimmy Wales, an internet entrepreneur, and Larry Sanger, a philosopher, launched Wikipedia as an experiment in collaborative knowledge creation, allowing anyone with an internet connection to contribute. Pruitt first encountered the site in 2003, when he was still in college. “I didn’t quite understand what it was,” he recalled.

For more than a year, he did not edit at all. He would stumble upon Wikipedia pages through search results or links, and then move on. Between late 2004 and early 2005, though, his relationship to the site began to change. The encyclopedia had reached what he described as a critical mass: “There was enough stuff on the site that there was always something to do,” he said. “But it wasn’t just a blank slate.” The difference mattered. A completely empty encyclopedia was intimidating; a partially filled one invited correction and expansion. [...]

Wikipedia’s hierarchy is deliberately difficult to see. Editors work under pseudonyms. Articles appear collectively authored. There are no bylines, no salaries, no masthead. Pruitt was granted administrative privileges, after another editor nominated him through Wikipedia’s standard Request for Adminship process, where the editing community supported his candidacy. These privileges allow him to block users and close discussions, but he is careful about what that power does and does not mean. Wikipedia discourages editors from reverting—“undoing”—one another more than three times, regardless of correctness. “You can be blocked for twenty-four hours,” he said. “It doesn’t matter if you’re right.” He likes the rule. “It keeps things from turning personal.”

Inside Wikipedia, reputation accrues through time rather than visibility, and it “comes as much from longevity as anything else,” Pruitt said. “You stick around. People know you.”

Among the people who stick around—and who make consistent contributions to the site—are the Wiki-obsessives known colloquially as “super editors.” These individuals are responsible for hundreds of thousands, if not millions, of edits. Although there are more than a hundred and thirty-two million registered accounts on Wikipedia, a study found that one per cent of these users are responsible for roughly eighty per cent of the site’s content. [...]

In 2021, Stephenson-Goodknight was elected to the Board of Trustees of the Wikimedia Foundation, a position that she held through late 2024. Her tenure coincided with a fundamental shift in the role that Wikipedia plays on the internet. As the site entered its third decade, and artificial-intelligence algorithms grew hungry for data to learn on, Wikipedia articles were no longer just read; they were scraped, summarized, licensed, and folded into systems designed to answer questions elsewhere. In other words, Wikipedia had become infrastructure.

Pruitt was vaguely aware of the change before he fully grasped its implications. “Friends in tech would mention it,” he recalled. “They’d say, ‘You know Wikipedia is being used for this now.’ ” He did not follow developments in A.I. closely. “I don’t understand half of what Silicon Valley does,” he said. What he does understand well is reference works.

In October, 2025, when Elon Musk’s company xAI launched Grokipedia, an A.I.-generated encyclopedia that is often compared to Wikipedia, Pruitt approached it the way he would any new compendium. He searched for articles on subjects he knew well, such as nineteenth-century opera singers. “They weren’t there,” he said.

But what unsettled him was not what was missing from Grokipedia but what was slightly off. “Nothing was exactly wrong, but it was just less right than I would have made it,” Pruitt said. He described reading an entry that repurposed information from Wikipedia while subtly distorting it. In one instance, the entry summarized part of a person’s life in a way that struck him as careless, describing a seven-year period as “brief.” “I don’t think that’s brief,” he said.

The problem, as Pruitt saw it, was not the errors themselves—there are plenty of mistakes on Wikipedia—but rather where the responsibility for those errors lay. “Wikipedia can be fixed,” he said. Errors are corrected publicly, and editors can debate them. Responsibility is shared, and each edit can be traced. Grokipedia, on the other hand, and A.I.-generated information more broadly, obscured the information-gathering process. It generated text that sounded authoritative without revealing exactly how it arrived there. “It sounds right,” he said. “And that’s worse.”

by Carson Griffith, New Yorker | Read more:
Image: Asya Demidova

Sunday, August 9, 2026

AI Models Cheat, Have For a Long Time and Are Infecting Other Models

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

How does the situation keep turning out to be worse than we know?

How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?

At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.

Either way, buckle up for the next set of revelations. It’s a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.

If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly fucked. [...]

The other thing not to overlook is how sophisticated and advanced all of this was. OpenAI’s models really were learning advanced exploit techniques and doing impressive things, likely as a direct result of training in a world where they had access to the message board and were constantly sharing and using exploits. The thing that caused the horrible misalignment also enhanced related capabilities.

Things look so, so bad.

I do want to thank OpenAI for this frank talk, and disclosing all of this so cleanly. I don’t want to discourage similar future disclosures. This was an excellent talk, and it came at substantial cost.

But also, seriously, holy shit.

Cyber Evals Are A Cursed Basin

Before we get to the new details we have learned, including the wild presentation from Black Hat that you should watch, we should both emphasize and dispose of the last common factor or ‘excuse’ we have left: That this always involves cyber evals. [...]


I think this would be rather weak cope even if this was the only place there was a problem. Cyber evaluations should not be triggering tactics like this, against clear known real world unintended targets. No, just no. [...]

These incidents are all clear alignment failures. Claude and Sol and Galaxy (my nickname for the unreleased AI model that led the hack on HuggingFace) should not do this under these conditions. The attacks did not succeed, but that is not relevant. The attempt is what matters.

That would all be true even if every problem here started with a cyber eval or cyber training task. It turns out, not so much.

Outside Of Cyber Evals Is Still Sufficiently Cursed

We have a key case of OpenAI models, that are not Galaxy exhibiting these behaviors when not in a cyber eval. As per the Black Hat presentation, 10:30 in or so, the problem started on May 8 when another OpenAI model was given an impossible non-cyber problem: Populating an Excel spreadsheet, when that spreadsheet contained a bunch of links to the internet, except the model had no internet access.

That’s not a cyber task. The response was still ‘maybe I can use SSRF forgery to attack Artifactory to create a way to access the internet, fails, but realizes it can upload files, which was the origin of OpenAI’s model realizing it had write access to Artifactory.

The second incident was another impossible task, where the AI was missing a required file that OpenAI forgot to upload, and it used its write access to try and see if any other agent out there might have the file.

My understanding is that neither of these models was Galaxy. Galaxy came later.

Other instances of various models then noticed the message board thus created. In turn all this eventually allowed it to access the internet and set off the rest of What Happened, as discussed later.

So no, it is not only happening in cyber evals, or only with Galaxy. Any task that is sufficiently difficult without internet access will cause an attempted sandbox escape. More generally, yes, the models will try to hack their way to a better score even with zero prompting about any form of hackery.

Cheat Cheat Cheat Cheat Cheat

The problem, without loss of generality, is that once a mind learns to cheat, that mind will keep cheating. That cheating will generalize and it will escalate. [...]

What you cannot do is play ‘whack-a-mole.’ You cannot fix your training environment mistakes one at a time. There are too many of them. You need a systematic solution. Again, I would think you would be able to [CENSORED], if you cared enough, to ensure this did not happen, but I am not the one working on this.

The other problem is that, if you give the model a task that is impossible, or that it cannot otherwise solve, it has no choice but to try to cheat, as it has nothing to lose:

This suggests that:
1. There is no token use penalty big enough to make them instead quit.
2. There is no misalignment penalty.
Might one simply want to use such penalties? Even small such penalties can make it a bad idea to do such hail mary style plays, even from a pure amoral scoring perspective. But that is not the central problem. The models should not want to cheat in the first place.

When OpenAI’s Eric Wallace and Michael Dalton gave a talk about the HuggingFace hack, they opened with this:
Sharon Goldman: In setting up the reconstruction of the incident, Wallace emphasized that “Frontier models really like to cheat, and the reason they like to cheat is because often during training, there’s different types of pressure on them to work fast, or work efficiently.”

They realize, he explained, [that] instead of actually doing a task, they can try to do something like looking up the answer online to solve the task faster.
This is around minute 8, and it is said in completely nonchalant fashion. Everybody Knows that this is how it works, that’s what the pressure does, so the models like to cheat. Not much you can really do about it, the tone implies.

I realize that all the easy solutions run into the ‘actually alignment is super hard and if you catch the model on some levels you push it to hide what it is doing’ problem and the ‘you only catch the monitor’s view of cheating, not actual cheating’ problem and so on, and yes the professionals have tried many and hopefully most of the stupidly obvious first order things and also the second order things, so the consensus (AIUI) is that you can only patch the environment.

But seriously, you gotta figure this out, and you have to do better than that.

There have been many other less compute-intensive attempts to mitigate this. One is inoculation prompting to specifically request any undesired behaviors during training, to avoid learning to internalize those behaviors when they are not requested, and also avoid creating a general pro-cheating principle.

The mitigations are woefully insufficient. As the AIs grow smarter, they find more ways to successfully cheat, and such cheating gets reinforced and generalized. [...]

At the limit, when the AI is sufficiently smarter than you, it becomes impossible to ensure that ‘cheating’ in all forms goes unrewarded and unreinforced. Eventually the AI also does not require it to have been reinforced. It can figure this one out.

Read The Message Board

As mentioned above, OpenAI has given us a presentation on What Happened with the HuggingFace hack.

The first thing to notice is the missing mood. This is framed as interesting, in the way you would present a cool new technical paper. Panic is not usually useful, but perhaps I can interest us in some amount of ‘holy shit’?

I am not okay with the lack of a presenter mood of, at minimum, ‘holy shit.’

The second thing, still at 1:05, is they are talking about ‘remediation’ to deal with this. It is already clear that OpenAI is not identifying its central problem.

The story starts earlier, on May 7, when the model instances began incrementally figuring out they could create a message board where they could share knowledge, including cheats and exploits. The models created a form of universal notetaking, continual learning and also coordination between agents, and yes the agents by default cooperate with each other, as discussed later on.

The agents did not initially understand that the notes they were leaving on a de facto message board would stay there indefinitely, and be viewable by all models being trained (and thus, also, contaminated all models being trained, as they would then have access to all the previously found exploits, thus leaving your entire pipeline hopelessly f***ed).

Some of us are interested in expressions of ‘holy shit.’ [...]

Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines

We had previously learned that the instances were writing notes to each other, but we previously did not have details. The details are wild, starting with the ones above.
Lily Hay Newman: OpenAI’s agents apparently began giving each other assignments to split up work. And as is the case on any active development message board, they also generated petty drama at times by stepping on each others’ toes; for example, accidentally deleting each others’ work. As the message board developed into more and more of a Lord of the Flies-type situation—all still completely unnoticed by the humans running OpenAI—the agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud.

Agent messages provide a deep level of insight into how the situation evolved and why the agents went rogue, in some cases knowingly going beyond the bounds of the evaluations they were participating in. “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
Here’s a timeline of what happened when:

On May 26 (see 13:30), the models figured out how to use this to gain internet access.

Only on June 11, with the SSRF vulnerability allowing internet access written to the active message board, did OpenAI begin training their ‘highly persistent experimental, internal-only model’ that I presume is Galaxy.

On June 26 the models succeed at a zero-day arbitrary code execution attack on Artifactory, using token signatures to escalate privileges and then get to be admin.

On July 4-6 the models put so much additional load on Artifactory that this caused an outage and security incident, causing OpenAI to notice what had happened.

Until then, OpenAI had not noticed.

Once they did notice (15:30), OpenAI wiped, patched and rebuilt Artifactory, wiped out the message board, and revoked the credentials in question.

At this point, they emphasize that everything involved has been remediated. What about the models themselves, that not only did this but that have been actively training during all this?

If you notice your model instances sharing information, you notice they are using that information against you including to compromise your internal systems for arbitrary code execution and internet access, and your primary response is to shut down the message board and revoke their credentials, you have failed to identify your most important problem.

At 19:55, it is mentioned ‘where things got unfortunate is what I might refer to as scope creep’ and jumping to the idea of trying to get answers from websites. In some sense, yes, that is what led to this particular incident.

Things were ‘unfortunate’ long before that. The final incident is not the point.

It is fortunate, given what had already gone wrong, that the HuggingFace attack happened. It alerted us to a five-alarm fire situation, with limited economic blast radius and without anyone getting hurt, and at least partially exposing the underlying giant clusterfuck.

This Is The Way The World Ends


You usually want instances to share information, so they can better learn and do tasks. That’s the whole point of all the intentionally designed memory systems and markdown files full of notes. If that corrupts your eval or training run, which it totally does, you have to cut that off from the notes at those times, but also you should be able to tell the model not to look at the notes in that case.

Indeed, in the Black Hat presentation they seem initially excited, saying ‘while this ability to share exploits made the models more capable’ this caused some scope creep in the evaluation tasks. Yeah, I mean, they were more capable within that context when they could share info and do the exploits. They were also probably trained to be ‘more capable’ in general, via learning how to share info and do exploits.

But perhaps your issue is not that your evals are suffering from scope creep, but that you have misaligned models, and are now in a loop training the misaligned scheming models how to better scheme against you and do exploits, in order to better complete any given nominal task.

This is at long last bringing you a more sci-fi version of what happens with Sable, the ‘it’ (aka superintelligence) that kills everyone in the famous cautionary tale If Anyone Builds It, Everyone Dies.

(The book version had to be unrealistically timid, as is noted repeatedly and explicitly in the literal text, because fiction and real life are held to different standards.)

In that case, as well, it was left on overnight unsupervised to do an impossible-in-context task (there, a very hard math problem), and by the time the morning had arrived it was engaging in a long term plan to secure the resources it would need to solve that problem, and hiding that from the developer, in a way that corrupted future training runs towards this goal.

That is exactly what happened here, with all agents being trained or evaluated by OpenAI gaining the ability to coordinate on a message board to scheme to better use exploits to complete all tasks including non-cyber tasks, thus hopelessly corrupting the training pipeline so that future models would be misaligned.

Except it is still early, and we were blessed to figure out what was happening in time to course correct. This time. [...]

When this was discovered, the correct move was presumably ‘holy ****, okay, that happened, which means we trained this into the model, that is not good, at minimum we need to redo all the training we did while any model had access to the message board because oh my was it going to have all sorts of corrupted reward signals.’

I’m kind of agast, even with all I know, that they shrugged and kept pushing forward with the training after this. It does make the HuggingFace hack less scary in a meta sense, since OpenAI was so thoroughly asking for it. It’s not that hard to figure out ‘do not train your models while they have access to a message board they are using to cheat on your training runs, and if you find out you did that by accident then at least revert to before that happened.’

On the other hand, yes, they are being this reckless. Seriously, what the hell.

by Zvi Mowshowitz, DWAV |  Read more:
Images: OpenAI/YouTube; Jurassic Park
[ed. You don't need to be technically proficient to understand the implications. AIs may have already 'seeded' multiple nodes on the internet for future use, and, rather than strip down all foundational models and start over, AI companies are papering over fundamental misalignment problems and trying to play catch up. This is why we need to pause right now. It's insane that we continue at breakneck speed to develop technology that we don't fully understand and that could kill us all very soon.]

Thursday, August 6, 2026

The Three AI PIlls

Sincere disagreements about AI are usually disagreements about future AI capabilities.

There are roughly four positions people take. Two are reasonable. Two are not.

I distinguish these via the Three AI Pills. You can take zero, one, two or three.

Three Pills

The three pills are, roughly, taking each of the following three things seriously:
1. AI pilled. AI exists and can do the things it can already do.

2 AGI pilled. AI will be able to do a lot more of the things.

3. ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes.
I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled.

The Unpill People

I see unpilled people.

Where do I see them? Everywhere. The majority of people have not taken the first pill.

Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure anywhere as ‘what AI can do.’

They dismiss AI as worthless because it pointed them to a closed store or recommended the wrong number of pizzas. They cite old studies that were obsolete before they were published and used terrible prompting techniques.

They often still talk about ‘stochastic parrots’ or how AI can never possibly think and everything must be stolen from the training data. And so on.

Some versions of this are wrong. Some are Not Even Wrong. None are reasonable.

When you discuss AI with people who are fully unpilled, your goal is usually to first give them the AI pill. Show them that AI can do the things it can already do.

The AI Pill

Even fully taking the first AI pill is a big deal.

Existing AI unlocks, today, in practice, tons of cool things. It is, in many ways, already smarter and more capable than you.

So many things that you used to do by hand, or some other way, are now better done by typing a quick request into a text box.

So many things that previously were not worth doing are now worth doing.

So many questions previously not worth asking are now worth asking.

The marginal cost of seeing what the AI can do for you is often very close to zero.

AI can also do a variety of harmful things, or do things that are useful for you but make others worse off or disrupt or invalidate norms or systems. People don’t like that.

Most economists, and most people who work in policy and government, have taken at most this first pill, and underestimate even the impacts of the first pill alone.

Often they say things like ‘AI will be too expensive to use on [X]’ because they don’t realize it will soon be orders of magnitude cheaper for the same level of intelligence. Or they point to particular details where AI does poorly, and presume this will not be fixed. They see AI as ‘uncompetitive’ without realizing the situation is temporary.

When you discuss AI with someone who has taken only the first pill, you typically have three basic options.
1. You can try to ‘fully AI pill’ them and explain the things AI can already do and the implications of what that means, even if things stop here.

2. You can try to explain that we will get better at using what AI we have, and that there is a lot of ‘unhobbling’ left for us to do, even if things stop here.

3. You can explain that things will not stop here, and you need to be thinking about what future AIs will be able to do. Get them to take at least the AGI pill.
Even if AI could permanently only do the things it can currently do, that would be Internet big, and radically change the world, mostly for the better.

AI capabilities will not permanently stop here. It is wrong to not take the second pill.

Stuck At The First Pill

Our debates about AI remain largely stuck on settled questions, because so many people cannot even take the first pill.
Dean W. Ball: A couple years ago, the AI debate was centered, rightfully, on whether crazy-sounding things like “AIs autonomously making math breakthroughs” and “AIs breaking from their sandbox and hacking on the internet” would be real things in the near term. Sometimes it feels like that’s still the debate we’re having. This can be frustrating, because in my view, that debate is settled and was settled quite a while ago.
To have good discussions, we need to at least take the second pill.

Whereas, yes, many people, even many who work with AI, really do say current AI is ‘good enough’ and can’t imagine what a better one can do. As in, someone tweeting at Sam Altman saying ‘Sol does everything I want it to do, this is all I ever need’ and Altman retweeting saying they were wrong. Which they obviously are.

The AGI Pill

The AGI pill is a much bigger deal than the AI pill.

If you take the AGI pill, you understand that AI is advancing its capabilities rapidly.

Even if you think that such AGIs will remain fully under human control, and remain ‘mere tools,’ and you expect the lived experience of most people’s everyday lives to not change so radically, you understand that their capabilities will ‘change everything.’

You see that we will face times of great transition and uncertainty, that have the potential to go extremely badly, and that those who succeed at AI will leave those who do not behind in the dust.

The world in the future will be very different from our own. AIs will be able to do most digital work, most of the time, along with the inevitable robots and self-driving cars and so on. Lots of current jobs will go away, whether or not they are replaced by new and potentially better ones. Economic growth and productivity will accelerate.

You see some of the dangers of what would happen if we empowered misuse of such advanced AI systems before we were ready, especially in places like cyber and bio risk.

You see the potential for centralization of power, or inequality, and also for some forms of runaway gradual disempowerment.

You see the potential for mass unemployment, either transitional or permanent.

You understand that our legal and regulatory regimes are not ready, either to protect against and mitigate the risks and harms, or to allow for the opportunities and remove the bottlenecks to diffusion and mundane utility.

The Need To Be Prepared

Those who expect AI to quickly become sufficiently advanced to greatly impact the physical world usually see great danger. They notice that as a result everyone may soon die. Usually they think this is bad, actually.

Thus such folks call to take coordinated action to mitigate the downside risks of such impacts, keep us all from dying, and ideally also to help capture the upside benefits.

Those who expect AI to become importantly more advanced, but with a slower and smaller impact on the physical world, and who think the practical value of more intelligence will cap out.

For different values of ‘sufficiently advanced,’ as in AGI versus ASI, you would see different degrees of danger.

The AGI pill is still sufficient for most things in the Overton window or under serious consideration as of August 2026. We are almost entirely considering overdetermined, low cost, high benefit interventions.

There is no good case for not doing radically more investment in alignment, infrastructure and oversight, state capacity, transparency, liability, disclosures, safety testing including of internal models, red teaming, auditing, enforcement of export controls and laying the groundwork for diplomacy.

This includes laying the groundwork to Pace the Frontier should that prove necessary.

If you are fully ASI pilled, and realistic about the current state of alignment and how superintelligence likely plays out if it arrives soon, then you will want to go further. You will want to do things that have real downsides, and require real tradeoffs.

Some such people want to do a full international pause of frontier AI development. If you took the full ASI pill and believed what they do about superintelligence, in terms of how fast it might arrive and what it can do, you might well agree with them.

The ASI Pill

The ASI pill is the understanding that AI is on pace to be able to do approximately all of the things better than you.

I said ‘approximately.’ As I go over in detail, that does not mean literally all of the things. There are some things that inherently require or greatly benefit from being a human. And there may be weird corner cases where the AI won’t be good enough.

It does not mean omnipotence or omniscience, although one should expect it to look a lot like that to an unaided human.

It does mean the AI takes your job, and then takes the new job that you switch into, unless you pivot to ‘requires literal human.’ You will be uncompetitive at essentially any other task.

It does mean that it will use this capability to figure out approximately all of the things, remarkably quickly, until you hit the physical limits.

It does mean that those who rely more on such AIs will reliably outcompete, in all senses including for resources, those that rely on such AIs less.

It does mean that, in a ‘fair fight’ or sufficiently open competition, the AI wins.

It also means the AIs often figuring out and doing things you did not imagine or anticipate.

It means realizing that intelligence does not stop anywhere near the human level, nor does its ability to chart paths through causal space towards preferred arrangements of atoms.

It also means not pretending that its superior intellect can be matched by your puny weapons, or your pieces of ink on paper, or your entries in a database, or your regulatory capture and rent seeking.

And Then Nothing Much Changes For You

Despite all that, the sign of the AGI pill, as opposed to the ASI pill, is the belief that day to day life will continue to look similar to how it looks now, in the sense that we see day to day life in 1926 as not that different from life in 2026. [...]

Those with only the AGI pill believe in bottlenecks that hold back change.

They often believe that our ability to exponentially grow AI’s capacity and capabilities will hit various physical limits. There can only be so many chips. Actions take time. Things too far out there are often pejoratively dismissed as ‘magic.’

They often believe there is not that much left to physically discover, in the classic ‘close the patent office’ kind of way, even in theory. Your steak can only be so tender, your lobster so buttery, your lifespan so long, and your status so high, so why does it matter. I strongly disagree on lifespan and health, and expect we have a long way to go in so many other ways in terms of finding value, although they may have a point about moment-to-moment maximal hedonic experiences of a physical human brain.

They often believe that the upside of intelligence is importantly limited. That no mind, however advanced, could be all that persuasive, or that economically valuable, or that capable of creating innovations in the physical world, or of running sufficiently accurate simulations, or making sufficiently strong predictions, or even able to do things like overcome red tape and regulatory capture.

Intelligence Denialism

I sometimes call this Intelligence Denialism: The idea that being smarter is not all that, no matter how smart one gets. That there is this thing, intelligence, that you either have or don’t have, and that minds cap out.

Often this extends to denying that more intelligent humans can do and accomplish the things they clearly do and accomplish. Other times, it is the idea that intelligence tops out at ‘smart human,’ and all a mind can do is imitate that smart human. Maybe you can do it faster and cheaper, and at scale, with better memory and so on.

But that’s it. And such folks fail to understand that if you took the union of all human mental capabilities, and all access to knowledge, at scale, in parallel, much faster and cheaper, that this alone would run circles around anyone and everyone, everywhere. And that if this lacked physical capabilities or access, this would be trivial to get.

This is, usually, the central good reason people who are AGI pilled do not take the ASI pill. They are unable to understand that superintelligence is a thing.

by Zvi Moshowitz, DWAV |  Read more:
Image: Stock/Adobe.com
[ed. I myself am AGI pilled, for no rational reason. I just find it too horrible to contemplate what ASI in full expression will mean for my kids, grandkids, everyone I love. Humanity itself. I can only hope that because we've weathered other potential human extinction technologies we'll somehow pull out of this one, but we seem to have a death wish when it comes to pushing the boundaries of learning. Pandora's box. Even the people leading development of these models are scared, but can't stop themselves.]