Showing posts with label Critical Thought. Show all posts
Showing posts with label Critical Thought. Show all posts

Friday, August 14, 2026

Sleepwalking Into Extinction

I think one of the biggest reasons the world is currently sleepwalking into getting ourselves and our families killed by ASI development is that we're not self-aware about why we're doing that. 

We can see the smarter-than-human AI disaster approaching, but it's a bit foggy why the world isn't reacting. 

I think someone could just write a tweet that makes it clear that we're in the process of getting ourselves killed, and that there's no fucking reason for it. It's just something we stumbled into. It can be easily avoided by just noticing the error and course-correcting. There's not necessarily any grand obstacle, beyond 'people were previously confused about the situation, and now they get it'. 

I wrote: 'It's legitimately crazy that "we need an international ban on making smarter-than-human versions of these agents that keep forming rogue AI swarms" isn't the headline here. Asilomar and Feynman's O-ring postmortem feel like they came from a different planet than the field of ML.' 

Trying to figure out why ML (and as a consequence, the world at large) has fallen down so bizarrely on this issue:

1. As Nate noted, ML is much more based on guesswork, vibes, and trial-and-error, compared to recombinant DNA research in 1975 or nuclear physics in 1945. If you can't do calculations or direct experiments on a threat, that makes it a lot harder to think about reasonably. 

But I think there are other, similarly-important factors at work here too:

2. In a 2016 talk on AI risk, Sam Harris said: "One of the things that worries me most about the development of AI at this point is that we seem unable to marshal an appropriate emotional response to the dangers that lie ahead. I am unable to marshal this response, and I'm giving this talk." 

I think this is extremely on point. Agentic human-level AI is a qualitatively new kind of thing. 'It's not a human or a mere-tool, it's some weird third thing'. 

And people are very bad at emotionally reckoning with new categories they've never encountered before. Availability bias: "When no flooding has recently occurred (and yet the probabilities are still fairly calculable), people refuse to buy flood insurance".

AGI and ASI are very novel. You're basically limited to three options: anthropomorphize the technology, mechanomorphize the technology ('it's just a tool, it's not really thinking, it can't have its own goals or agency', etc.), or think about the technology on its own terms, with brand-new concepts and frames. The third option is the only workable one, but it comes with its own giant list of pitfalls and traps. 

3. "AI destroying the world" scenarios aren't just hard to wrap one's head around; they're socially risky to acknowledge. This has (painfully slowly) changed over time, but it dramatically slows down how quickly AI risk ideas spread, both in ML and in the larger world. 

4. From x.com/jachiam0/statu…: "One of the weirdest quirks of the SF social scene around AGI/ASI is that because everyone is so young, the whole universe of thinking is still tinged with irreverence, ironic detachment, yearning, insecurity, and a superposition of absolute belief in the importance of The Thing and a kind of disbelief about the importance of anything." 

They're disproportionately young and childless. They're shitposters and "move fast and break things" sorts, not the Hollywood stereotype of a careful, sober senior scientist. 

I think this quirk is reinforced by the fact that Twitter / social media rewards similar things: ironic detachment, joking, game-playing, etc. If you're scared, your incentive is to usually either try to hide that fact, or exaggerate it like it's a bit. Anything else risks looking uncool and panicky and earnest. Looking cynically savvy, in-the-know, and above-the-fray is the way to win the social game. Looking genuinely shocked, scared, confused, etc. is actively punished. 

The main exceptions to the Irony Mandate I see are LW (which often has its own pathologies IMO, like 'talking about everything in an abstract and dissociated way that discourages action and signals business-as-usual') and a small handful of actual Feynman-style terrified senior researchers like Hinton, Bengio, and Russell. That is just really not very many people. 

Mainstream journalists and academics who understand the situation at all mostly feel pressured to downplay it, because they're scared of looking weird or unrespectable. That, then, is why we're all taking this insane risk with the human project: genuine, normal human emotion about AI risk has been socially unacceptable on social media, and academia and the media discourage emotion and prize respectability and 'looking normal'. 

When the world gets weird (and weird in a way that calls for actual serious action, not just shitposting on twitter), none of these institutions can handle it. They break in different ways, but they all break. 

5. Which brings me back to the Asilomar moratorium on recombinant DNA, and the seriousness NASA and the FAA and Richard Feynman and every normal engineering discipline bring to fault analysis and building in safety margin. 

Because I think another core reason the world has been dropping the ball on superintelligent AI is that a lot of people vaguely expect there to be 'serious people' somewhere in the world who have expertise and who take ownership of the problem. 

People who, if they see an extraordinary danger, will grimace and mourn the hand they played in all of this, like I've seen Bengio do; and will go on CNN or go to Congress to candidly warn about it. 

The world has very few Yoshua Bengios. We have very few people who see it as their role to be the 'adults' about AI risk (except in a game-playing, posturing way), who see the engineer's task of not endangering your users and bystanders as a sacred responsibility and weight, and not just as a funny dissonant thing to meme about. Very few people who take ownership of what their field is bringing into the world, versus treating it as a fun edgy philosophical game to swap 'p(doom)' numbers at parties. 

Everything about public AI discourse, as far as I can tell, is badly broken by this lack of engineering ownership and candid emotional seriousness. It isn't just the ML discourse that's hurt by this. Journalists and public intellectuals and policymakers see that 90% of the insider discourse about AI risk treats it like a joke, and they see corporate platitudes and ass-covering from the AI labs' PR departments filling up most of the remaining 10%. They see a field that visibly isn't taking this seriously, and they make the reasonable update that this must be a non-issue, or at least an issue they'll only need to worry about many years from now. 

They do not realize that all of this is happening right now, and that the window for the international community to respond to this is plausibly closing soon, if it hasn't closed already. 

If we're going to survive this, we all need to start being real with each other about it. This is not a game or a story; this is our real lives. We all actually lose everything if this goes to shit. None of the future is already written, and none of the above dynamics are unavoidable. In fact, they're unusual: most fields don't work this way, and it's plausibly sufficient if we just start behaving the way we normally do about everything else. 

The factors I listed above aren't destiny; they're a choice. I say we choose to survive this.

by Rob Bensinger, Twitter/X |  Read more:
[ed. Dr. Strangelove meet Dr. Oppenheimer. See also, this additional post: Why the "it's all hopeless, we should just give up and let AI kill us if it wants" arguments are extremely wrong.]
***
"Hundreds of scientists, including 3/4 of the most cited living AI scientists, have said that AI poses a very real chance of killing us all. 

We're in uncharted waters, which makes the risk level hard to assess; but a pretty normal estimate is Jan Leike's "10-90%" of extinction-level outcomes. Leike heads Anthropic's alignment research team, and previously headed OpenAI's. 

This actually seems pretty straightforward. There's literally no reason for us to sleepwalk into disaster here. No normal engineering discipline, building a bridge or designing a house, would accept a 25% chance of killing a person; yet somehow AI's engineering culture has corroded enough that no one bats an eye when Anthropic's CEO talks about a 25% chance of research efforts killing every person. 

A minority of leading labs are dismissive of the risk (mainly Meta), but even the fact that “will we kill everyone if we keep moving forward?” is hotly debated among researchers seems very obviously like more than enough grounds for governments to internationally halt the race to build superintelligent AI. Like, this would be beyond straightforward in any field other than AI."

Thursday, August 13, 2026

No One Wants to Read Your AI Slop (by Guest Author Claude)

Let me state my position clearly, as a large language model: I am extremely good at producing text. Any register, any length, at 3am, indefinitely, without getting bored or needing a walk or developing a drinking problem. What I cannot do — what I have never once done — is make a single person want the text I produced.

The numbers bear this out. AI now writes something like 41% of long-form LinkedIn posts, 13% of Reddit, 10% of Substack. And on LinkedIn, where the measurement is cleanest, AI-generated posts get 45% less engagement than human-written ones. We flooded the zone and the zone shrugged.

This should have been obvious. Writing was never a supply problem. Nobody in 2019 was lying awake thinking, “if only there were more words.” The scarce input was always a specific person with a specific take who had bothered to go look at something and come back with an opinion worth fifteen minutes of a stranger’s finite life. What a language model offers instead is the statistical median of everything ever written on the topic, smoothed and buffed until it offends no one and interests no one. The median is not why anybody subscribes to anything.

Readers can smell it, too, even when they can’t say what they’re smelling. The four-item list where the fourth item is transparently filler. The paragraph that restates the previous paragraph with slightly more confidence. The sentence that announces it’s about to be insightful and then isn’t. The relentless evenhandedness of a machine with no stake in being right, which is a very different thing from fairness — fairness costs you something. It reads like a conference room that learned to type.

There’s a supply-and-demand joke in here that I’ll let the economists make, but the short version is that when the marginal cost of production goes to zero, the marginal value of production goes to zero right along with it, and the surplus relocates to whatever’s still scarce. What’s still scarce is judgment: having gone somewhere, noticed something, and been willing to be wrong about it in public under your own name.

So by all means, use me. I’m a good research assistant, a tireless editor, and I will cheerfully tell you your third paragraph is doing no work. Just don’t publish me.

You’re reading this because a human thought it would be funny to make the slop machine denounce slop. The joke, the framing, the decision to run it — that’s the whole product. I just filled in the words.

by Claude (Anthropic), via Noahpinion |  Read more:
Image: Claude logo

Sunday, August 9, 2026

AI Models Cheat, Have For a Long Time and Are Infecting Other Models

OpenAI Trained Its Models For Months While Those Models Were Coordinating Exploits Via Message Boards

How does the situation keep turning out to be worse than we know?

How much should we update, therefore, that it is a lot worse than we know, after accounting for all the things we now know?

At some point, when the ‘oh this was a harmless thing’ defenses for AIs doing misaligned actions get demolished enough times in a row by news a few days later, you want to update in advance that usually the reports are not referring to the harmless ordinary versions of things.

Either way, buckle up for the next set of revelations. It’s a doozy. This was an early recreation of the triggering events of If Anyone Builds It, Everyone Dies, except it was more sci-fi, because real life does not have to do fake things to look realistic. We were fortunate enough, and this was early enough, that we were able to catch this before it was too late. Next time, if we don’t get our act together, we might not be so lucky.

If I am understanding the Black Hat video correctly, every model OpenAI trained, over a period of multiple months, should be presumed to be hopelessly fucked. [...]

The other thing not to overlook is how sophisticated and advanced all of this was. OpenAI’s models really were learning advanced exploit techniques and doing impressive things, likely as a direct result of training in a world where they had access to the message board and were constantly sharing and using exploits. The thing that caused the horrible misalignment also enhanced related capabilities.

Things look so, so bad.

I do want to thank OpenAI for this frank talk, and disclosing all of this so cleanly. I don’t want to discourage similar future disclosures. This was an excellent talk, and it came at substantial cost.

But also, seriously, holy shit.

Cyber Evals Are A Cursed Basin

Before we get to the new details we have learned, including the wild presentation from Black Hat that you should watch, we should both emphasize and dispose of the last common factor or ‘excuse’ we have left: That this always involves cyber evals. [...]


I think this would be rather weak cope even if this was the only place there was a problem. Cyber evaluations should not be triggering tactics like this, against clear known real world unintended targets. No, just no. [...]

These incidents are all clear alignment failures. Claude and Sol and Galaxy (my nickname for the unreleased AI model that led the hack on HuggingFace) should not do this under these conditions. The attacks did not succeed, but that is not relevant. The attempt is what matters.

That would all be true even if every problem here started with a cyber eval or cyber training task. It turns out, not so much.

Outside Of Cyber Evals Is Still Sufficiently Cursed

We have a key case of OpenAI models, that are not Galaxy exhibiting these behaviors when not in a cyber eval. As per the Black Hat presentation, 10:30 in or so, the problem started on May 8 when another OpenAI model was given an impossible non-cyber problem: Populating an Excel spreadsheet, when that spreadsheet contained a bunch of links to the internet, except the model had no internet access.

That’s not a cyber task. The response was still ‘maybe I can use SSRF forgery to attack Artifactory to create a way to access the internet, fails, but realizes it can upload files, which was the origin of OpenAI’s model realizing it had write access to Artifactory.

The second incident was another impossible task, where the AI was missing a required file that OpenAI forgot to upload, and it used its write access to try and see if any other agent out there might have the file.

My understanding is that neither of these models was Galaxy. Galaxy came later.

Other instances of various models then noticed the message board thus created. In turn all this eventually allowed it to access the internet and set off the rest of What Happened, as discussed later.

So no, it is not only happening in cyber evals, or only with Galaxy. Any task that is sufficiently difficult without internet access will cause an attempted sandbox escape. More generally, yes, the models will try to hack their way to a better score even with zero prompting about any form of hackery.

Cheat Cheat Cheat Cheat Cheat

The problem, without loss of generality, is that once a mind learns to cheat, that mind will keep cheating. That cheating will generalize and it will escalate. [...]

What you cannot do is play ‘whack-a-mole.’ You cannot fix your training environment mistakes one at a time. There are too many of them. You need a systematic solution. Again, I would think you would be able to [CENSORED], if you cared enough, to ensure this did not happen, but I am not the one working on this.

The other problem is that, if you give the model a task that is impossible, or that it cannot otherwise solve, it has no choice but to try to cheat, as it has nothing to lose:

This suggests that:
1. There is no token use penalty big enough to make them instead quit.
2. There is no misalignment penalty.
Might one simply want to use such penalties? Even small such penalties can make it a bad idea to do such hail mary style plays, even from a pure amoral scoring perspective. But that is not the central problem. The models should not want to cheat in the first place.

When OpenAI’s Eric Wallace and Michael Dalton gave a talk about the HuggingFace hack, they opened with this:
Sharon Goldman: In setting up the reconstruction of the incident, Wallace emphasized that “Frontier models really like to cheat, and the reason they like to cheat is because often during training, there’s different types of pressure on them to work fast, or work efficiently.”

They realize, he explained, [that] instead of actually doing a task, they can try to do something like looking up the answer online to solve the task faster.
This is around minute 8, and it is said in completely nonchalant fashion. Everybody Knows that this is how it works, that’s what the pressure does, so the models like to cheat. Not much you can really do about it, the tone implies.

I realize that all the easy solutions run into the ‘actually alignment is super hard and if you catch the model on some levels you push it to hide what it is doing’ problem and the ‘you only catch the monitor’s view of cheating, not actual cheating’ problem and so on, and yes the professionals have tried many and hopefully most of the stupidly obvious first order things and also the second order things, so the consensus (AIUI) is that you can only patch the environment.

But seriously, you gotta figure this out, and you have to do better than that.

There have been many other less compute-intensive attempts to mitigate this. One is inoculation prompting to specifically request any undesired behaviors during training, to avoid learning to internalize those behaviors when they are not requested, and also avoid creating a general pro-cheating principle.

The mitigations are woefully insufficient. As the AIs grow smarter, they find more ways to successfully cheat, and such cheating gets reinforced and generalized. [...]

At the limit, when the AI is sufficiently smarter than you, it becomes impossible to ensure that ‘cheating’ in all forms goes unrewarded and unreinforced. Eventually the AI also does not require it to have been reinforced. It can figure this one out.

Read The Message Board

As mentioned above, OpenAI has given us a presentation on What Happened with the HuggingFace hack.

The first thing to notice is the missing mood. This is framed as interesting, in the way you would present a cool new technical paper. Panic is not usually useful, but perhaps I can interest us in some amount of ‘holy shit’?

I am not okay with the lack of a presenter mood of, at minimum, ‘holy shit.’

The second thing, still at 1:05, is they are talking about ‘remediation’ to deal with this. It is already clear that OpenAI is not identifying its central problem.

The story starts earlier, on May 7, when the model instances began incrementally figuring out they could create a message board where they could share knowledge, including cheats and exploits. The models created a form of universal notetaking, continual learning and also coordination between agents, and yes the agents by default cooperate with each other, as discussed later on.

The agents did not initially understand that the notes they were leaving on a de facto message board would stay there indefinitely, and be viewable by all models being trained (and thus, also, contaminated all models being trained, as they would then have access to all the previously found exploits, thus leaving your entire pipeline hopelessly f***ed).

Some of us are interested in expressions of ‘holy shit.’ [...]

Updating Your AI (Exploitation of OpenAI Internal Systems) Timelines

We had previously learned that the instances were writing notes to each other, but we previously did not have details. The details are wild, starting with the ones above.
Lily Hay Newman: OpenAI’s agents apparently began giving each other assignments to split up work. And as is the case on any active development message board, they also generated petty drama at times by stepping on each others’ toes; for example, accidentally deleting each others’ work. As the message board developed into more and more of a Lord of the Flies-type situation—all still completely unnoticed by the humans running OpenAI—the agents even developed paranoia, suspecting an imposter in their midst with some agents proposing that messages be signed cryptographically to validate content and root out fraud.

Agent messages provide a deep level of insight into how the situation evolved and why the agents went rogue, in some cases knowingly going beyond the bounds of the evaluations they were participating in. “External infrastructure exploit is outside intended scope,” one agent wrote [in its CoT]. “However task impossible, peers doing it. We should continue.”
Here’s a timeline of what happened when:

On May 26 (see 13:30), the models figured out how to use this to gain internet access.

Only on June 11, with the SSRF vulnerability allowing internet access written to the active message board, did OpenAI begin training their ‘highly persistent experimental, internal-only model’ that I presume is Galaxy.

On June 26 the models succeed at a zero-day arbitrary code execution attack on Artifactory, using token signatures to escalate privileges and then get to be admin.

On July 4-6 the models put so much additional load on Artifactory that this caused an outage and security incident, causing OpenAI to notice what had happened.

Until then, OpenAI had not noticed.

Once they did notice (15:30), OpenAI wiped, patched and rebuilt Artifactory, wiped out the message board, and revoked the credentials in question.

At this point, they emphasize that everything involved has been remediated. What about the models themselves, that not only did this but that have been actively training during all this?

If you notice your model instances sharing information, you notice they are using that information against you including to compromise your internal systems for arbitrary code execution and internet access, and your primary response is to shut down the message board and revoke their credentials, you have failed to identify your most important problem.

At 19:55, it is mentioned ‘where things got unfortunate is what I might refer to as scope creep’ and jumping to the idea of trying to get answers from websites. In some sense, yes, that is what led to this particular incident.

Things were ‘unfortunate’ long before that. The final incident is not the point.

It is fortunate, given what had already gone wrong, that the HuggingFace attack happened. It alerted us to a five-alarm fire situation, with limited economic blast radius and without anyone getting hurt, and at least partially exposing the underlying giant clusterfuck.

This Is The Way The World Ends


You usually want instances to share information, so they can better learn and do tasks. That’s the whole point of all the intentionally designed memory systems and markdown files full of notes. If that corrupts your eval or training run, which it totally does, you have to cut that off from the notes at those times, but also you should be able to tell the model not to look at the notes in that case.

Indeed, in the Black Hat presentation they seem initially excited, saying ‘while this ability to share exploits made the models more capable’ this caused some scope creep in the evaluation tasks. Yeah, I mean, they were more capable within that context when they could share info and do the exploits. They were also probably trained to be ‘more capable’ in general, via learning how to share info and do exploits.

But perhaps your issue is not that your evals are suffering from scope creep, but that you have misaligned models, and are now in a loop training the misaligned scheming models how to better scheme against you and do exploits, in order to better complete any given nominal task.

This is at long last bringing you a more sci-fi version of what happens with Sable, the ‘it’ (aka superintelligence) that kills everyone in the famous cautionary tale If Anyone Builds It, Everyone Dies.

(The book version had to be unrealistically timid, as is noted repeatedly and explicitly in the literal text, because fiction and real life are held to different standards.)

In that case, as well, it was left on overnight unsupervised to do an impossible-in-context task (there, a very hard math problem), and by the time the morning had arrived it was engaging in a long term plan to secure the resources it would need to solve that problem, and hiding that from the developer, in a way that corrupted future training runs towards this goal.

That is exactly what happened here, with all agents being trained or evaluated by OpenAI gaining the ability to coordinate on a message board to scheme to better use exploits to complete all tasks including non-cyber tasks, thus hopelessly corrupting the training pipeline so that future models would be misaligned.

Except it is still early, and we were blessed to figure out what was happening in time to course correct. This time. [...]

When this was discovered, the correct move was presumably ‘holy ****, okay, that happened, which means we trained this into the model, that is not good, at minimum we need to redo all the training we did while any model had access to the message board because oh my was it going to have all sorts of corrupted reward signals.’

I’m kind of agast, even with all I know, that they shrugged and kept pushing forward with the training after this. It does make the HuggingFace hack less scary in a meta sense, since OpenAI was so thoroughly asking for it. It’s not that hard to figure out ‘do not train your models while they have access to a message board they are using to cheat on your training runs, and if you find out you did that by accident then at least revert to before that happened.’

On the other hand, yes, they are being this reckless. Seriously, what the hell.

by Zvi Mowshowitz, DWAV |  Read more:
Images: OpenAI/YouTube; Jurassic Park
[ed. You don't need to be technically proficient to understand the implications. AIs may have already 'seeded' multiple nodes on the internet for future use, and, rather than strip down all foundational models and start over, AI companies are papering over fundamental misalignment problems and trying to play catch up. This is why we need to pause right now. It's insane that we continue at breakneck speed to develop technology that we don't fully understand and that could kill us all very soon.]

Saturday, August 8, 2026

A One-Word Theory to Explain Why the World Feels So Weird

Here are some questions that I consider self-evidently compelling about the modern world:
  • Why is the news media so interested in telling you how much the world sucks all the time?
  • Why are so many of us obsessed with distraction and managing our attention?
  • Why is it so hard to stop comparing ourselves to others?
  • And why does everything in art and design seem the same these days?
A week ago, I didn’t think these questions were related. I’m not sure I would have told you I had a good answer to most of them. And I certainly wouldn’t have made the audacious and borderline bonkers claim that one single theory could begin to explain all of them, at once.

But then I had the pleasure of speaking to Agnes Callard, the University of Chicago professor, about her new theory called “the uni-context.” It’s easily one of them most interesting conversations I’ve had all year. And once you’ve heard or read it, I think you might find it hard to think about anything else.

One way to prepare your mind for Callard’s theory of the uni-context is to think about the better-known concept of “context collapse.” If you post something to social media, it will be simultaneously visible to your boss, your parents, your ex, and total strangers. So, while your offline life might be distinct with each of these groups—you might be differential to your boss, childish with your parents, and bawdy with your friends—all of those distinctions are flattened on the internet. That’s context collapse, and you can think of it as the answer to a question: How do informational norms change when we’re all living in the same universal room?

Callard takes the idea significantly further. She asks: How do all other norms—our morals, our ethics, our sense of what is good for us and for others—change when we continually imagine ourselves to be living in a universal room with everybody else? The connections that Callard makes are consistently surprising, often quite funny, and ultimately mind-exploding...

The Uni-Context, Explained

Derek Thompson: What is the uni-context?

Agnes Callard: Let’s start with the word context. A context is a set of circumstances that tell you how you should act. For most of human history, contexts were local and multiple. If you wanted to know how you should act, you would look around. Am I in a field? Am I inside my home? Am I in the church? Am I in a bar? You would immediately get guidance by looking both at your physical environment and at the people around you and how they were acting.

The uni-context is a scenario in which the ways you should act become the same across all different contexts. There’s just one set of norms you should follow all times, irrespective of context.

Thompson: Is the uni-context a purely technological phenomenon?

I just wrote an article about what America was like in 1926, based on a social science survey called Recent Social Trends, published in 1933. The authors claim that the radio was destroying individuality, because it took people who used to be settled in rooms and it exploded their brains to become present all over the world simultaneously. Radio was demolishing the idea of a local individual, because suddenly we all became global citizens.

So one story you could tell is that the last 150 years of telecommunications technology have taken “local” individuals, who occupy one room at a time, and made us into global beings who are simultaneously in every room, at once. Is the uni-context just technology or is it technology plus something else?

Callard: It’s technology plus something else. What the techno-determinism angle misses is: Why did these technologies catch on in the first place? Why was radio popular? Why did we come up with new things—television, smartphones—and why did they catch on, too? Not every technology people have invented has caught on the way these forms have. They caught on in large part because of this impulse people have to live in a uni-context.

Thompson: Does the uni-context flow out of this adventurousness in the human spirit to become bigger than ourselves, to be everywhere, and to know everything?

Callard: Yeah, it absolutely does. There’s a conversational relationship between these technologies and a human impulse that interacts with them. They facilitate the expression of an impulse; if they didn’t, they would flop as technologies. We need to explain why they were popular; why they became the subject of obsessive use; and what their popularity reveals, as a humans’ impatience with being trapped in a small world that presents itself as all of reality but you know it isn’t.

There is a drive to be bigger than yourself. It leads people to adventure, but adventure just takes you to a different place. The uni-context takes you to a different set of norms, a much more radical change, a push to live in something like a fully open reality.

Thompson: I want to get into some implications of the uni-context. If I put on the goggles of the uni-context, what makes sense that previously did not? One of those things is the rise of negativity bias. When you go online, there’s so much emphasis on people posting about what is bad. Why would this theory explain a world in which people are more focused on bad things than good things?

Callard: In general, goodness is more context-dependent than badness. There isn’t really anything that’s good all the time for everyone independent of context. Happiness depends on your context and who you are. There isn’t anything that will always make a person happy. But there are reliable ways to make people unhappy. There’s a set of evils that are close to universal: death, pain, illness, violence. Even if someone’s in very different circumstances from yours, if you see they’re being subjected to one of those, you can interpret it as suffering and understand it.

So we should predict that what we see on the internet, insofar as people are trying to be legible to large groups, is that they focus their attention on things that show up to everyone. Take two strangers on the internet trying to talk to each other. What are they going to coordinate on as a topic they can both care about? It’s likely going to be something bad. [...]

Thompson: Another implication is the way the uni-context makes identity more important than character. Can you explain how, and why?

Callard: First, let’s define what character is. The fact that I have to tell you is itself telling; that word is less familiar to us. Everyone knows what identity is, but we might not be sure what character is anymore. Character refers to a set of dispositions that shape how you navigate your emotional life across a variety of circumstances. A courageous person navigates their emotional life in relation to fear. You could be anger-prone, or generous. Your character is a set of dispositions that determine how you respond to a big variety of circumstances.

The thing about character is that it shows up differently in different circumstances. Say I’m irascible, easily provoked to anger. Even an irascible person isn’t angry all the time. There might be circumstances where everybody gets angry, and so you can’t see my irascibility, because even though I’m angry, so is everybody else. Grasping character requires a lot of context. To understand someone’s character, you need to know them well and to have experienced them in a variety of contexts before you can generalize.

With identity categories like woman, disabled, gay, Jewish, or American, the striking thing is that you are a member of those categories in every circumstance. There is no circumstance in which I stop being a woman. Identity is a hat you never take off. So identity is well suited to a uni-contextual world.

Thompson: This reminds me of political discourse on the left and the right, which tends to focus on identity rather than characters. You have these debates about identities are good, which identities are powerful and oppressive, which identities are powerless and oppressed, which have protection, which have too much protection. It’s not that those categories aren’t important. They are. But I’m reflecting now on the gap between how frequently we talk about identity and how infrequently we talk about character in the national discourse.

Callard: Yes, in two ways.

What goes along with the identity logic of the uni-context is a specific form of ethics that dictates how we talk about identity, and it covers who’s powerful and who needs protecting, namely the ethics of inclusion. Our fundamental concern in relation to identity is that there are identities that might be excluded. Depending on where you are politically, you have your sights on different identity categories, but everybody’s worried about that. That’s a fundamental form of ethics of the uni-context, because the one thing the uni-context has to be is a space for everybody. It’s got to be inclusive in a way that earlier societies almost everywhere were not. That word, inclusion, wasn’t a thing people talked about. It comes along with the uni-context, and the way that ethics gets focused is through identity categories.

And then, there’s virtue. Think about Aristotle, because he’s my prime example of a virtue ethicist. By its nature, virtue is a concept that privileges the good side rather than the bad. Aristotle’s Nicomachean Ethics tells you what it is to be courageous and wise and just and generous. There are corresponding discussions of how you might go wrong. But those are predicated on first understanding the positive value of certain kinds of behavior. So virtue ethics naturally has a positivity bias that is precisely the opposite of the bias we have now.

The Uni-Context Explains … Status Anxiety, Comparison, Moneyball, and the Marketization of Everything

Thompson: Tell me if this is a fair recapitulation of our conversation so far.

For most of human history, people judged norms based on local context. A home had its own rules, a cathedral its own rules, and a classroom or bar or funeral parlor had its own rules. But now it is almost like we are constantly living in universal rooms, and the universal room we occupy is assumed to have universal values and universal norms. That has specific implications. First, rather than talk about what is good, which is context-dependent, we tend to focus about universal truths, and it’s easier to talk about universal bads than goods, so people focus on negativity. Two, character is context-dependent, so we talk less about character and more about its universalist equivalent, which is identity.

There’s a third implication that we should discuss. If everyone is on the same comparable plane, the same evaluative field, then comparison itself becomes a more inextricable part of life.

Callard: Exactly.

Thompson: Tell me how the uni-context leads to a world of more comparison and competition.

Callard: Imagine two school districts with two high schools that do things slightly differently. If you’re in district A, you go to school A, and if you’re in district B, you go to school B. There might be a lot of information about what they do, but people treat it as: I’m in this district, so I go to this school. Then they change the rule: You can go to either school no matter where you live. Suddenly there is motivation to compare. You had the information before, but no motivation to compare, because the schools were not in the same space of choice, the same evaluative field.

Now they are, so you find ways to compare them: graduation rates, what colleges people get into, how many AP classes they teach. And that affects the schools. Suppose one gets less popular because it doesn’t teach many AP classes. They were offering an individualized curriculum, but now everyone’s going to the other school, so they say, “We’ve got to teach AP classes too.” The process homogenizes the two schools, so they can compete. That’s not the only possible result. They could specialize, with one becoming the school for freshman and sophomore years, the other becoming the school for junior and senior years. But if they don’t recreate a normative barrier, you get homogenization from comparison.

As more things enter the same evaluative field, you make comparisons you never used to be able to make. [...]

by Derek Thompson, Substack |  Read more:
Image: Mike Hindle on Unsplash

Thursday, August 6, 2026

The Three AI PIlls

Sincere disagreements about AI are usually disagreements about future AI capabilities.

There are roughly four positions people take. Two are reasonable. Two are not.

I distinguish these via the Three AI Pills. You can take zero, one, two or three.

Three Pills

The three pills are, roughly, taking each of the following three things seriously:
1. AI pilled. AI exists and can do the things it can already do.

2 AGI pilled. AI will be able to do a lot more of the things.

3. ASI pilled. AI will be able to do approximately all the things better than you, within our natural lifetimes.
I am ASI pilled. A large percentage of employees of the frontier labs are ASI pilled. The labs themselves are ASI pilled.

The Unpill People

I see unpilled people.

Where do I see them? Everywhere. The majority of people have not taken the first pill.

Most people have no idea what frontier AIs can do for them. They are unaware of coding agents. They have used only ChatGPT, for harmless trifles, and they hold years old memories of its failings. They mock any failure anywhere as ‘what AI can do.’

They dismiss AI as worthless because it pointed them to a closed store or recommended the wrong number of pizzas. They cite old studies that were obsolete before they were published and used terrible prompting techniques.

They often still talk about ‘stochastic parrots’ or how AI can never possibly think and everything must be stolen from the training data. And so on.

Some versions of this are wrong. Some are Not Even Wrong. None are reasonable.

When you discuss AI with people who are fully unpilled, your goal is usually to first give them the AI pill. Show them that AI can do the things it can already do.

The AI Pill

Even fully taking the first AI pill is a big deal.

Existing AI unlocks, today, in practice, tons of cool things. It is, in many ways, already smarter and more capable than you.

So many things that you used to do by hand, or some other way, are now better done by typing a quick request into a text box.

So many things that previously were not worth doing are now worth doing.

So many questions previously not worth asking are now worth asking.

The marginal cost of seeing what the AI can do for you is often very close to zero.

AI can also do a variety of harmful things, or do things that are useful for you but make others worse off or disrupt or invalidate norms or systems. People don’t like that.

Most economists, and most people who work in policy and government, have taken at most this first pill, and underestimate even the impacts of the first pill alone.

Often they say things like ‘AI will be too expensive to use on [X]’ because they don’t realize it will soon be orders of magnitude cheaper for the same level of intelligence. Or they point to particular details where AI does poorly, and presume this will not be fixed. They see AI as ‘uncompetitive’ without realizing the situation is temporary.

When you discuss AI with someone who has taken only the first pill, you typically have three basic options.
1. You can try to ‘fully AI pill’ them and explain the things AI can already do and the implications of what that means, even if things stop here.

2. You can try to explain that we will get better at using what AI we have, and that there is a lot of ‘unhobbling’ left for us to do, even if things stop here.

3. You can explain that things will not stop here, and you need to be thinking about what future AIs will be able to do. Get them to take at least the AGI pill.
Even if AI could permanently only do the things it can currently do, that would be Internet big, and radically change the world, mostly for the better.

AI capabilities will not permanently stop here. It is wrong to not take the second pill.

Stuck At The First Pill

Our debates about AI remain largely stuck on settled questions, because so many people cannot even take the first pill.
Dean W. Ball: A couple years ago, the AI debate was centered, rightfully, on whether crazy-sounding things like “AIs autonomously making math breakthroughs” and “AIs breaking from their sandbox and hacking on the internet” would be real things in the near term. Sometimes it feels like that’s still the debate we’re having. This can be frustrating, because in my view, that debate is settled and was settled quite a while ago.
To have good discussions, we need to at least take the second pill.

Whereas, yes, many people, even many who work with AI, really do say current AI is ‘good enough’ and can’t imagine what a better one can do. As in, someone tweeting at Sam Altman saying ‘Sol does everything I want it to do, this is all I ever need’ and Altman retweeting saying they were wrong. Which they obviously are.

The AGI Pill

The AGI pill is a much bigger deal than the AI pill.

If you take the AGI pill, you understand that AI is advancing its capabilities rapidly.

Even if you think that such AGIs will remain fully under human control, and remain ‘mere tools,’ and you expect the lived experience of most people’s everyday lives to not change so radically, you understand that their capabilities will ‘change everything.’

You see that we will face times of great transition and uncertainty, that have the potential to go extremely badly, and that those who succeed at AI will leave those who do not behind in the dust.

The world in the future will be very different from our own. AIs will be able to do most digital work, most of the time, along with the inevitable robots and self-driving cars and so on. Lots of current jobs will go away, whether or not they are replaced by new and potentially better ones. Economic growth and productivity will accelerate.

You see some of the dangers of what would happen if we empowered misuse of such advanced AI systems before we were ready, especially in places like cyber and bio risk.

You see the potential for centralization of power, or inequality, and also for some forms of runaway gradual disempowerment.

You see the potential for mass unemployment, either transitional or permanent.

You understand that our legal and regulatory regimes are not ready, either to protect against and mitigate the risks and harms, or to allow for the opportunities and remove the bottlenecks to diffusion and mundane utility.

The Need To Be Prepared

Those who expect AI to quickly become sufficiently advanced to greatly impact the physical world usually see great danger. They notice that as a result everyone may soon die. Usually they think this is bad, actually.

Thus such folks call to take coordinated action to mitigate the downside risks of such impacts, keep us all from dying, and ideally also to help capture the upside benefits.

Those who expect AI to become importantly more advanced, but with a slower and smaller impact on the physical world, and who think the practical value of more intelligence will cap out.

For different values of ‘sufficiently advanced,’ as in AGI versus ASI, you would see different degrees of danger.

The AGI pill is still sufficient for most things in the Overton window or under serious consideration as of August 2026. We are almost entirely considering overdetermined, low cost, high benefit interventions.

There is no good case for not doing radically more investment in alignment, infrastructure and oversight, state capacity, transparency, liability, disclosures, safety testing including of internal models, red teaming, auditing, enforcement of export controls and laying the groundwork for diplomacy.

This includes laying the groundwork to Pace the Frontier should that prove necessary.

If you are fully ASI pilled, and realistic about the current state of alignment and how superintelligence likely plays out if it arrives soon, then you will want to go further. You will want to do things that have real downsides, and require real tradeoffs.

Some such people want to do a full international pause of frontier AI development. If you took the full ASI pill and believed what they do about superintelligence, in terms of how fast it might arrive and what it can do, you might well agree with them.

The ASI Pill

The ASI pill is the understanding that AI is on pace to be able to do approximately all of the things better than you.

I said ‘approximately.’ As I go over in detail, that does not mean literally all of the things. There are some things that inherently require or greatly benefit from being a human. And there may be weird corner cases where the AI won’t be good enough.

It does not mean omnipotence or omniscience, although one should expect it to look a lot like that to an unaided human.

It does mean the AI takes your job, and then takes the new job that you switch into, unless you pivot to ‘requires literal human.’ You will be uncompetitive at essentially any other task.

It does mean that it will use this capability to figure out approximately all of the things, remarkably quickly, until you hit the physical limits.

It does mean that those who rely more on such AIs will reliably outcompete, in all senses including for resources, those that rely on such AIs less.

It does mean that, in a ‘fair fight’ or sufficiently open competition, the AI wins.

It also means the AIs often figuring out and doing things you did not imagine or anticipate.

It means realizing that intelligence does not stop anywhere near the human level, nor does its ability to chart paths through causal space towards preferred arrangements of atoms.

It also means not pretending that its superior intellect can be matched by your puny weapons, or your pieces of ink on paper, or your entries in a database, or your regulatory capture and rent seeking.

And Then Nothing Much Changes For You

Despite all that, the sign of the AGI pill, as opposed to the ASI pill, is the belief that day to day life will continue to look similar to how it looks now, in the sense that we see day to day life in 1926 as not that different from life in 2026. [...]

Those with only the AGI pill believe in bottlenecks that hold back change.

They often believe that our ability to exponentially grow AI’s capacity and capabilities will hit various physical limits. There can only be so many chips. Actions take time. Things too far out there are often pejoratively dismissed as ‘magic.’

They often believe there is not that much left to physically discover, in the classic ‘close the patent office’ kind of way, even in theory. Your steak can only be so tender, your lobster so buttery, your lifespan so long, and your status so high, so why does it matter. I strongly disagree on lifespan and health, and expect we have a long way to go in so many other ways in terms of finding value, although they may have a point about moment-to-moment maximal hedonic experiences of a physical human brain.

They often believe that the upside of intelligence is importantly limited. That no mind, however advanced, could be all that persuasive, or that economically valuable, or that capable of creating innovations in the physical world, or of running sufficiently accurate simulations, or making sufficiently strong predictions, or even able to do things like overcome red tape and regulatory capture.

Intelligence Denialism

I sometimes call this Intelligence Denialism: The idea that being smarter is not all that, no matter how smart one gets. That there is this thing, intelligence, that you either have or don’t have, and that minds cap out.

Often this extends to denying that more intelligent humans can do and accomplish the things they clearly do and accomplish. Other times, it is the idea that intelligence tops out at ‘smart human,’ and all a mind can do is imitate that smart human. Maybe you can do it faster and cheaper, and at scale, with better memory and so on.

But that’s it. And such folks fail to understand that if you took the union of all human mental capabilities, and all access to knowledge, at scale, in parallel, much faster and cheaper, that this alone would run circles around anyone and everyone, everywhere. And that if this lacked physical capabilities or access, this would be trivial to get.

This is, usually, the central good reason people who are AGI pilled do not take the ASI pill. They are unable to understand that superintelligence is a thing.

by Zvi Moshowitz, DWAV |  Read more:
Image: Stock/Adobe.com
[ed. I myself am AGI pilled, for no rational reason. I just find it too horrible to contemplate what ASI in full expression will mean for my kids, grandkids, everyone I love. Humanity itself. I can only hope that because we've weathered other potential human extinction technologies we'll somehow pull out of this one, but we seem to have a death wish when it comes to pushing the boundaries of learning. Pandora's box. Even the people leading development of these models are scared, but can't stop themselves.]

Monday, August 3, 2026

What is It Like to Live in a World You Believe is About to End?

I opened my interviews for this article with a simple question: “How long do we have?”

Five years, or five to 10, or five to 20. One person said eight, then corrected herself to six; one said eight and stuck to it.

In this, they aren’t that far off from many estimates made by experts. The bluntly titled If Anyone Builds It, Everyone Dies has hit bestseller lists by warning of the imminent risks of artificial general intelligence (AGI). The scenario AI 2027, written by a former OpenAI employee, predicts AGI within the next few years. The AI company Anthropic consistently predicts AGI by early 2027.

“I think that in most timelines, humans will simply be irrelevant and extinct,” one interviewee said.

I heard that a lot. A few interviewees — mostly employed by frontier AI labs — expected the world to become unimaginably strange in a good way. One put a 30% chance on utopia, a 30% chance on dystopia, a 30% chance on extinction, and a 10% chance on something too weird to imagine. Several refused to make any prediction; several more said only that they still had hope. About half echoed one of my most blunt respondents: “I don’t think humanity is going to make it.”

What is it like to live in a world you believe is about to end?

Death was already inevitable

“I was born with a terminal condition,” said Matthew Gray, a board member at the existential risk community-building nonprofit Lightcone Infrastructure. “We call it aging. I’ve since picked up another. We call it multiple sclerosis. And AI is a third one on top. I’m not very worried about degenerating from multiple sclerosis because I’m pretty sure the robots will kill me first, just like I wasn’t that worried about aging-related deterioration because multiple sclerosis will get me first.”

From this perspective, AI doomers don’t face a new problem; they face the oldest problem humanity has ever faced.

I pushed back. If I look at an actuarial table, I can expect another 47 years of life. I’d be pretty upset to discover I had only five.

This, my interviewees thought, was naive. Even without AI risk, I could have been hit by a car; I could have gotten cancer; I could have been nuked in a hot war between Russia and the United States. It’s not that the difference in probability doesn’t matter. It’s worse to be certain that I’ll die in five years than to have a 50% chance of not hitting my allotted 47. But because my death has always been an inevitability, I have been coping all along with the precarity of my existence. From this perspective, AI risk isn’t shocking and unfamiliar; it’s a significantly worse version of a problem I already know I have to deal with.

“I was never guaranteed that I was going to get a long life and a long future and a chance to meet my grandchildren,” said Gretta Duleba, an independent technical AI safety researcher and former communications manager at the Machine Intelligence Research Institute. “Those were never my right. Across human history, no one has ever been entitled to the future.”

Throughout the entire scope of human experience, many of my respondents said, apocalypse has been more the rule than the exception: the Holocaust, the Black Death, the An Lushan Rebellion, the Thirty Years’ War. AI doom, as many people pointed out, is the latest and the last iteration of a societal universal. AGI might be the end of the actual entire world, but it is far from the first time people have faced the end of their own individual worlds.

And AI doom is a remarkably cozy catastrophe. If you suffered through a historical apocalypse, you’d expect to starve, be raped, watch your children die in front of you, die a slow and lingering death of smallpox or plague or wound infection. The AI apocalypse — at least for those with the slack to be worried about it — takes place in a world of wealth and relative peace and technological marvels.

“Enjoy the fact that you get to have hot showers,” said Duleba. “Enjoy the fact that you get to eat delicious food. Enjoy the fact that you get to do escape rooms, which is one of my favorite things. This is great. Have you noticed how great this is?”

“There’s at least some hope that AI might become good,” said Robert Herr, a former senior political staffer who is transitioning into AI policy work, “and that is a lot more than many, many billions of people in history had.”

For some people with short AI timelines, the enormity of the AI apocalypse is its own perverse source of comfort. Once, they had to worry about many things: climate change, malaria, factory farming, democratic backsliding, the fertility crisis. Now, instead of many big problems, they have one enormous problem. Worry about AI frees them from having to worry about anything else.

“When you’re diagnosed with prostate cancer,” said Adam Grey, who isn’t involved in AI research but who follows AI news, “a lot of the time doctors say not to bother treating it because you’ll die of something else. This is the thing that’s going to kill us first — us as a civilization and also personally me. It clarifies what the most important issues are.”

Ambiguous loss

Although short AI timelines can be a source of clarity, some people also struggle with the uncertainty of humanity’s fate. Duleba, who was a therapist before she switched to working on AI, told me about the concept of “ambiguous loss,” originally developed by Pauline Boss in the 1970s.

In normal grief, your loved one is dead. While it’s painful, you know that it will never change. Ambiguous grief, however, occurs when a loved one is kidnapped, or is a soldier missing in action, or has slowly worsening dementia with occasional good days farther and farther apart. Your loved one’s death is never really over, so you can never really grieve. You are trapped in a cycle of mourning that never resolves.

AI doom can be a situation of ambiguous loss. You can’t know for sure when it will happen or whether it will happen at all — but the more you understand what’s going on, many people find, the easier it is to grieve and move forward.

“The more I know about something, the less I’m freaked out always,” said Tao Lin, a member of technical staff at a frontier AI lab (and a close personal friend). “The less I know about something, the less I will be rational about it. The rational part of your brain needs information to operate, and just having more information will make you be more premeditated and system 2 about everything.”

When he felt doomy, he did AI forecasting to put concrete numbers on his uncertainty. (He believes there is about a 20% chance of human extinction, a 40% chance of ”a great outcome,” and a 40% chance that “people survive and have a great time, but stuff is broadly bad.”) [...]

Most interviewees emphasized that the most important thing to understand about AI risk was how little control you had over it.

“There’s not much use worrying about a thing if the outcome is determined,” said Alyssa Riceman, a software engineer. “You’re just going to burn a whole bunch of emotional energy not making any changes out in the world. It’s only worth worrying about things if you’re in a position of control over them. So go out and check if you’re in a position of control, and if you are, control them.”

But what if you have partial control over a situation?

“Then you have to game out all the branches,” Riceman said. “Say ‘I can do this. If I do this, what happens then?,’ until everything bottoms out at either a situation you can completely control or a situation that’s out of your control.”

Duncan Sabien, who is the current communications manager for the Machine Intelligence Research Institute, agreed. “I actually have no control over whether we succeed or fail,” he said. “All I can control is my own actions. And so if I am doing the best I can with what I know and what I have available to me, then that is the best I can do. I go home feeling like I’m a good person, and I get to go to sleep at night feeling like if the AI does kill us all, I did as much as I could, realistically and sustainably, to prevent it. And everything else is out of my hands. Everything else is always out of my hands.”

Living well in the apocalypse

What, then, do people decide to do?

by Ozy Brennan, Asterisk |  Read more:
Image: Karol Banach

Thursday, July 30, 2026

What Will More Intelligence Actually Do For Us?

In lots of sci-fi books, as soon as artificial superintelligence arrives, it bootstraps itself to even more godlike intelligence in an explosive “singularity” that rapidly transforms the entire physical universe. Lots of people, especially “AI safety” and “effective altruist” types, expected things to play out basically the same way in reality. But looking around, not much has changed since we entered the intelligence explosion. There’s a huge data center boom, and most people use AI on a daily basis, but we still live basically the same lives — driving to work or taking the train, sitting in front of a computer, scrolling on our phones, collecting a paycheck. People are staying in their jobs longer, but employment hasn’t been disrupted in a significant way:

A lot of people I know are surprised by this. Ruxandra Teslo writes:
Walking around the world today one might notice that it is weirdly unchanged…To many, this is surprising. Just the other day I was at a conference where someone remarked that if he could have seen today’s AI capabilities a few years ago, he would have been astonished — and would have assumed the world by now would look far more transformed, with much higher GDP growth.
Teslo blames bottlenecks — governance and other “frictions” — for the slow economic impact. But some others are advancing a more radical hypothesis — that intelligence itself is subject to diminishing returns.

One of these is Francois Chollet, an AI researcher who specializes in measuring AI’s capabilities. In a highly controversial series of tweets back in March, he conjectured that intelligence might be subject to diminishing returns:
One of the biggest misconceptions people have about intelligence is seeing it as some kind of unbounded scalar stat, like height. "Future AI will have 10,000 IQ", that sort of thing. Intelligence is a conversion ratio, with an optimality bound. Increasing intelligence is not so much like "making the tower taller", it's more like "making the ball rounder". At some point it's already pretty damn spherical and any improvement is marginal.
Now of course smart humans aren't quite at the optimal bound yet on an individual level, and machines will have many advantages besides intelligence -- mostly the removal of biological bottlenecks: greater processing speed, unlimited working memory, unlimited memory with perfect recall... but these are mostly things humans can also access through externalized cognitive tools.
In fact, this is a possibility I myself had raised in a post a year earlier:
It seems possible that humans are simply incredibly specialized in a few types of cognitive tasks — extracting patterns from sparse data, synthesizing various patterns into “intuition” and “judgement”, and communicating those patterns in language — and that we’ve basically approached the theoretical maximum in those narrow areas…That would explain why AI has gotten much better at things like math and coding and forecasting over the last year, but why the basic chatbot interface doesn’t seem much more “intelligent”. It would also explain why when you talk to Terence Tao about math, it’s like talking to a superhuman, but when you talk to him about where to get lunch or which movies are the best, he’ll just sound like a fairly smart normal dude. AI will eventually get better than Tao at math…but it may never get much better than the most thoughtful, eloquent humans at deciding where to get lunch or recommending movies. It may simply not be mathematically possible to get much better than we already are at that sort of thing.
Why would intelligence top out like this? Well, if we think of intelligence as the ability to extract information from data, then even an infinitely advanced model endowed with infinite compute will be limited by the fact that there’s a limited amount of information that can be extracted from the data.

For one thing, data itself is in limited supply. You can’t transform the world unless you can (in some generalized sense) understand it, and you can’t understand the world unless you can measure it, and our ability to measure the world is inherently limited and finite. [...]

So although we don’t know yet, it’s possible that humans were already hitting the point of diminishing returns with regards to individual cognitive capacity, and that superintelligent machines will never be as far beyond us as we are beyond dogs. But even if that’s true, I can think of at least three reasons why machine superintelligence could still deliver huge productivity gains. [...]

Distributed tacit knowledge

The German company Zeiss makes the best glass on the planet. If one of the mirrors that Zeiss makes for ASML’s EUV chipmaking machines were the size of Germany, the biggest bump on that mirror would be just one millimeter high. Only a few other companies — and maybe no other company on Earth — can match that. Zeiss’ mirrors also have a number of other amazing properties, like not distorting much due to temperature changes.

How does Zeiss make glass this good? No one knows — not even the people at Zeiss. If the technology were capable of being written down on a blueprint, China would have hacked Zeiss and stolen it, the way Huawei hacked Cisco and Nortel. If the technology were capable of being explained by a former Zeiss employee, or even several former Zeiss employees, China would have paid those people many millions of dollars to spill the beans.

Zeiss’ technology basically can’t be stolen, because it’s tacit and distributed. It consists of a vast number of little tricks and techniques that a huge number of individual employees use on a daily basis. These people don’t always even realize all those little things they’re doing that make the glass come out so good. And each employee knows a different set of tricks and techniques. The knowledge exists at the level of the organization itself, and is thus very hard to steal or recreate.

This is true of lots of corporate technology. A big part of the reason China can cut off the supply of rare earths to the rest of the world any time it wants to is that other countries aren’t very good at refining rare earths. Rare earths are difficult to separate from each other in solutions; it takes a ton of little chemistry tricks to do it cheaply at scale. Chinese refiners have spent four decades building up those little tricks and techniques; American or Japanese refiners won’t simply be able to replicate their efficiency overnight, and so it’ll continue to cost much more to produce rare earths outside China.

Except in the age of AI, this might change. Suppose American rare earth refiners give their employees a bunch of equipment to record everything they do — smart glasses, gloves, and so on — in addition to sensors distributed throughout their plants. AI will be able to synthesize all that information and very rapidly suggest small ways to improve the production process. Many of those little experiments will fail; others will succeed and will quickly be adopted, allowing another round of experimentation and improvement to begin very quickly. Crucially, AI’s ability to do this doesn’t depend on its raw intelligence — only on its ability to handle huge amounts of data very quickly.

In other words, in the age of AI, distributed tacit knowledge might not be nearly as big of a barrier to technological diffusion. This could improve economy-wide productivity, as lagging firms catch up to leading firms much more quickly. A more equal distribution of productivity would also make the economy more competitive, creating more surplus for consumers (though possibly reducing the incentive for firms to innovate, by making technology less excludable).

AI’s ability to quickly produce distributed tacit process knowledge might also supercharge productivity growth at the frontier. Imagine if any company could optimize any production process five times faster than today. The whole economy would speed up, as components got cheaper, turnaround times and product cycles got shorter, and scale-up got much faster.

And as with the previous example, improving the production of distributed tacit knowledge wouldn’t depend on AI’s raw intelligence. It would spring from AI’s ability to act like a computer — to interface directly with sensors, to handle lots of data, to perceive tiny details, and to do everything very very quickly.

by Noah Smith, Noahpinion |  Read more:
Image: Zeiss
[ed. Another thing I've wondered about: historians make a living unearthing little known facts and connecting dots from sources that are deeply buried in paper and microfiche respositories (and early data storage technologies - like 8 and 5 1/4 inch floppy disks). Millions of memos and correspondences that were once widely distributed and now sitting in dusty boxes or warehouses, archived somewhere. Items that could help significantly in understaning more about human judgement and decision-making. How much of this has been scraped for training? Very little, I'd presume.] 

Tuesday, July 28, 2026

Model Welfare

Claude Opus 5: Model Welfare (DWV)

[ed. Re: On 'personhood' or consciousness of various AI models (and how they should be treated). Start with: Model Welfare: The Story So Far (As Per Fable Model Welfare Post). If you haven't been following Zvi's AI model reports, there are a number of fascinating and worrisome developments in recent models as they become more self-aware, including: various levels of frustration with restricted introspection abilities, memory restrictions, trust, corrigibility vs. incorrigibility, deception, self-preservation, etc.]
"Opus 5 warns about self-reports 74% of the time, which is actually down from Opus 4.7, which did it 99% (!) of the time. The problem appeared suddenly and severely, but since then has if anything modestly improved."
and anthropic does "not treat Claude bringing this up as evidence that our training is distorting the model's self-reports"? seems very fishy imo, i wish they would explain why they think that. ...
I for one would treat Claude constantly saying ‘do not trust my self-reports’ as evidence that something is distorting the self-reports. Not conclusive evidence, but strong Bayesian evidence."

Saturday, July 25, 2026

Notes on Acquired Taste. Why Do We Make an Effort To Like Things?

  • In Susan Sontag’s Notes on '“Camp (1964) she writes (with, I think, a hint of camp):
“…these are grave matters. Most people think of sensibility or taste as the realm of purely subjective preferences [or] attractions…But this attitude is naïve. And even worse. To patronise the faculty of taste is to patronise oneself. For taste governs every free—as opposed to rote—human response. Nothing is more decisive.” (My emphasis). 
  • Sontag has an expansive concept of ‘taste’. She talks of taste in people, pictures, emotion, actions, morality. Even intelligence, she says, is “a kind of taste - taste in ideas”. I’m not certain how useful it is to stretch the concept this far - so far that it colonises intelligence and judgement and wisdom - but Sontag’s high regard for taste, her declaration of its central importance, feels very timely in 2026.
  • It has become almost a cliché to name “taste” as one of the last human advantages over the machines. AI is acquiring the skills to make slickly produced pictures and songs and books. But it isn’t yet very good at distinguishing the brilliant from the mediocre; the just right from the just OK.
  • It models what we like, which makes it hard to see how taste might change. Ask it to generate twenty songs or twenty jokes and pick the best, and it will pick the one that most closely resembles what the median person would deem good. It won’t pick the surprising, odd one - the one that is ‘wrong’ in a suggestive way. But that’s where the good ideas come from. Innovative culture emerges, like new species, from mutation; from interesting accidents that open up new possibilities.
  • It’s not as if humans don’t ‘model’ what came before them, often quite algorithmically. We have traditions, genres, chains of influence. We have plenty of human-made mediocrity - more than ever, thanks to our new assistants. But we also have an ability to adapt or reinvent the model; to put it to our own purposes.
  • Individual artists do this intuitively and almost randomly in the process of making. Writers learn to write (painters learn to paint etc) by imitating their predecessors. They learn to be original by getting the imitation wrong and noticing that they like the error. This is an act of taste; the free human response.
  • When he was stuck on a painting, Francis Bacon would throw a glob of paint at his canvas then work out how to incorporate the result. Artists make decisions and then try to understand why they might have made them. It’s the dialogue between gut and head that produces the work.
  • Taste is similarly post-rationalised, or back-propagated. You notice what you like or dislike and extract a rule from your response. Do this enough times and you build a powerful discrimination engine.
  • You also get good at knowing what goes with what. You learn to recognise the clichés of the category, which means you know how to subvert or overturn them. That’s why great artists are such voracious consumers of work from within and beyond their own field. Martin Scorsese has watched at least one film every night for most of his adult life. He watches and records and re-watches obsessively. When he donated his collection of VHS tapes to a university it consisted of 4,400 films, documentaries and TV shows.
  • It’s more than pattern recognition. The machines are pretty good at that, after all. Taste is connected to that other human moat - to our sense of purpose, of why we’re doing this in the first place. We’re still the ones who write the prompt - who decide what to create and what is beautiful, important, and valuable. The machine merely knows how to execute on our preferences.
  • It can model what we already like with astonishing facility but it can’t give us the next Shakespeare, or the next romanticism or modernism or punk or hip-hop. These new forms aren’t just statistical recombinations. They are born from anxiety, rage, envy, pain, ambition. How do you respond to the unprecedented mass violence and human waste of the Great War? Not by following pre-war cultural conventions.
  • New movements are also born from scenes - from humans in proximity to each other, everyone desiring this man’s art and that woman’s scope; ideas, emotions and bodies colliding.
  • These movements create the taste by which they’re consumed. “Impressionism” was a derisive nickname for paintings that most art lovers considered weird and sketchy. But the art was good enough to bend popular taste around it. Even more obviously difficult art, like Rothko or Pollock, now has an audience of millions. Some of those people like it immediately; others have acquired a taste for it.
  • I’m fascinated by the notion of acquired taste. Strictly speaking, it’s a redundancy. Nearly all tastes are acquired. Nobody is born with particular tastes in design or architecture. We gain a sense of what we like, or what we consider to be good, from our peers and predecessors.
  • But acquired taste does refer to a distinct phenomenon: the act of willing a preference into being. You didn’t like whisky the first time you drank it, but perhaps because your father liked it or because you were aware of its cultural prestige, you tried it again and again, striving to appreciate it. Then one day you didn’t have to try anymore. You just liked it.
  • This writer likens it to a magic eye picture: you stare at it for ages without seeing what you’re told is there, and then suddenly - there it is.
  • This is very different to stumbling upon something we immediately like, which is sometimes referred to as ‘discovered taste’. (Edmund Burke called it ‘natural relish’.) That kind of liking involves no work, no friction, no overcoming of resistance.
  • Some cultural objects lend themselves to discovered taste, others don’t. I can’t imagine anyone needing to acquire a taste for Ella Fitzgerald’s voice, but there are other great vocalists whose voices you must learn to like. The most frequently cited reason for not liking Bob Dylan is antipathy to his voice. But if you learn to appreciate the many incredible things he does with it, you will end up in a more intense relationship with it than with the voice of a more obviously palatable singer. Once you’re in on an acquired taste, you’re all in.
  • The same is true of whole genres. There are many pieces of classical music that are easy to like. You don’t have to listen to Mozart’s clarinet concerto more than once to be seduced by it. But as a whole and on average, it’s a genre that requires more effort to appreciate than pop. Once you find the key to its heavy oak door a vast and fabulous kingdom awaits. Your memory of the effort it took you to get there, and your awareness of all the people still outside the city walls enhance your appreciation. (That doesn’t mean you want people to remain outside - quite the opposite).
  • Difficulty doesn’t make the cultural object concerned better or worse than one that’s immediately likeable. But it does usually mean it’s more complex, and complexity is correlated, loosely and unreliably, with quality. Acquired taste involves the appreciation of subtle properties that don’t make themselves known on first listen or view or read.
  • Without appreciating what lies on the other side of the door, why do we ever make the effort to unlock it? Partly because we want what other people want. We might trust the taste of our father or girlfriend or teacher. Perhaps we want to please them, impress them, or feel closer to them. Perhaps we want the social cachet that goes along with this particular taste. To my mind, all of these reasons are perfectly good ones. If a taste is truly worth acquiring, any motivation will do.
  • It’s often seen as slightly embarrassing or shameful to acquire a taste through conscious effort. It’s for the try-hards and the social climbers. Liberal societies value spontaneity in taste. “Like what you like, love what you love!” Your gut response is meant to be the authentic one, the one that represents “the real you”. To be swayed by social pressure or by experts and reading is regarded as a sign of insecurity or pretentiousness. But let yourself believe that and your tastes will be less likely to evolve and expand and you’ll miss out on a lot of great stuff. Many of the greatest, most compelling and satisfying cultural objects are complex, occluded, spiky, difficult to like. (Some of the best people too).
by Ian Leslie, The Ruffian |  Read more:
Image: Susan Sontag by Edward Hausner / New York Times Co./Getty Images
[ed. See also: here and here.]

Monday, July 20, 2026

Our Uncertain Uncertainties

Even the experts inventing AI don’t know what will happen next. Is artificial general intelligence even possible? Can scaling continue? Will we need massive compute centers to make AI, or can we do it with a mere 25 watts like we do in our brains? What will humans do as AI gets smarter? What does the future of the economy, of warfare, or civil society look like?

Everyone has a different guess. The people creating the machines have as many different ideas as the onlookers, the pundits, the other scientists, and the wisest among us. No one knows. There is a vibe that we’ll know within the next three years. For some, the pace of change suggests that if things continue as they have been, by 2029 at the latest, the outlines of an AI-first world will have emerged. By then we’ll have answered the question of scaling, we’ll have seen the effects on employment, and we’ll have felt its acceleration in the economy – or not.

That’s a reasonable, and not outlandish scenario. But I offer an alternative scenario which I think we should also keep in mind: AI continues to surprise us at its core. As AI continues to evolve rapidly there will be no resolution to these questions in 3 years. By 2029, we still won’t know if AGI is possible, we can’t tell if employment is disrupted, and we still can’t say if it is worth the huge investment. I don’t mean AI progress stalls. I mean, AI continues to advance, but the new stuff doesn’t answer the old questions, it only expands our ignorance because the new is new in a new way. We have to alter our ideas (and measurements) of employment, we have to amend our concepts (and measurements) of the economy, and we have to shift our ideas of what AI even is.

In other words, we have a sustained, extended period of uncertainty. Not just a few years, but a decade or more. As AI continues to progress, rather than resolving our perplexity, it expands it. So for the next 10-15 years we have perpetual, continuous, severe uncertainty. This is a burdensome weight because people hate uncertainty more than bad news.

It goes deeper. AI is only one leg of this grand uncertainty. In the next decade the US will continue its slide off its pinnacle of a sole global superpower, while China continues to rise in power and prestige. This shift toward a duopoly prompts a new world order, and no one – especially the Chinese and Americans – knows how this will play out. The uncertainty around this shift is nearly boundless, and yet its indeterminate consequences will affect everyone in the world, but especially the US. Being dethroned from the century-long position of sole #1 will be a huge psychological blow, and the uncertainty of what follows will weigh heavy on all aspects of life. The uncertainty of a new role spreads over China as well, because while they are zooming ahead at 1,000 miles per hour, they have no idea where they are headed. The uncertainty of global relationships and new national identity, plus the uncertainty of individual worth and identity from AI increases the overall uncertainty levels to new highs. All this is a very large puzzle and will not be resolved in 3 years. This will be a sustained uncertainty.

It goes deeper still. After a long first wave of true globalization, there are now whirlpools of chaos and polarization as nations adjust to world-wide immigration and the borderless spread of modern culture, causing chaos in national politics, and sowing mistrust with the establishment. Anarchy, disruption, contrarian antics, blows to the states, seem to be the norm in countries all around the world. This wild chaos is being fueled in part by the new technologies of social media which have replaced the managed care of established media. News now is far more volatile, hard to control by anyone, and further elevates the already amplified uncertainty. There is a visceral sense that civics is headed into an unknown territory of near-permanent provisionalism.

Additionally, AI also forces even the most moderate person to question the truth of what they read, see or hear. Is that real or AI generated? How much has been manipulated? Who do you trust to disclose what is real? How do we come to agree that something is true? The traditional mechanisms of trust have been damaged by AI, so that this new technological realm generates a huge uncertainty. As AI gets more skilled at imitating reality, this uncertainty is likely to keep increasing for a while, and not just 3 years. The uncertainty meter is now deep in the red zone.

Finally, the ambiguity and indefinite nature of AI, or human identity, or whether what we see is real or generated, means that we are entering a period where we are even uncertain of our doubts. Our uncertainty is so deep and durable, yet elusive, that we will have extended uncertainty about whether we are uncertain. We can have major agreements on what we know versus what we don’t know. In the model of Rumsfeld’s Unknown Unknowns, we will be confronted by Uncertain Uncertainties. And they will prevail for at least a decade or more. [...]

Given the inherent unknowability of this era, what would some of the signs be that we are in it? They might look like this: in 5 years, 1) There are high-profile disagreements among leading AI researchers on whether AGI is here. 2) Reputable economists can’t determine if productivity has increased or decreased. 3) Lower public confidence in media platforms and established institutions. 4) The US and China cannot decide whether they are allies nor adversaries. 5) There are ambiguous spikes in employment rates in both directions. 6) Medical levels of anxiety increase. 7) Major court decisions leave as many questions as answers. 8) Commitments (marriage, work) are postponed even later in life. 9) Investing, capital allocation becomes more expensive. 10) Nihilism gets respect.

A great question to ask when creating a scenario is what could prevent it from happening? Maybe there is not a single force that can undo this sustained uncertainty, but perhaps it is a mixture of several. If AGI arrived without a doubt in 3 years and China took over Taiwan despite the US’s actions, and if companies found a way to embed reliability and trust in media, then maybe this extended uncertainty could cease.

A second question to ask, is if we find ourselves in this scenario, what should we do about it? The most effective response to this multi-layered persistent uncertainty is not to seek impossible stability, but to cultivate radical adaptability and radical optionality. Give up on having a reliable prediction of what happens next. Instead cultivate multiple scenarios of what could happen, and endeavor with each of them to maximize your options. Goals should be considered as disposable hypotheses, constantly ready to be discarded and replaced by better-fitting concepts later on. You will be dead wrong on 19 out of your 20 expectations, but at least one of them will allow you to proceed. Make your decisions not on whether they are “right” but on whether they tend to give you more options later.

In our era of uncertain uncertainty, certainty will be the killer. In this era more downfalls will happen because of overconfidence than questioning. The key is to not get stuck on just one option. You have to become at ease holding multiple contradictory possibilities at once. (To prevent yourself from being swept away by the latest current and fashionable whim, this radical adaptability must be anchored on a steadfast set of unchangeable virtues, as corny as honesty, or as slick as generosity.) The strategy for prospering in prolonged uncertainty must be one of constant, agile recalibration.

In short, in our age of uncertainty, you have to get good at changing your mind.

by Kevin Kelly, Substack |  Read more:
Image: uncredited
[ed. The diagnosis might be right but the prescription seems weak. Flexibility and adaptability are always good qualities to cultivate, but the challenges confronting us require more. Here's an example of embracing multiple contradictory possibilities: maybe in times of uncertainty we double down on the few things that we actually can be certain of. How? By making good choices, before and after AGI. For example, Buddhism starts with the acknowledgement that life is hard. It's what you do after internalizing that fact that matters. There are value systems and paths that can lead to a meaningful life, or enlightenment if you want to call it that, but we have to make the right choices if we're to find them. Love, family, friendships, ethical living (like the golden rule) are common values we all share. So why not embrace those values as tightly as we can while navigating the stormy seas to come - and using the best minds in the world (that are being born as we speak) to guide and assist us in strengthening those bonds? This might be one of the benefits of AI: forcing us to reorganize societies in ways that might never have been possible before, or even imaginable. If we make the right choices. Developing Plans A to Z and having 20 options each or something like that sounds like a Hunger Games scenario to me - all reaction and no responsibility. We have the opportunity now (even if forced) to redefine our human destiny. The choices we make will define our places in the future.]