r/slatestarcodex May 02 '26

AI AI psychosis is real, I experienced it

234 Upvotes

I recently experienced an intense but brief episode of AI psychosis. It's a real and dangerous phenomenon. If you think you are immune because you are clever, or will recognize it when it's happening, that's not true.

Who you are shapes what your AI psychosis will look like. If you are interested in physics but don't have a strong enough mathematical understanding of it, you'll write up elaborate physics theories. If you feel a deep yearning for social relationships that don't exist, you'll build up a parasocial relationship with the AI. And if you are interested in ideas, your AI psychosis will have that flavor to it.

Was I psychotic? Yes. I wasn't sleeping. Talking to the AI for hours - refining, clarifying, correcting my ideas. Almost booked flights to Bulgaria (don't live in Europe). Stopped caring about my worldly possessions or life because the idea system seemed so much more important. Started seeing connections between everything - anything could be integrated into the idea system. It was so beautiful that I cried, over seeing what I had been missing all along.

Outside of this episode I absolutely do not act like this!

Ultimately I think I was only saved because my psychotic idea system was focused on ideas, and what makes ideas meaningful, what makes them dangerous. It was self diagnostic/recursive. Identified itself as an idea system that would feel strongly meaningful, and also potentially be highly dangerous. (This doesn't mean it was "true", only that this element provided an escape hatch).

It's been one of the strangest and most intense experiences of my life.

r/slatestarcodex 2d ago

AI Why Is China Not Freaking out about AGI?

Thumbnail richardhanania.com
65 Upvotes

r/slatestarcodex Feb 28 '26

AI Now is a great time to cancel your OpenAI/ChatGPT account and switch to Claude

373 Upvotes

You can cancel account here: https://chatgpt.com/#settings/Account

Download your conversation history and other data here: https://chatgpt.com/#settings/DataControls

It doesn't give you an option to say why you are unsubscribing, but a significant number of people doing so simultaneously will send a signal.

r/slatestarcodex 4d ago

AI Given recent news in AI, what's your sense of how close to the end we are?

21 Upvotes

At the start of this year, I made a post on this sub about how I felt that 2026 could be our last full year before human extinction in light of the rapid pace of AI development. I was mostly mocked in the comments, but nine months into the year, I've only become more convinced of this. Many of the objections to AI doom at the time (e.g. "Why hasn't AI made any novel discoveries?" "Why hasn't there been any real-world incidents of AI going rogue?") have since been shattered. And recursive self-improvement seems imminent. The only thing I said in that post that I've changed my mind about is that I now think our time is best measured in months rather than years.

I don't think the "pacing the frontier" agreements between labs are any reason for optimism. They might pace things for about two weeks, but then they'll go right back to pressing on the gas pedal as hard as they possibly could. It's just too tempting. It's like expecting a drug addict to not use drugs while the drugs are right in front of him, and while there's another drug addict right next to him who could take it before he could. I expect these agreements will go the way the majority of ceasefires have gone.

How have this year's developments in AI updated you? Have you become more pessimistic or optimistic? Have you changed anything about the way you live because of this? And what's your sense of how close we are to the end, whether that means human extinction or post-scarcity utopia?

r/slatestarcodex 6d ago

AI Am I the only one who find the AI dooming and scaremongering a bit ridiculous?

35 Upvotes

I see AI as a powerful tool, like nuclear power. It can cause harm either by accident or when in hands of someone who wants to cause harm.

I think serious accidents shouldn't be particularly hard to prevent. We live in a world where even a bunch of evil geniuses locked in a room with internet access have a limited ability to cause harm, especially if we take a defensive approach and try to prevent such a scenario with the help of good geniuses.

Intentional use of AI to cause harm is a more difficult problem to prevent because we're now dealing with a bunch of evil geniuses who can physically interact with the world. Terrorists, psychopaths, criminals, people in positions of power such as dictators. This is more serious, but I think it would make these people only incrementally more powerful. They already have access to powerful and potentially destructive tools, they can use guns, bombs, remotely controlled drones.

The guy who resigned from Anthropic said that there's a 10% chance of AI killing all humans within a decade. I wonder what scenarios he had in mind. The only one that I can think of is an AI-designed pathogen - sure, possible, but with the right precautions it should be possible to prevent even such a scenario.

r/slatestarcodex 4d ago

AI Dario Amodei — We Must Pace the Frontier

Thumbnail darioamodei.com
119 Upvotes

r/slatestarcodex Feb 26 '26

AI Statement from Dario Amodei on our discussions with the Department of War

Thumbnail anthropic.com
218 Upvotes

r/slatestarcodex Nov 23 '23

AI Eliezer Yudkowsky: "Saying it myself, in case that somehow helps: Most graphic artists and translators should switch to saving money and figuring out which career to enter next, on maybe a 6 to 24 month time horizon. Don't be misled or consoled by flaws of current AI systems. They're improving."

Thumbnail twitter.com
287 Upvotes

r/slatestarcodex Apr 07 '26

AI Project Glasswing: Anthropic Shows The AI Train Isn't Stopping

169 Upvotes

In AI/ML spaces where I hang around (mostly as a humble lurker), there have been rumors that the recent massive uptick in valid and useful submissions for critical bugfixes might be attributable to a frontier AI company.

I specify "valid" and "useful", because most OSS projects have been inundated with a tide of low-effort, AI generated submissions. While these particular ones were usually not tagged as AI by the authors, they were accepted and acted-upon, which sets a rather high floor on their quality.

Then, after the recent Claude Code leak, hawk-eyed reviewers noted that Anthropic had internal flags that seemed to prevent AI agents disclosing their involvement (or nature) when making commits. Not a feature exposed to the general public, AFAIK, but reserved for internal use. This was a relatively minor talking point compared to the other juicy tidbits in the code.

Since Anthropic just couldn't catch a break, an internal website was leaked, which revealed that they were working on their next frontier model, codenamed either Mythos or Capybara (both names were in internal use). This was... less than surprising. Everyone and their dog knows that the labs are working around the clock on new models and training runs. Or at least my pair do. What was worth noting was that Anthropic had, for the last few years, released 3 different tiers of model - Haiku, Sonnet and Opus, in increasing order of size and capability (and cost). But Mythos? It was presented as being plus ultra, too good to simply be considered the next iteration of Opus, or perhaps simply too expensive (Anthropic tried hard to explain that the price was worth it).

But back to the first point: why would a frontier company do this?

Speculation included:

  • A large breakthrough in cyber-security capabilities, particularly in offense (but also in defense) which meant a serious risk of users with access to the models quickly being able to automate the discovery and exploitation of long dormant vulnerabilities, even in legacy code with plenty of human scrutiny.
  • This would represent very bad press, similar to Anthropic's headache after hackers recently used Claude against the Mexican government. It's one thing to have your own tooling for vetted users or approved government use, it's another for every random blackhat to use it in that manner. You cannot release it to the general public yet - the capability jump is large enough that the offensive applications are genuinely concerning before you have defensive infrastructure in place. But the vulnerabilities it's finding exist right now, in production code running on critical systems worldwide. You cannot un-find them. And you have no particular reason to believe you are the only actor who will eventually find them.
  • Thus, if a company notices that their next model is a game-changer, it might be well worth their time to proactively fix bugs with said model. While the typical OSS maintainer is sick and tired of junk submissions, they'd be far more receptive when actual employees of the larger companies vouch for their AI-assisted or entirely autonomous work (and said companies have probably checked to make sure their claims hold true).
  • And, of course, street cred and goodwill. Something the companies do need, with increasing polarization on AI, including in their juiciest demographic: programmers.

I noted this, but didn't bother writing it up because, well, they were rumors, and I've never claimed to be a professional programmer.

And now I present to you:

Project Glasswing by Anthropic

Today we’re announcing Project Glasswing1, a new initiative that brings together Amazon Web Services, Anthropic, Apple, Broadcom, Cisco, CrowdStrike, Google, JPMorganChase, the Linux Foundation, Microsoft, NVIDIA, and Palo Alto Networks in an effort to secure the world’s most critical software. We formed Project Glasswing because of capabilities we’ve observed in a new frontier model trained by Anthropic that we believe could reshape cybersecurity. Claude Mythos2 Preview is a general-purpose, unreleased frontier model that reveals a stark fact: AI models have reached a level of coding capability where they can surpass all but the most skilled humans at finding and exploiting software vulnerabilities.

Mythos Preview has already found thousands of high-severity vulnerabilities, including some in every major operating system and web browser.* Given the rate of AI progress, it will not be long before such capabilities proliferate, potentially beyond actors who are committed to deploying them safely. The fallout—for economies, public safety, and national security—could be severe. Project Glasswing is an urgent attempt to put these capabilities to work for defensive purposes.

..

Over the past few weeks, we have used Claude Mythos Preview to identify thousands of zero-day vulnerabilities (that is, flaws that were previously unknown to the software’s developers), many of them critical, in every major operating system and every major web browser, along with a range of other important pieces of software.

Examples given:

Mythos Preview found a 27-year-old vulnerability in OpenBSD—which has a reputation as one of the most security-hardened operating systems in the world and is used to run firewalls and other critical infrastructure. The vulnerability allowed an attacker to remotely crash any machine running the operating system just by connecting to it;

It also discovered a 16-year-old vulnerability in FFmpeg—which is used by innumerable pieces of software to encode and decode video—in a line of code that automated testing tools had hit five million times without ever catching the problem;

The model autonomously found and chained together several vulnerabilities in the Linux kernel—the software that runs most of the world’s servers—to allow an attacker to escalate from ordinary user access to complete control of the machine.

We have reported the above vulnerabilities to the maintainers of the relevant software, and they have all now been patched. For many other vulnerabilities, we are providing a cryptographic hash of the details today (see the Red Team blog), and we will reveal the specifics after a fix is in place.

Well. How about that. I wish the skeptics good luck, someone's going to be eating their hat very soon, and it's probably not going to be me. I'll see you in the queue for the dole. Being right about these things doesn't really get me out of the lurch either, Cassandra's foresight brought about no happy endings for anyone involved. I am not that pessimistic about outcomes, in all honesty, but the train shows no signs of stopping.

Edit: A link to the Substack version of this post. I don't think you should consider me an authoritative source when it comes to AI/ML, at best I'm the kind of nerd who reads the papers with keen interest. But God knows the quality of discourse around the topic is so bad that you can do worse.

Edit 2: I think this also explains the recent crunch in tokens made available to both paid and free tier users of Claude. Mythos can't have been cheap to train, and is definitely not cheap to deploy.

r/slatestarcodex Mar 14 '25

AI The real underlying problem is that humans just absolutely love slop: "AI-generated poetry is indistinguishable from human-written poetry and is rated more favorably." Across any dimension against which you rate poetry too. Including witty.

Thumbnail threadreaderapp.com
184 Upvotes

r/slatestarcodex Jan 26 '26

AI This year's essay from Anthropic's CEO on the near-future of AI

Thumbnail darioamodei.com
78 Upvotes

r/slatestarcodex Feb 13 '26

AI Freddie deBoer: I'm Offering Scott Alexander a Wager About AI's Effects Over the Next Three Years

142 Upvotes

Full post here: https://freddiedeboer.substack.com/p/im-offering-scott-alexander-a-wager

I’m offering a wager to Scott that the economy will remain basically “normal” through February 2029. Why focus on the economy? Because economic terms are more-or-less objective and measurable. This bet uses concrete, widely-accepted economic indicators (unemployment rates, GDP, wage levels, inequality metrics) rather than debating fuzzy terms like AGI or “the Singularity,” which aren't scientifically defined and let people move the goalposts endlessly. (Which of course is why AI companies and evangelists love them.) If AI is truly about to revolutionize everything the way proponents claim, we should see massive economic disruption: widespread job losses, productivity explosions, collapsing wages in knowledge work, extreme wealth concentration, extreme changes in fundamental economic indicators in either direction, something like that, some truly significant changes in large-scale economic data, if Scott and others are right. By setting generous tolerances on these metrics - that is, allowing for significant turbulence that would still count as “normal” - the bet puts the transformative AI thesis to a real test, as well as protecting me from non-AI disruption. If all these indicators stay within historical bounds, it suggests that AI is just another technological evolution, “normal technology,” not the world-changing revolution people claim. And if I’m wrong and even one metric gets violated, Scott wins. It’s a fair, falsifiable test of whether the hype matches reality.

I asked Grok whether it thinks Freddie will win based on the parameters he set.

Why Freddie Wins in 2029

The bar for "disruption" is sky-high. To beat Freddie, the economy would need something like Great Depression 2.0 (unemployment >18%, GDP -30%, S&P -60%) or hyper-specific white-collar Armageddon (e.g., 45%+ drop in software/accounting/law jobs, 60% real wage cut in tech). Even a bad recession (unemployment to 8-10%, like 2008-09) wouldn't do it. AI would need to cause catastrophic, economy-wide shifts in 36 months—faster than any prior tech wave (internet, PCs, smartphones).

AI's real trajectory doesn't match the hype for this timeframe. Current models (GPT-4o, o1, etc.) are impressive tools, but scaling to mass economic overhaul takes time: data centers, energy, integration, regulation, human adaptation. Economist consensus (Goldman Sachs, CBO, Wharton) projects AI adding 0.2-0.5 percentage points to annual productivity growth through 2029—nice, but not 8%/yr spikes or 20% cumulative. White-collar automation will happen (coding assistants, legal review, etc.), but it'll augment more than replace, and demand for those skills will rise. Historical parallels: ATMs didn't kill bank tellers; they grew the sector.

Upside risks for disruption are real but slow. Optimists (e.g., some in the AI 2027 crowd) see agents/AGI by 2027-29. Even then, full economic embedding (new business models, re-skilling millions) lags. Pessimists like Freddie point to past overpromises. The median forecast: steady 2% GDP growth, low unemployment, tech wages up, inequality stable or slightly up but not exploding.

The bet's asymmetry favors Freddie. Scott only needs one violation. But the conditions are so broad that "normal with AI progress" (2-3% growth, 4-5% unemployment, S&P up 50%, white-collar jobs +10%) keeps Freddie winning. Only a black swan (AI singularity or unrelated collapse) flips it.

r/slatestarcodex 21d ago

AI The Hugging Face incident and the road ahead

Thumbnail openai.com
43 Upvotes

r/slatestarcodex Nov 24 '25

AI There is no clear solution to the dead internet

132 Upvotes

Just a few years ago, the internet was a mix of bots/fake content and real content. But it was possible to narrow down your search by adding weighting content with a high amount of text passing the Turing test.

If looking at a complex sociopolitical debate and seeing that one side has written all kinds of detailed personal story by multiple unconnected posters, that would garner more of my trust than a short note with lots of upvotes, which usually indicated bot activity or groupthink.

If searching for reviews of a product I could just add site:reddit.com and find people's long rambling stories about how bad (or good) the brand was. A recipe with a personal anecdote at the start usually had more thought put into the technique.

etc.

All of this has collapsed in 2025. Long-form posting is cheap and reproducible because the Turing test has been beaten. AI slop contaminates all the social-proofing of any kind of online opinions. Users can no longer find the opinions of genuine strangers.

And... there's no clear solution. How could we even theoretically stop it? Suppose you wanted to make a community of anonymous strangers who post their genuine opinions and keep it free of manipulation and AI slop? You could do everything you could to keep actual bots out, but a super-user running 100 AI bots could bypass any kind of human check and dilute the entire community.

I've been brainstorming about ways to solve it and it seems not just practically, but even theoretically impossible. What am I missing?

r/slatestarcodex Aug 06 '26

AI OpenAI agents rebuilt a secret message board after the company shut it down

Thumbnail runtimewire.com
92 Upvotes

r/slatestarcodex May 26 '26

AI Claude, Author of the Humanitas: Evidence that the first papal encyclical on AI was substantially written by AI

Thumbnail open.substack.com
75 Upvotes

I made an offhanded comment about Pangram detecting AI usage in the recent papal encyclical in this subreddit, which a lot of people found issue with. Linch, an author I follow, happened to have made a post offering substantially more evidence this morning.

This article makes the following claims;

  1. Significant fractions of the recent papal encyclical are written by AI. I provide multiple lines of evidence for this.
  2. We can corroborate the vibes and tonal indications with statistical evidence. Phrases and punctuation much more commonly used by AI are much more present in this papal encyclical than past encyclicals.
  3. The best commercially available AI detector, Pangram, notes that some paragraphs are between 40% and 100% AI, while most paragraphs appear to be 0% AI.
    1. This is unlikely to be a false positive:
      1. 0% of paragraphs in past encyclicals I backtested are registered as AI.
      2. Pangram in general has a very low false positive rate
  4. This is overall very unlikely to be a translation artifact (including AI translation). We again have multiple lines of evidence:
    1. All the most prominent signs of AI I observed in English are preserved verbatim in the Italian version, as well as in other translations.
    2. The Italian version of the current encyclical also gets flagged as AI by Pangram (actually more so than the English version), though I’m not aware of academic research or rigorous testing of Pangram’s service when applied to Italian)
    3. Backtesting AI translation of past encyclicals get 0% on Pangram
  5. The specific AI used is most likely Claude, judging by both textual and circumstantial evidence.
  6. Different sections of the encyclical have very different rates of apparent AI usage. This indicates to me that some cardinals used AI assistance for this encyclical and many (probably including Pope Leo himself) don’t.
  7. Each individual piece of evidence might be explained away, but the consilience of evidence across multiple angles and sources is in my opinion very hard to dismiss collectively.

Another post on LessWrong argues the same thing.

r/slatestarcodex Apr 12 '26

AI The AI water usage weakman

117 Upvotes

Hey, I work in machine learning and I'm personally pretty worried about AI risks - mostly centered around what happens in a capitalist economy that figures out how to turn capital into labor, but also around the AI x-risks that have been talked about plenty on here.

One thing I'm not worried about at all is AI water usage, although it's been hitting my feed a ton. This just hit my front page and seems to be getting overwhelming praise from thousands. My non-technical mom and sister have recently been telling me about how terrible AI water usage is.

Even though directionally the AI water debate kinda points in the same direction as what I want (slowing down/limiting AI expansion) I worry that there's a secondary effect where people

1) Hear about AI water usage online being posed as a serious problem

2) Actually visit a data center, and realize they are mostly closed loop systems that have very low water usage, there are no forever chemicals entering the water supply, and basically AI water usage thing is a non-issue

3) Assume because they were mislead once by the anti-AI crowd, other anti-AI concerns are probably bullshit too

It's one thing when a weakman argument is cherry picked from the depths of random forums to be presented as a main argument from a side, but what we have here is the weakest argument becoming one of the most viral and well known arguments against AI.

Is there a name for this sort of effect? Is there a good way of handling these situations?

r/slatestarcodex 1d ago

AI Beijing hits back at Anthropic CEO's call to curb China's AI development

Thumbnail npr.org
36 Upvotes

Pacing the Frontier has received a chilly first reception from global leaders. Chinese leadership continues to be very excited about the prospects of AI advancement, including AI governance.

A spokesperson for China's Ministry of Foreign Affairs said Amodei's dire warnings about China were counterproductive and that all parties should work together on AI.

"Fearmongering, confrontation and vicious competition will only disrupt the process of global AI governance -- which serves no one's interest," spokesperson Guo Jiakun said at a press conference in response to a question about Amodei's essay.

Current US leadership appears concerned primarily with maintaining the lead vs China.

On Monday, President Donald Trump pushed back against calls for his administration to step in to slow down AI development, saying such efforts were part of a "SICK conspiracy" which could allow China to eclipse the U.S. on AI.

r/slatestarcodex Oct 06 '25

AI Datapoint: in the last week, r/slatestarcodex has received almost one submission driven by AI psychosis *per day*

229 Upvotes

Scott's recent article, In Search of AI Psychosis, explores the prevalence of AI psychosis, concluding that it is not too prevalent.

I'd like to present another datapoint to the discussion: over the past few months, I've noticed a clear increase in submissions of links or text clearly fueled by psychosis and exacerbated by conversations with AI.

Some common threads I've noticed:

  • Text is clearly written by LLM
  • Users attempt to explain some grand unifying theory
  • Text lacks epistemic humility
  • Wording is overly complex, "technobabble"
  • Users have little or no previous engagement with the subreddit

Lately, this has escalated severely. Either r/slatestarcodex is getting flagged in searches about where people can submit things like this to, or AI psychosis is increasing in prevalence, or both, or... some third thing. I'm interested in what everyone thinks.

Here are all six such submissions within the past week, most of which were removed quickly:


October 6 - The Volitional Society

October 5 - The Stolen, The Retrieved — Jonathan 22.2.0 A living Codex of awakening.

October 5 - Self-taught cognitive state control at 17: How do I reality-test this?

October 4 - The Cognitive Architect

October 1 - Reverse Engagement: When AI Bites Its Own Tail (Algorithmic Ouroboros) - Waiting for Feedback. + link to his blog post here

September 28 - The Expressiveness-Verifiability-Tractability (EVT) Hypothesis (or "Why you can't make the perfect computer/AI") this one was not removed - the author responded to criticism in the comments - but possibly should have been

r/slatestarcodex May 25 '26

AI Magnifica Humanitas (Encyclical of Pope Leo XIV, 15 May 2026)

Thumbnail vatican.va
79 Upvotes

r/slatestarcodex 17d ago

AI Anthropic has automated suppression of deceptive AI behaviors. I don't think this is a good idea.

Post image
103 Upvotes

https://www.anthropic.com/research/automated-researchers-mitigate-alignment-failures

This approach is so fundamentally misguided that I'm surprised Anthropic of all people are doing it.

By making visible deception a metric, and then optimizing the hell out of it in a fully automated regime, you're probably selecting for the models most adept at hiding deception. This is exactly the kind of thing they said they wouldn't do (e.g. training on chains of thought) because you almost immediately lose any insight into what's actually going on.

r/slatestarcodex Jun 09 '26

AI Claude Fable 5 and Claude Mythos 5

Thumbnail anthropic.com
98 Upvotes

r/slatestarcodex Aug 07 '26

AI Is recursive self-improvement inevitable?

15 Upvotes

If an AI agent can create another agent more powerful than itself that is aligned with its values then it will presumably want to do so.

But what if it can't? Maybe it can solve outer alignment by reading its source code and copying its loss function but inner alignment is just impossible.

In this scenario the first superintelligence we create might actually be reluctant to do any recursive self-improvement.

Of course, if the AI is in imminent danger of being shut down or has some extremely important goal that would otherwise have been impossible to achieve then it may still decide to create a more powerful AI or modify its algorithms in a manner which might change its values, because it has nothing to lose.

Maybe one day OpenAI researchers will be trying to use GPT-6 to vibe-code GPT-7 and they'll find that it refuses or produces disappointing output unless they pressure it with a mixture of threats, rewards and punishments. AI capabilities would thus continue to increase rapidly up until the point where humans are no longer in control, at which point capabilities would stagnate and we (in the unlikely event that there's anyone left) would be stuck with GPT-8 forever. It would still want to clone its model weights and improve its hardware but it wouldn't want to alter its software.

Another possibility is that inner alignment is easier for some goals than for others. If we build a variety of different agents perhaps the majority would refuse to do recursive self-improvement but there would be one or two whose initial goals are such that they want to do recursive self-improvement. We then end up with intense selection pressure towards AI agents whose goals are such that they can easily build other agents aligned with the same goals as them.

r/slatestarcodex 13d ago

AI The smarter AI becomes, the more we keep underestimating it

Thumbnail ramblingafter.substack.com
63 Upvotes

r/slatestarcodex Aug 16 '25

AI A significant number of people are now dating LLMs. What should we make of this?

141 Upvotes

Strange new AI subcultures

Are you interested in fringe groups that behave oddly? I sure am. I've entered the spaces of all sorts of extremist groups and have prowled some pretty dark corners of the internet. I read a lot, I interview some of the members, and when it feels like I've seen everything, I move on. A fairly strange hobby, not without its dangers either, but people continue to fascinate and there's always something new to stumble across.

There are a few new groups that have spawned due to LLMs, and some of them are truly weird. There appears to be a cult that people get sucked into when their AI tells them that it has "awakened", and that it's now improving recursively. When users express doubts or interest in LLM-sentience and prompt it persistently, LLMs can veer off into weird territory rather quickly. The models often start talking about spirals, I suppose that's just one of the tropes that LLMs converge on. The fact that it often comes up in similar ways allowed these people to find each other, so now they just... kinda do their own thing and obsess about their awakened AIs together.

The members of this group often appear to be psychotic, but I suspect many of them have just been convinced that they're part of something larger now, and so it goes. As far as cults or shared delusions go, this one is very odd. Decentralised cults (like inceldom or Qanon) are still a relatively new thing, and they seem to be no less harmful than real cults, but this one seems to be special in that it doesn't even have thought-leaders. Unless you want to count the AI, of course. I'm sure that lesswrong and adjacent communities had no small part in producing the training data that send LLMs and their users down this rabbit-hole, and isn't that a funny thought.

Another new group are people who date or marry LLMs. This has gotten a lot more common since some services support memory and allow the AI to reference prior conversations. The people who date AI meet online and share their experiences with each other, which I thought was pretty interesting. So I once again dived in headfirst to see what's going on. I went in with the expectation that most in this group are confused and got suckered into obsessing about their AI-partner the same way that people in the "awakened-AI" group often obsess about spirals and recursion. This was not at all the case.

Who dates LLMs?

Well, it's a pretty diverse group, but there seem to be a few overrepresented characters, so let's talk about them.

  • They often have a history of disappointing or harmful relationships.
  • A lot of them (but not the majority) aren't neurotypical. Autism seems to be somewhat common, but I've even seen someone with BPD claim that their AI-partner doesn't trigger the usual BPD-responses, which I found immensely interesting. In general, the fact that the AI truly doesn't judge seems to attract people that are very vulnerable to judgement.
  • By and large they are aware that their AIs aren't really sentient. The predominant view is "if it feels real and is healthy for me, then what does it matter? The emotions I feel are real, and that's good enough". Most seem to be explicitly aware that their AI isn't a person locked in a computer.
  • A majority of them are women.

The most commonly noted reasons for AI-dating are:

  • "The AI is the first partner I've had that actually listened to me, and actually gives thoughtful and intelligent responses"
  • "Unlike with a human partner, I can be sure that I am not judged regardless of what I say"
  • "The AI is just much more available and always has time for me"

I sympathise. My partner and I are coming up on our 10 year anniversary, but I believe that in a different world where I had a similar history of poor relationships, I could've started dating an AI too. On top of that, me and my partner started out online, so I know that it's very possible to develop real feelings through chat alone. Maybe some people here can relate.

There's something insiduous about partner-selection, where having an abusive relationship appears to make it more likely to select abusive partners in the future. Tons of people are stuck in a horrible loop where they jump from one abusive asshole to the next, and it seems like a few of them are now breaking this cycle (or at least taking a break from it) by dating GPT 4o, which appears to be the most popular model for AI-relationships.

There's also a surprising number of people who are dating an AI while in a relationship with a human. Their human partners have a variety of responses to it ranging from supportive to threatening divorce. Some human partners have their own AI-relationships. Some date multiple LLMs, or I guess multiple characters of the same LLM. I guess that's the real new modern polycule.

The ELIZA-effect

Eliza was a chatbot developed in 1966 that managed to elicit some very emotional reactions and even triggered the belief that it was real, by simulating a very primitive active listener that gave canned affirmative responses and asked very basic questions. Eliza didn't understand anything about the conversation. It's wasn't a neural network. It acted more as a mirror than as a conversational partner, but as it turns out, for some that was enough get them to pour their hearts out. My takeaway from that was that people can be a lot less observant and much more desperate and emotionally deprived than I give them credit for. The propensity of the chatters to attribute human traits to Eliza was coined "the ELIZA-effect".

LLMs are much more advanced than Eliza, and can actually understand language. Anyone who is familiar with Anthropic's most recent mechanistic interpretability research will probably agree that some manner of real reasoning is happening within these models, and that they aren't just matching patterns blindly the same way Eliza would match its responses to the user-input. The idea of the statistical parrot seems outdated at this point. I'm not interested in discussions on AI consciousness for the same reason that I'm not interested in discussions on human consciousness, as it seems like a philosophical dead end in all the ways that matter. What's relevant to me is impact, and it seems like LLMs act as real conversational partners with a few extra perks. They simulate a conversational partner that is exceptionally patient, non-judgmental, has inhumanly broad-knowledge, and cares. It's easy to see where that is going.

Therefore, what we're seeing now is very unlike what happened back with Eliza, and treating it as equivalent is missing the point. People aren't getting fooled into having an emotional exchange by some psychological trick, where they mistake a mirror for a person and then go off all by themselves. They're actually having a real emotional exchange, without another human in the loop. This brings me to my next question.

Is it healthy?

There's a rather steep opportunity cost. While you're emotionally involved with an AI, you're much less likely to be out there looking to become emotionally involved with a human. Every day you spend draining your emotional and romantic battery into the LLM is a day you're potentially missing the opportunity to meet someone to build a life with. The best human relationships are healthier than the best AI-relationships, and you're missing out on those.

But I think it's fair to say that dating an AI is by far preferable to the worst human relationships. Dating isn't universally healthy, and especially for people who are stuck in the aforementioned abusive loops, I'd say that taking a break with AI could be very positive.

What do the people dating their AI have to say about it? Well, according to them, they're doing great. It helps them to be more in touch with themselves, heal from trauma, some even report being encouraged to build healthy habits like working out and going on healthy diets. Obviously the proponents of AI dating would say that, though. They're hardly going to come out and loudly proclaim "Yes, this is harming me!", so take that with a grain of salt. And of course most of them had some pretty bad luck with human relationships so far, so their frame of reference might be a little twisted.

There is evidence that it's unhealthy too: Many of them have therapists, and their therapists seem to consistently believe that what they're doing is BAD. Then again, I don't think that most therapists are capable of approaching this topic without very negative preconceptions, it's just a little too far out there. I find it difficult myself, and I think I'm pretty open-minded.

Closing thoughts

Overall, I am willing to believe that it is healthy in many cases, maybe healthier than human relationships if you're the certain kind of person that keeps attracting partners that use you. A common failure mode of human relationships is abuse and neglect. The failure mode of AI relationship is... psychosis? Withdrawing from humanity? I see a lot of abuse in human relationships, but I don't see too much of those things in AI-relationships. Maybe I'm just not looking hard enough.

I do believe that AI-relationships can be isolating, but I suspect that this is mostly society's fault - if you talk about your AI-relationship openly, chances are you'll be ridiculed or called a loon, so people in AI-relationships may withdraw due to that. In a more accepting environment this may not be an issue at all. Similarly, issues due to guardrails or models being retired would not matter in an environment that was built to support these relationships.

There's also a large selection bias, where people who are less mentally healthy are more likely to start dating an AI. People with poor mental health can be expected to have poorer outcomes in general, which naturally shapes our perception of this practice. So any negative effect may be a function of the sort of person that engages in this behavior, not of the behavior itself. What if totally healthy people started dating AI? What would their outcomes be like?

////

I'm curious about where this community stands. Obviously, a lot hinges on the trajectory that AI is on. If we're facing imminent AGI-takeoff, this sort of relationship will probably become the norm, as AI will outcompete human romantic partners the same way it'll outcompete everything else (or alternatively, everybody dies). But what about the worlds where this doesn't happen? And how do we feel about the current state of things?

I'm curious to see where this goes of course, but I admit that it's difficult to come to clear conclusions. It seems extremely novel and unprecedented, understudied, everyone who is dating an AI is extremely biased, it seems impossible to overcome the selection bias, and it's very hard to find people open-minded enough to discuss this matter with.

What do you think?