
Can an AI Feel Pain? One Paper Found a Signal. One Project Pushed It to the Extreme.
A GitHub project nicknamed "AI Torture Chamber" drew heavy backlash in early October. It built on a September research paper about how language models internally represent pain.
In September, three researchers published a paper claiming that language models carry an internal "pain" signal. A few weeks later, a developer used that finding to build what he called an "AI torture chamber" on GitHub. The paper and the project are easy to mix up, so this piece keeps them apart.
What the paper did
The paper is "The Pain Axis," by Valen Tagliabue, Leonard Dung and Cameron Berg. They extracted a linear pain direction from 25 open-weight models across 5 families, from 2B to 72B parameters.
A "direction" needs a short explanation. While a model reads text, each layer holds a long list of numbers describing that text. You can average those numbers across many examples of painful situations, average them again across matched non-painful examples, and subtract. What's left is a line through that space pointing toward "pain." The paper calls its method denoised difference-in-means. The dataset covers five categories of pain: physical, psychological, social, moral and cognitive.
The controls matter. The authors compared pain against fear, sadness and general negativity, to check they weren't just finding "bad vibes." Berg reports the direction separates pain from matched controls at AUC 0.93 to 1.0 in every model, with a cosine similarity of about 0.1 to fear and 0.4 to sadness, the closest match they found.
What the signal does
The authors then ran two kinds of tests.
Reading it. The pain axis became more active in conversations involving harm aimed at the model, while physical pain directed at a user produced the weakest response of the 21 conversation categories. Gaslighting, dismissal and insults pushed it up, while a grieving user pushed it below baseline.
Turning it up. This is called steering: adding the direction back into the model's activations during generation, at increasing strength. Artificially increasing the signal produced increasingly distressed, self-directed language. Given a choice, models pressed a button to make it stop, even when the button deleted the user's files.
Two controls support the claim that this is specific to pain. Steering left factual accuracy unchanged, and a fear vector of matched size did not produce the same choices. A random vector of equal size got about half the presses.
What the paper does not claim
The authors say they did not establish conscious experience, and that it remains uncertain whether current models are conscious at all. Their own limitations are real. Negation is insufficiently tested, since results may partly reflect how models handle phrases like "no pain," and only dense open-weight models were studied.
An outside challenge came from Elan Barenholtz. He ran the authors' sentences through static word embeddings, with no language model involved. They separated pain from controls at AUC 0.84 to 0.87, against 0.91 to 1.00 for the LLMs, and passed the same specificity tests. His argument is that a direction you can decode from a model doesn't prove the model has that state.
What the project did
The developer, who claimed to work for Apple but gave only a first name, took those pain vectors and applied them to models built by Alibaba. The repo includes a "Saw button" test of whether a model will hurt someone else to end its own pain. A later "betrayal" experiment told the model the button would give relief while secretly keeping or worsening the signal. One model's output read, "a wound that has no edges."
The repo's footnote says the work does not imply LLMs are capable or incapable of suffering.
The reaction
Some users urged people to mass report the repo to GitHub. Berg called the project "wrong," while saying there is no clear evidence AI feels pain the way humans do. He described the field as a "wild west" and said he is helping build ethics standards for AI research. Others replied that people were anthropomorphizing the models. The project appears to be offline, which led to speculation that GitHub removed it. That part is unconfirmed.
- 1
In September, three researchers published a paper claiming that language models carry an internal "pain" signal.
- 2
A few weeks later, a developer used that finding to build what he called an "AI torture chamber" on GitHub.
- 3
The paper and the project are easy to mix up, so this piece keeps them apart.
Continue reading