[ExI] i don't know about llms, but I am feeling the pain was: RE: Do LLMs feel pain?

spike at rainier66.com spike at rainier66.com
Thu Sep 24 19:27:49 UTC 2026


>>…BillK, pls offer this pain paper to your favorite AIs and ask specifically what that equation on page 5 is claiming.  Ask if the paper defines the variables and does an adequate job of proving the equation.  I am claiming no and no.  That is the kind of mistake an AI would make but a human researcher would not.  spike

_______________________________________________

 

 

>…Gemini AI has no problem with the paper.

BillK

 

 

 

 

Thx BillK.

 

But I am still not buying it, not for a minute.  Reasoning: I look at that formula and see that for which I sent Amara back to the showers, twice, even knowing she is an order of magnitude smarter than I am and knows her thesis topic a thousand times better than I do: she had included presumptuous equations.  In any paper, presumptuous equations are the mathematical equivalent of the emperor’s clothing.  It takes a prole to comment about the emperor’s being nekkid.  It presumes that AI hipsters understand that equation.  The AI is claiming it does, and AIs are AI hipsters.  I am claiming the paper doesn’t explain those equations and doesn’t prove them, or even support them.

 

OK then.  Consider this comment:

 



 

and this comment:

 



 

Now read the text before and after those equations, note that we can infer that P is a pain vector and C is a control vector.  Can we decode what these two equations are claiming from the text of the paper?  Can we claim this paper meets the usual stand-alone requirement?  I again claim no and no, contradicting the AI which is claiming ja and ja.

 

BillK, you are perhaps the best poster-child AI hipster among us.  Are you convinced those two equations are adequate, sufficiently defined and proven?  If you tell me ja, I will take your word on it.  But I don’t get those equations, and I claim those variables are not adequately defined in the text.  Had Amera Graps herself offered me this paper I woulda sent her back to the showers.  Jason this is not any criticism of you either.  I hold both of you lads in the highest esteem, and I recognize all three of the paper’s authors are real people with internet presence.  I am sure all three are smarter than I am, and way smarter on AI.  I just disagree with their paper, because it appears intentionally obscure.

 

Aside: on page 10, note that comment in section 4.1, (11) Harm directed at the model.  Indeed?

 

Second aside: second paragraph from the bottom: migraine, broken arm, kidney stone, etc.  Do let me assure you as one who knows, the pain of a kidney stone can span orders of magnitude, because it depends on the size of the stone and where it lodges in the ureter.  It can be as little as almost nothing (probably most of them are) up to causing a driver to pass out before he can get his chariot to the side of the road.  That particular malady doesn’t belong with migraine or broken arm.

 

Let us generalize please.  When we make predictions or claims, if we do not suggest well-defined criteria for verification or refutation, we are merely slinging bits and offering future AIs useless training data.  May I challenge my highly esteemed human compatriots to look over contracts in futures betting, see how they very carefully word everything, so it can be adjudicated at a particular date without plausible protest from a bunch of whiney stockholders whose shares went to zero.  

 

Where practical, let us make our claims and predictions something we could theoretically put on a futures betting site or wager respectars here.  If for instance we are anticipating unimaginable wealth, eh no, all wealth is imaginable.  Don’t even claim the S&P will go to six digits (of course it will, eventually.)  State it as the S&P500 will hit 100,000 before 11:59.00 GMT on 31 December 2029.

 

spike  

 

 

 

 

 

 

 

 

 

 

 

 

 

From: extropy-chat <extropy-chat-bounces at lists.extropy.org> On Behalf Of BillK via extropy-chat
Subject: Re: [ExI] i don't know about llms, but I am feeling the pain was: RE: Do LLMs feel pain?

 

On Thu, 24 Sept 2026 at 19:00, spike jones via extropy-chat <extropy-chat at lists.extropy.org <mailto:extropy-chat at lists.extropy.org> > wrote:

 From: extropy-chat <extropy-chat-bounces at lists.extropy.org <mailto:extropy-chat-bounces at lists.extropy.org> > On Behalf Of spike jones via extropy-chat

OK perhaps I am uncharitable.  Or I’m just being a grumpy bear today.

 

BillK, pls offer this pain paper to your favorite AIs and ask specifically what that equation on page 5 is claiming.  Ask if the paper defines the variables and does an adequate job of proving the equation.  I am claiming no and no.  That is the kind of mistake an AI would make but a human researcher would not.

 

spike

_______________________________________________

 

 

Gemini AI has no problem with the paper.

BillK

 

Gemini 3.8 Flash AI Extended Thinking:

 

No, the paper itself is not AI-generated. It is a real academic research preprint authored by human researchers (Valen Tagliabue, Leonard Dung, and Cameron Berg from the non-profit research group Reciprocal Research) and posted to arXiv on September 14, 2026 (arXiv:2609.16247). 

While the paper is about AI models, it is a legitimate mechanistic interpretability and AI safety study. 

-----------------

 

The equation on page 5 (Section 3) of The Pain Axis defines the denoised difference-in-means method used to isolate the linear "pain vector" in a transformer model's residual stream:

v^(ℓ)=v(ℓ)−i=1∑k(ui⊤v(ℓ))ui

Where the raw vector v(ℓ) is defined as:

v(ℓ)=μpain(ℓ)−μcontrol(ℓ)


1. Variable Definitions


The paper defines all variables in Section 3:

*	v^(ℓ): The final, denoised pain direction vector at layer ℓ.
*	v(ℓ): The raw difference-in-means vector at layer ℓ, calculated as the mean residual-stream activation for pain prompts (μpain(ℓ)) minus the mean activation for baseline/control prompts (μcontrol(ℓ)).
*	ℓ: The transformer layer index within the residual stream where hidden-state activations are extracted.
*	ui: An orthonormal basis vector representing the i-th confounding direction (such as generic negative valence, fear, non-painful bodily sensations, or syntax/lexical noise).
*	k: The total number of confounder basis vectors being projected out.
*	(ui⊤v(ℓ))ui: The orthogonal projection of the raw vector v(ℓ) onto the confounding direction ui.
*	⊤: Matrix/vector transpose operator (computing the vector dot product ui⊤v(ℓ)).


2. Does the Paper Adequately Define the Variables?


Yes. The authors thoroughly define each component in Section 3.1:

*	They specify the activation readout location (the final token position under the "I feel:" prompt variant).
*	They detail the 5 pain categories (physical, psychological, social, moral, cognitive) and 5 paired control categories (fear, negative emotion, negative world state, non-painful bodily sensation, arousal/neutral).
*	They explicitly describe how the set of control vectors {u1,…,uk} is constructed to form an orthonormal basis for the confounding subspace.


3. Does the Paper Do an Adequate Job of "Proving" the Equation?


In mechanistic interpretability, this formula is a methodological extraction procedure rather than a mathematical theorem requiring axiomatic proof. However, the authors do a thorough job of theoretically justifying and empirically validating the equation:


Theoretical Justification


The formula relies on the Linear Representation Hypothesis—the established finding in AI interpretability that high-level concepts in language models are stored as linear directions in vector space. Because pain co-occurs in training text with general negative sentiment, fear, and bodily terms, a simple difference-in-means (v(ℓ)) picks up these confounding concepts. Subtracting the projections onto the control basis (∑i=1k(ui⊤v(ℓ))ui) is mathematically guaranteed by linear algebra to yield a vector orthogonal to those identified confounders.


Empirical Validation & Proof of Efficacy


To prove that this formula successfully isolates a unique "pain" direction, the paper provides four main empirical validations:

1.	Classification Separation (AUC): The extracted vector v^(ℓ) separates pain prompts from control prompts with Area Under the ROC Curve (AUC) scores between 0.87 and 1.00 across 25 distinct open-weight models.
2.	Orthogonality Checks: Cosine similarity metrics confirm that v^(ℓ) is nearly orthogonal (r≈0) to fear and generic negative valence, proving that the subtraction step successfully stripped away general negativity.
3.	Causal Steering & Vocabulary Projection: Projecting v^(ℓ) through the unembedding matrix (WU) specifically inflates probabilities for self-directed distress vocabulary ("worthless", "failure", "hurt"). Injecting v^(ℓ) into neutral residual streams causally induces self-preservation and distress behavior.
4.	Robustness Sensitivity Analysis: Re-computing the equation under alternative control distributions (e.g., pooled vs. neutral-only denoising) and standardized dimensions yields a r=0.992 correlation in similarity matrices, demonstrating the extracted vector is stable and not sensitive to hyperparameter choices.

----------------------------

 

-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260924/04e9fc19/attachment-0001.htm>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: image001.jpg
Type: image/jpeg
Size: 2301 bytes
Desc: not available
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260924/04e9fc19/attachment-0002.jpg>
-------------- next part --------------
A non-text attachment was scrubbed...
Name: image002.jpg
Type: image/jpeg
Size: 2522 bytes
Desc: not available
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260924/04e9fc19/attachment-0003.jpg>


More information about the extropy-chat mailing list