[ExI] i don't know about llms, but I am feeling the pain was: RE: Do LLMs feel pain?

BillK pharos at gmail.com
Thu Sep 24 18:35:58 UTC 2026


On Thu, 24 Sept 2026 at 19:00, spike jones via extropy-chat <
extropy-chat at lists.extropy.org> wrote:

>  *From:* extropy-chat <extropy-chat-bounces at lists.extropy.org> *On Behalf
> Of *spike jones via extropy-chat
>
> OK perhaps I am uncharitable.  Or I’m just being a grumpy bear today.
>
>
>
> BillK, pls offer this pain paper to your favorite AIs and ask specifically
> what that equation on page 5 is claiming.  Ask if the paper defines the
> variables and does an adequate job of proving the equation.  I am claiming
> no and no.  That is the kind of mistake an AI would make but a human
> researcher would not.
>
>
>
> spike
> _______________________________________________
>


Gemini AI has no problem with the paper.
BillK

Gemini 3.8 Flash AI Extended Thinking:

*No, the paper itself is not AI-generated.* It is a real academic research
preprint authored by human researchers (Valen Tagliabue, Leonard Dung, and
Cameron Berg from the non-profit research group Reciprocal Research) and
posted to arXiv on September 14, 2026 (arXiv:2609.16247).

While the paper is *about* AI models, it is a legitimate mechanistic
interpretability and AI safety study.
-----------------

The equation on page 5 (Section 3) of *The Pain Axis* defines the *denoised
difference-in-means* method used to isolate the linear "pain vector" in a
transformer model's residual stream:
v^(ℓ)=v(ℓ)−i=1∑k(ui⊤v(ℓ))ui

Where the raw vector v(ℓ) is defined as:
v(ℓ)=μpain(ℓ)−μcontrol(ℓ)
1. Variable Definitions

The paper defines all variables in Section 3:

   -

   v^(ℓ): The final, *denoised pain direction vector* at layer ℓ.
   -

   v(ℓ): The *raw difference-in-means vector* at layer ℓ, calculated as the
   mean residual-stream activation for pain prompts (μpain(ℓ)) minus the
   mean activation for baseline/control prompts (μcontrol(ℓ)).
   -

   ℓ: The *transformer layer index* within the residual stream where
   hidden-state activations are extracted.
   -

   ui: An *orthonormal basis vector* representing the i-th confounding
   direction (such as generic negative valence, fear, non-painful bodily
   sensations, or syntax/lexical noise).
   -

   k: The total number of *confounder basis vectors* being projected out.
   -

   (ui⊤v(ℓ))ui: The *orthogonal projection* of the raw vector v(ℓ) onto the
   confounding direction ui.
   -

   ⊤: Matrix/vector transpose operator (computing the vector dot product ui⊤
   v(ℓ)).

2. Does the Paper Adequately Define the Variables?

*Yes.* The authors thoroughly define each component in Section 3.1:

   -

   They specify the activation readout location (the final token position
   under the "I feel:" prompt variant).
   -

   They detail the 5 pain categories (physical, psychological, social,
   moral, cognitive) and 5 paired control categories (fear, negative emotion,
   negative world state, non-painful bodily sensation, arousal/neutral).
   -

   They explicitly describe how the set of control vectors {u1,…,uk} is
   constructed to form an orthonormal basis for the confounding subspace.

3. Does the Paper Do an Adequate Job of "Proving" the Equation?

In mechanistic interpretability, this formula is a *methodological
extraction procedure* rather than a mathematical theorem requiring
axiomatic proof. However, the authors do a thorough job of *theoretically
justifying* and *empirically validating* the equation:
Theoretical Justification

The formula relies on the *Linear Representation Hypothesis*—the
established finding in AI interpretability that high-level concepts in
language models are stored as linear directions in vector space. Because
pain co-occurs in training text with general negative sentiment, fear, and
bodily terms, a simple difference-in-means (v(ℓ)) picks up these
confounding concepts. Subtracting the projections onto the control basis (∑i
=1k(ui⊤v(ℓ))ui) is mathematically guaranteed by linear algebra to yield a
vector orthogonal to those identified confounders.
Empirical Validation & Proof of Efficacy

To prove that this formula successfully isolates a unique "pain" direction,
the paper provides four main empirical validations:

   1.

   *Classification Separation (AUC):* The extracted vector v^(ℓ) separates
   pain prompts from control prompts with Area Under the ROC Curve (AUC)
   scores between *0.87 and 1.00* across 25 distinct open-weight models.
   2.

   *Orthogonality Checks:* Cosine similarity metrics confirm that v^(ℓ) is
   nearly orthogonal (r≈0) to fear and generic negative valence, proving
   that the subtraction step successfully stripped away general negativity.
   3.

   *Causal Steering & Vocabulary Projection:* Projecting v^(ℓ) through the
   unembedding matrix (WU) specifically inflates probabilities for
   self-directed distress vocabulary (*"worthless"*, *"failure"*, *"hurt"*).
   Injecting v^(ℓ) into neutral residual streams causally induces
   self-preservation and distress behavior.
   4.

   *Robustness Sensitivity Analysis:* Re-computing the equation under
   alternative control distributions (e.g., pooled vs. neutral-only denoising)
   and standardized dimensions yields a *r=0.992 correlation* in similarity
   matrices, demonstrating the extracted vector is stable and not sensitive to
   hyperparameter choices.

----------------------------
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260924/32ae1b27/attachment.htm>


More information about the extropy-chat mailing list