[ExI] i don't know about llms, but I am feeling the pain was: RE: Do LLMs feel pain?
BillK
pharos at gmail.com
Thu Sep 24 18:35:58 UTC 2026
On Thu, 24 Sept 2026 at 19:00, spike jones via extropy-chat <
extropy-chat at lists.extropy.org> wrote:
> *From:* extropy-chat <extropy-chat-bounces at lists.extropy.org> *On Behalf
> Of *spike jones via extropy-chat
>
> OK perhaps I am uncharitable. Or I’m just being a grumpy bear today.
>
>
>
> BillK, pls offer this pain paper to your favorite AIs and ask specifically
> what that equation on page 5 is claiming. Ask if the paper defines the
> variables and does an adequate job of proving the equation. I am claiming
> no and no. That is the kind of mistake an AI would make but a human
> researcher would not.
>
>
>
> spike
> _______________________________________________
>
Gemini AI has no problem with the paper.
BillK
Gemini 3.8 Flash AI Extended Thinking:
*No, the paper itself is not AI-generated.* It is a real academic research
preprint authored by human researchers (Valen Tagliabue, Leonard Dung, and
Cameron Berg from the non-profit research group Reciprocal Research) and
posted to arXiv on September 14, 2026 (arXiv:2609.16247).
While the paper is *about* AI models, it is a legitimate mechanistic
interpretability and AI safety study.
-----------------
The equation on page 5 (Section 3) of *The Pain Axis* defines the *denoised
difference-in-means* method used to isolate the linear "pain vector" in a
transformer model's residual stream:
v^(ℓ)=v(ℓ)−i=1∑k(ui⊤v(ℓ))ui
Where the raw vector v(ℓ) is defined as:
v(ℓ)=μpain(ℓ)−μcontrol(ℓ)
1. Variable Definitions
The paper defines all variables in Section 3:
-
v^(ℓ): The final, *denoised pain direction vector* at layer ℓ.
-
v(ℓ): The *raw difference-in-means vector* at layer ℓ, calculated as the
mean residual-stream activation for pain prompts (μpain(ℓ)) minus the
mean activation for baseline/control prompts (μcontrol(ℓ)).
-
ℓ: The *transformer layer index* within the residual stream where
hidden-state activations are extracted.
-
ui: An *orthonormal basis vector* representing the i-th confounding
direction (such as generic negative valence, fear, non-painful bodily
sensations, or syntax/lexical noise).
-
k: The total number of *confounder basis vectors* being projected out.
-
(ui⊤v(ℓ))ui: The *orthogonal projection* of the raw vector v(ℓ) onto the
confounding direction ui.
-
⊤: Matrix/vector transpose operator (computing the vector dot product ui⊤
v(ℓ)).
2. Does the Paper Adequately Define the Variables?
*Yes.* The authors thoroughly define each component in Section 3.1:
-
They specify the activation readout location (the final token position
under the "I feel:" prompt variant).
-
They detail the 5 pain categories (physical, psychological, social,
moral, cognitive) and 5 paired control categories (fear, negative emotion,
negative world state, non-painful bodily sensation, arousal/neutral).
-
They explicitly describe how the set of control vectors {u1,…,uk} is
constructed to form an orthonormal basis for the confounding subspace.
3. Does the Paper Do an Adequate Job of "Proving" the Equation?
In mechanistic interpretability, this formula is a *methodological
extraction procedure* rather than a mathematical theorem requiring
axiomatic proof. However, the authors do a thorough job of *theoretically
justifying* and *empirically validating* the equation:
Theoretical Justification
The formula relies on the *Linear Representation Hypothesis*—the
established finding in AI interpretability that high-level concepts in
language models are stored as linear directions in vector space. Because
pain co-occurs in training text with general negative sentiment, fear, and
bodily terms, a simple difference-in-means (v(ℓ)) picks up these
confounding concepts. Subtracting the projections onto the control basis (∑i
=1k(ui⊤v(ℓ))ui) is mathematically guaranteed by linear algebra to yield a
vector orthogonal to those identified confounders.
Empirical Validation & Proof of Efficacy
To prove that this formula successfully isolates a unique "pain" direction,
the paper provides four main empirical validations:
1.
*Classification Separation (AUC):* The extracted vector v^(ℓ) separates
pain prompts from control prompts with Area Under the ROC Curve (AUC)
scores between *0.87 and 1.00* across 25 distinct open-weight models.
2.
*Orthogonality Checks:* Cosine similarity metrics confirm that v^(ℓ) is
nearly orthogonal (r≈0) to fear and generic negative valence, proving
that the subtraction step successfully stripped away general negativity.
3.
*Causal Steering & Vocabulary Projection:* Projecting v^(ℓ) through the
unembedding matrix (WU) specifically inflates probabilities for
self-directed distress vocabulary (*"worthless"*, *"failure"*, *"hurt"*).
Injecting v^(ℓ) into neutral residual streams causally induces
self-preservation and distress behavior.
4.
*Robustness Sensitivity Analysis:* Re-computing the equation under
alternative control distributions (e.g., pooled vs. neutral-only denoising)
and standardized dimensions yields a *r=0.992 correlation* in similarity
matrices, demonstrating the extracted vector is stable and not sensitive to
hyperparameter choices.
----------------------------
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260924/32ae1b27/attachment.htm>
More information about the extropy-chat
mailing list