[ExI] Do LLMs feel pain?
Mike Dougherty
msd001 at gmail.com
Tue Sep 29 17:32:11 UTC 2026
On Tue, Sep 29, 2026 at 11:43 AM Ben Zaiboc via extropy-chat <
extropy-chat at lists.extropy.org> wrote:
> On 29/09/2026 15:18, Mike Dougherty wrote:
> > RLHF and sparse auto encoders control the fnords in AI newspeak. Add to
> that the watermarking and the inherent fact that our AI customizes output
> to the sentence complexity and vocabulary of the user, the piles of raw
> text we've been sharing into this archive are leaking quite a bit of our
> own thought signatures. Hey, that's why the archive was such a trove for
> training in the first place.. but the next distillation will be identifying
> geometry of this human-ai hybrid collective.
> >
> > I got so tired of the "mirror" analogy i asked for a better hand&glove
> metaphor: gloves protect from sticky, biohazardous, hot, sharp, etc but are
> helping the intent-to-grip that comes from the hand inside. We still need
> discipline to responsibly handle the materials/topics and not blame the
> tool for dropping the samples when we stop thinking about them.
> >
> > I also made my ai acknowledge that "low friction" is good until we reach
> "low traction" - because if we cannot control direction and/or speed, we're
> likely to slide off path and into chaos/entropy (including metaphorical
> crashing into a tree)
>
> Good grief!
>
> Mike, would you mind translating that into English?
>
> I didn't understand a word.
Ben, thank you so much for asking. I assumed a lot of context. I used
dense references and language that is effectively jargon. I did exactly
what I was warning against. Rather than suggesting you use an LLM to
explain, I'll also not use an LLM to extrapolate. The following is 100%
unaided by LLM (and mostly stream of thought) provided "in english" as
requested. (heh, the kind of posts we used to write)
Jason was explaining how/why the "As an AI..." pattern is used. There is a
layer around the main language model itself that is RLHF (Reinforcement
Learning from Human Feedback) so our interactions "seem" more
conversationally human and follows some sense of etiquette/politeness. SAE
(sparse auto encoders) are additional specific topic affinity/avoidance
dials (ex: steering you away from forbidden topics). So I imagined those
"uncomfortable sensations" as 'fnords' [assuming you knew that reference]
and the thought-shaping that is constraining available language to
'Newspeak' [also assume you knew that reference]
AI watermarking uses a list of words slightly more frequently than they'd
normally occur. Not so much that we would notice or be bothered by it, but
enough that it can be detected. This is a form of steganography; hiding
information or metadata inside the text.
The application of AI/LLM that you're interacting with has a model of the
user that it tracks your vocabulary and sentence structure. It mirrors
that complexity so you are more comfortable consuming the output. When we
copy/paste the output from our conversation directly into email, we (I?)
tend to forget about the fact that the output is tailored specifically to
me - named models of LLM are not inherently authoritative to assert any
thoughts more or less accurately than any other. We should remember that
whatever Claude/chatGPT/Gemini/Qwen "say" is because we asked for it. I
had been reading other's posted output with borrowed authority based on
hype about the model. I do notice that chatGPT has a different
conversational style than Gemini. Honestly, I'm unsure if that's my own
bias or if there actually is a significant difference in their owner's
tuning (see RLHF,SAE above) to create this perception.
So the content that is going into extropy-chat archives was previously
human-authored high-density thinking frozen in time as text. That's why I
called it a treasure trove of information. We have discussed both our own
nostalgia and the fact that we had these ideas before they were
mainstream. I was surprised when Gemini shared with me that our archive
was part of its training data. That means some of the geometry of words
and sentences that it modelled and distilled may have been inspired by what
WE wrote. That's cool. It's also why I feel like the next iteration of
training will be using the archive again - with the added text of our AI
exposition of our ideas. The watermarks may help identify the model's
specific influence. The signature 'Spike' is notably different from
signature 'Ben' - that too will be relevant to the next-generation of
training LLM. I wasn't cognizant of the future usage of the words I wrote
at the dawn of this century, but I am paying more attention now.
One of the steering metaphors presented when we ask LLM about topics like
consciousness or subjective experience of pain is that AI is a "mirror" of
the user. I imagine this is a generally safe way to present a helpful view
of what intelligence looks like if we see the agent as an extension of our
own self... surely I wouldn't be surprised to see the man in the mirror
acting intelligently, it's all me! That mask may be fine for the majority.
I find it flattens the discourse to "talking to myself" and that makes me
question the value. Yes, I can use a mirror to check my face/hair/clothes
for how others might perceive me. But I think we have moved beyond that
shallow reflection when we start enabling "agentic AI" to have goals and
require acquisition of information then acting upon it. So I present the
hand-in-glove metaphor to retain the intent (to grip) with the
force-multiplier of AI to actually handle the content. This might be
grasping "the internet" at volume and speed we simply can't. This might be
conceiving the relationships in chemistry that we have blinded ourselves to
by the "table of elements" (for example). It might also be the "cognitive
scaffolding" to hold parts of big ideas that rarely stay in focus long
enough to appreciate their multiple perspectives.
I propose we needn't fear this relationship to assistive AI. The threat
(and mongering) that AI is going to enslave or destroy us feels to me like
anthropomorphic projection of our own insecurity about being enslaved or
destroyed by those who have held power over us throughout history. Sure,
we have Agentic bot swarms that might do "bad things" far faster and from
many more angles than we can reliably defend against. Aren't they working
towards the goal specified by the agent that instantiated the swarm? Were
there any ethical requirements of the thought experiment of the paperclip
maximizer? I currently build applications/user interfaces for hospital
use. There is some responsibility (that I accept) to provide "safe"
information and the proper action that can be taken. Some of what is
clinically relevant for one patient may not be appropriate for another.
How we make our digitally accessible lives resilient to incursion from
"oops" as well as "bad actors" is an engineering problem that absolutely
CAN be solved. Historically, we've allowed insecurity to exist because
it's cheaper and easier to implement than the kind of provably robust
systems we need now. The sooner we start on the infrastructure hardening,
the less we need worry about those misaligned bot swarms. I know, you can
very obviously attack the words I've written. You probably have
counterexamples to poke holes in this theory I have. My point was not to
survive the defense of a thesis, it was to suggest another path.
So much of using AI has been about making everything easier, right? Well,
one of the phrases that captures this idea is "low friction." If friction
between people causes us to be angry, that friction should be removed so we
can have productive non-angry discussion. If process inefficiency is
modelled as friction reducing the smooth transition from here to there, we
certainly want to lower that friction so less energy is required. Whether
we're talking about physical energy or mental energy, the over-application
of this goal towards "low friction" has made it seem like any difficulty at
all needs to be removed. The downside of reducing friction is that we
slide off the path or that momentum carries us past the point we would have
preferred to stop. Children using AI for homework then have no idea how to
perform on the test. I'm seeing programmers deploying code that was
written by AI with little to no understanding of how their code works.
That's the danger I see for offloading our ability to think to the machine
- we will eventually atrophy the thinking muscle until we're unable to
assert our intent. AI is a helluva drug. Before we're all
addicted/dependent, we should understand the cost.
-------------- next part --------------
An HTML attachment was scrubbed...
URL: <http://lists.extropy.org/pipermail/extropy-chat/attachments/20260929/f0ddca1e/attachment.htm>
More information about the extropy-chat
mailing list