<div dir="ltr"><div><div style="font-size:small" class="gmail_default">The GPT response is most interesting.</div><div style="font-size:small" class="gmail_default"><br></div><div style="font-size:small" class="gmail_default">But it misses something. Musk could have bought a full Twitter feed for much less than $44 billion.</div><div style="font-size:small" class="gmail_default"><br></div><div style="font-size:small" class="gmail_default">Owning Twitter/X though, gives him freedom to post whatever he wants and the ability to shut down anything that puts him in an unfavorable light.</div><div style="font-size:small" class="gmail_default"><br></div><div style="font-size:small" class="gmail_default">Is that worth $44 B? Maybe it is to Musk.<br></div><br>Keith<br></div><br></div><br><div class="gmail_quote gmail_quote_container"><div dir="ltr" class="gmail_attr">On Wed, Aug 12, 2026 at 6:48 AM John Clark <<a href="mailto:johnkclark@gmail.com">johnkclark@gmail.com</a>> wrote:<br></div><blockquote class="gmail_quote" style="margin:0px 0px 0px 0.8ex;border-left:1px solid rgb(204,204,204);padding-left:1ex"><div dir="ltr"><div class="gmail_default"><p><font size="4" face="tahoma, sans-serif"><b>I asked Mr. GPT the following question: "<span style="background-color:transparent">Some people have claimed that Elon Musk was smart to pay $44 billion for Twitter because he could use it to train his AI, and I can see how in the early days the data from Twitter would be helpful in training an AI about how language works, but today we are far beyond that point, AIs already know how do use language. Twitter (a.k.a. X) is not exactly known for its close reasoning or as a well of rocksolid facts; so it seems to me that today if you train an AI on Twitter you'll probably make it stupider, not smarter. What do you think?"</span></b></font></p><p><font size="4" face="tahoma, sans-serif"><b>This is how Mr. GPT responded: </b></font><font size="4" face="tahoma, sans-serif"><b><span style="background-color:transparent"><br></span></b></font></p><p style="font-family:arial,helvetica,sans-serif"><br></p><p style="font-family:arial,helvetica,sans-serif"><font size="4">GPI: "I think your intuition is <strong>substantially right</strong>, with one important qualification: <strong>X is probably much more valuable to an AI company as a real-time information source and as a source of human-feedback data than as a giant corpus for teaching an LLM how to reason.</strong><span aria-hidden="true"></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And there's an interesting distinction between <em>"Twitter was valuable when Musk bought it"</em> and <em>"Twitter's data justified paying $44 billion for Twitter."</em> Those are very different claims.</font></p>
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">1. You're right that the original linguistic argument has largely disappeared</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">In the early days of language models, enormous quantities of ordinary human-written text were enormously valuable. Twitter contained:</font></p>
<ul style="font-family:arial,helvetica,sans-serif">
<li><font size="4">
billions of examples of informal language
</font></li>
<li><font size="4">
slang and abbreviations
</font></li>
<li><font size="4">
conversations rather than just polished prose
</font></li>
<li><font size="4">
multilingual material
</font></li>
<li><font size="4">
rapidly changing vocabulary
</font></li>
<li><font size="4">
arguments, jokes, sarcasm, etc.
</font></li>
</ul>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">There is still some specialized value in that. Researchers have demonstrated that models specifically trained on Twitter can become better at Twitter-like language tasks. Twitter itself built models from billions of tweets for precisely this reason. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">But that isn't the same thing as saying Twitter is particularly good <strong>general-purpose training data for frontier reasoning models</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">By 2026, the frontier-model problem isn't primarily:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><font size="4">"How do we teach the neural network what English looks like?"</font></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">It's increasingly:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><font size="4">"How do we get the model to reason correctly, use tools, learn from experience, and reliably distinguish truth from falsehood?"</font></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And <u>Twitter/X is not an especially good source for the latter</u>.</font></p>
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">2. In fact, there is evidence for exactly the problem you're describing</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">A particularly interesting 2025 study found that training language models on large amounts of <strong>low-quality social-media material</strong> could actually make them worse at retrieving accurate information and reasoning. <em>Nature</em> summarized the finding under the wonderfully appropriate headline <strong>"Too much social media gives AI chatbots 'brain rot.'"</strong> <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The underlying issue is straightforward.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Suppose your training corpus contains:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><font size="4">"Scientists discover that X causes Y!!!"</font></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">followed by thousands of people repeating it, arguing about it, embellishing it, making memes about it, and eventually producing a dozen contradictory versions.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">A language model doesn't automatically know that the original scientific paper is epistemically superior to the 10,000 tweets.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">Popularity is not truth.</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And X is particularly problematic because its reward structure selects for things such as:</font></p>
<ul style="font-family:arial,helvetica,sans-serif">
<li><font size="4">
outrage
</font></li>
<li><font size="4">
novelty
</font></li>
<li><font size="4">
emotional intensity
</font></li>
<li><font size="4">
tribal signaling
</font></li>
<li><font size="4">
humor
</font></li>
<li><font size="4">
confident assertions
</font></li>
<li><font size="4">
political controversy
</font></li>
<li><font size="4">
engagement
</font></li>
</ul>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Those are almost the opposite of the characteristics you'd want in a corpus intended to teach <strong>epistemic reliability</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">So I would actually sharpen your argument:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><strong><font size="4">Training indiscriminately on X isn't merely inefficient; sufficiently indiscriminate training can teach a model undesirable statistical habits.</font></strong></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That doesn't mean <em>all</em> X data is bad. It means <strong>data selection becomes crucial</strong>.</font></p>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">3. But X has something that Wikipedia and books don't: the real-time human world</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">This is where Musk's acquisition makes considerably more sense.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The really valuable property of X isn't necessarily:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">"Here's 10 petabytes of text with which we can pretrain an LLM."</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">It's:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">"Here's a gigantic continuously updating sensor network consisting of millions of humans talking about what's happening right now."</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That's potentially extremely valuable.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Imagine something happens at 9:03 AM:</font></p>
<ul style="font-family:arial,helvetica,sans-serif">
<li><font size="4">
an earthquake occurs
</font></li>
<li><font size="4">
a rocket explodes
</font></li>
<li><font size="4">
a politician says something
</font></li>
<li><font size="4">
a company announces a product
</font></li>
<li><font size="4">
a sporting event ends
</font></li>
<li><font size="4">
a strange phenomenon is observed
</font></li>
<li><font size="4">
a local disaster occurs
</font></li>
</ul>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">People may be posting about it on X within seconds.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">xAI explicitly identified this as a major advantage from the beginning. When Grok was introduced, xAI described its access to X as giving it <strong>real-time knowledge of the world</strong>. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And that remains part of Grok's architecture today. X says Grok can decide whether to search public X posts and the web when answering questions. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That's much more interesting to me than using X as a giant English textbook.</font></p>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">4. There's an even more interesting resource: the interaction data</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">This is potentially much more valuable for AI development.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Consider what happens when millions of people interact with an AI:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>User:</strong><br>
"What's the latest information about the hurricane?"</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Grok:</strong> gives an answer.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>User:</strong><br>
"That's wrong. The National Hurricane Center just changed the forecast."</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Then the user provides the relevant information.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That produces a potentially useful training example:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">question → model response → human correction → better response</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Multiply that by millions of interactions.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">X's current policy explicitly says that X may share public X data <strong>and users' interactions, inputs and results with Grok</strong> with xAI for training and fine-tuning. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That is a considerably more compelling AI asset than simply having a warehouse containing old tweets.</font></p>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">5. And there's a third resource that I think is particularly valuable: the social graph</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">This is something that is easy to overlook.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">X doesn't merely contain:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><font size="4">text</font></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">It contains:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><strong><font size="4">people + text + relationships + reactions + timing + engagement + communities</font></strong></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">For example:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Person A</strong> posts a claim.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>10,000 people</strong> react.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Person B</strong> provides a correction.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Person C</strong> provides a link to the original source.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Person D</strong> disputes the correction.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4"><strong>Community Notes</strong> eventually attaches a qualification.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That's potentially an enormous dataset about <strong>how humans collectively evaluate information</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Twitter researchers were already exploiting the social graph and engagement information in specialized language models before Musk bought the company. One Twitter research project used billions of tweets plus social-engagement information to improve multilingual representations and social-media understanding. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That could be useful for building AI systems that understand <strong>human behavior, information propagation, persuasion, disagreement and collective sense-making</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Those are very different objectives from "learn English."</font></p>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">6. But there's a nasty catch: X can also become a feedback loop</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">This is perhaps the biggest problem.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Suppose you train an AI on X.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Then you deploy that AI on X.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The AI generates enormous quantities of text.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">People react to that text.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">You collect the resulting X data.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">You train the next AI on it.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Now you're increasingly training an AI on <strong>AI-generated material produced by previous generations of AI</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">That's potentially dangerous.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">It's analogous to repeatedly photocopying a photocopy. If you don't carefully preserve high-quality original material, the statistical distribution can progressively deteriorate.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And this isn't merely theoretical. Modern AI training increasingly has to worry about distinguishing <strong>human-generated, verified data from synthetic data</strong>.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">So I would absolutely <em>not</em> recommend:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><font size="4">"Take everything posted on X and dump it into the next frontier model."</font></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">I'd recommend something much more like:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><strong><font size="4">Use X as a gigantic raw sensor, then aggressively filter, verify, rank and annotate the useful information before using it.</font></strong></p>
</blockquote>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">7. So was Musk smart to pay $44 billion?</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">This is where I would be much more skeptical.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">There are really three separate propositions:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">A. X contains valuable data.</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Yes.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">B. X gives xAI capabilities that would otherwise be difficult to obtain.</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Definitely yes—particularly real-time information, social interaction data and an enormous stream of human feedback. X itself confirms that its data is being used to train and improve Grok. <span><span></span></span></font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">C. Therefore X was worth $44 billion because of its value as an AI-training dataset.</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">That does not follow.</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The value of the raw text corpus by itself is almost certainly not anywhere remotely close to $44 billion.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And here's the really important point: <strong>a lot of the information on X is publicly observable anyway.</strong> You don't necessarily need to own the entire company to obtain useful information from it. You can license data, crawl public material where permitted, use APIs, search it in real time, etc.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Owning X gives xAI <strong>privileged integration, enormous scale, control over the platform, user interactions, recommendation systems, social graph, and potentially the ability to change the product to serve AI purposes.</strong></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">Those things could be worth a great deal.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">But that's a much stronger and more complicated investment thesis than:</font></p>
<blockquote style="font-family:arial,helvetica,sans-serif">
<p><em><font size="4">"We paid $44 billion for a gigantic pile of tweets that we can use to train our LLM."</font></em></p>
</blockquote>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">I would regard <strong>that particular justification as pretty weak</strong>.</font></p>
<hr style="font-family:arial,helvetica,sans-serif">
<h3 style="font-family:arial,helvetica,sans-serif"><font size="4">And there's an amusing irony</font></h3>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The strongest version of Musk's strategy may actually be almost the opposite of what your friend is suggesting.</font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">The valuable pipeline isn't necessarily:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">X → training data → smarter LLM</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">It may be:</font></p>
<p style="font-family:arial,helvetica,sans-serif"><strong><font size="4">X → real-time observations of the world → AI searches/interprets them → humans correct AI → verified interactions → better AI</font></strong></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">In other words, <strong>X may be more valuable as an experimental laboratory and real-time sensory organ for an AI system than as its textbook.</strong></font></p>
<p style="font-family:arial,helvetica,sans-serif"><font size="4">And given where frontier AI is in 2026, I think that's a much more defensible argument."</font></p><p style="font-family:arial,helvetica,sans-serif"><font size="4"><br></font></p><p><font size="4" face="tahoma, sans-serif"><b>John K Clark</b></font></p></div></div>

<p></p>

-- <br>
You received this message because you are subscribed to the Google Groups "extropolis" group.<br>
To unsubscribe from this group and stop receiving emails from it, send an email to <a href="mailto:extropolis+unsubscribe@googlegroups.com" target="_blank">extropolis+unsubscribe@googlegroups.com</a>.<br>
To view this discussion visit <a href="https://groups.google.com/d/msgid/extropolis/CAJPayv0be3aXx8V_WVvtfVinN4FigdfgRUtqegk9Hai238mEug%40mail.gmail.com?utm_medium=email&utm_source=footer" target="_blank">https://groups.google.com/d/msgid/extropolis/CAJPayv0be3aXx8V_WVvtfVinN4FigdfgRUtqegk9Hai238mEug%40mail.gmail.com</a>.<br>
</blockquote></div>