The Data Era Is Over โ€” How the Intelligence of AIs Trained on the Same Web Splits into Three Designs

Why do AIs trained on almost the same web data perform so differently? The answer is not the volume of knowledge but three pieces of design: the filtering schema, the reward rule, and the reasoning logic. A fifty-something's intuition checked against public numbers from Epoch AI, FineWeb, phi-1, InstructGPT, LIMA, and o1.
Markdown sourceยทAnything to add or correct?

The Data Era Is Over โ€” How the Intelligence of AIs Trained on the Same Web Splits into Three Designs

The conclusion first. The gap in AI intelligence does not come from the volume of knowledge. AIs trained on nearly the same web split for three reasons: the filtering schema that selects what knowledge survives, the reward rule (RLHF) that sets the direction of the answer, and the reasoning logic (chain of thought) that fixes the order of thinking. And this conclusion is not my intuition alone; it is confirmed by published papers and benchmark numbers. Put together with the warning that the internet, this sea of knowledge, is drying up, and it becomes clear that the AI industry has already left the "who fed more data" stage and entered the "who chews better and in what order swallows it" stage.

This essay starts from one suspicion and passes through three gates. Along the way, several numbers I originally wrote down were rechecked against public sources and corrected to their accurate values, and expressions without a clear basis were replaced with verifiable figures. The conclusion did not change. If anything it became more solid.

1. The Suspicion: Why Do Students Who Read the Same Textbook Score Differently

AI companies all say they "trained on the vast web data of the internet." But the sea of knowledge is not infinite. In its 2024 paper "Will we run out of data?", the British AI research organization Epoch AI estimated the stock of public human text at roughly 4 x 10^14 tokens (400 trillion), and projected that if current trends continue, training dataset sizes will reach that available stock between 2026 and 2032, with a median estimate of 2028. That is the so-called data wall. My original note that "data runs out around 2026" was the earliest edge of that range. The precise reading is 2026 to 2032, median 2028.

The same movement shows up in search. Search engines now list "scaled content abuse," bulk-generated low-quality text, as a spam violation. That also reads as a consequence of the rising scarcity of fresh, high-quality data for search and training.

So the suspicion begins here. If nearly identical web data was scraped and trained on, shouldn't the AIs' intelligence be roughly the same?

For a person it makes sense. Two students in the same classroom with the same textbook end up first and last because their will, their study habits, and their capacity to absorb differ. A machine has neither will nor habits, so there is no reason for a difference in attitude. And yet the performance gap between ChatGPT, Claude, Gemini and the rest is noticeable enough to feel.

First, a refutation: volume of data does not explain performance

If knowledge volume equaled intelligence, the model with more parameters and more training data would always win. The published numbers say otherwise.

ModelScaleTraining and alignment dataKey resultSource
InstructGPT1.3B parametersFine-tuning with human feedback (RLHF)Preferred over 175B GPT-3 in human evaluation despite being 100x smallerOuyang et al., arXiv:2203.02155 (2022)
phi-11.3B parameters6B tokens of "textbook quality" web data plus 1B tokens of synthetic textbooks (8 A100s, 4 days)HumanEval pass@1 50.6%, MBPP 55.5%Gunasekar et al., arXiv:2306.11644 (2023)
LIMA65B parametersOnly 1,000 curated instruction-response pairs, no reinforcement learningEquivalent to or strictly better than GPT-4 in 43% of cases, 58% vs Bard, 65% vs DaVinci003Zhou et al., arXiv:2305.11206 (2023)
Llama 3 (8B/70B)8B and 70BOver 15 trillion tokens plus a filtering pipelineState of the art at that parameter scaleMeta, Llama 3 announcement (2024-04-18)

A model cut to one hundredth of the size (1.3B) with the right alignment rules beats a 175B model. A 65B model given only 1,000 examples matches or beats GPT-4 in 43% of cases. This did not come from feeding more data. It came from changing how the data is used. The LIMA authors put it the same way: almost all knowledge in a large language model is learned during pretraining, and only limited instruction tuning data is necessary to teach it to produce high quality output.

2. First Gate โ€” The Filtering Schema: What Is Kept and What Is Thrown Away

A large share of internet data is repeated boilerplate, auto-generated spam, ads, hateful language, and context-free fragments. Early AIs swallowed it whole, and that is exactly what Garbage In, Garbage Out means. High-performing models split here. All of them cast a net into the same sea; the tightness of the mesh and the shape of the frame are what differ.

The FineWeb dataset, released by Hugging Face in 2024, assembled 96 Common Crawl snapshots into 15 trillion tokens. A second pass with a separate educational-quality classifier produced FineWeb-Edu at 1.3 trillion tokens. In other words, out of an already filtered 15 trillion tokens, only about 8.7 percent of text with high educational value survived. And in the paper's own words, models pretrained on FineWeb-Edu showed dramatically better performance on knowledge- and reasoning-intensive benchmarks such as MMLU and ARC. More tokens were not added. The rule for deciding what stays in the same sea was changed.

StageSizeWhat it doesSource
Common Crawl raw snapshots96 snapshots for FineWebScrapes text from the web (raw material)FineWeb, arXiv:2406.17557 (2024)
FineWeb15 trillion tokensText extraction, base filters, deduplication, heuristic filtersSame paper
FineWeb-Edu1.3 trillion tokens (about 8.7% of FineWeb)One more pass through an educational quality classifierSame paper
phi-1 training data6B tokens plus 1B syntheticSelects only "textbook quality" web text and injects generated textbooksarXiv:2306.11644
Llama 3 filtering pipelineOver 15 trillion tokensHeuristic filters, NSFW filters, semantic deduplication, text quality classifiersMeta, Llama 3 announcement

The point worth noting is that this recipe is not published. The FineWeb paper says it directly: curation strategies for pretraining datasets are often treated as closely guarded trade secrets. Parameters are open-sourced; the data recipe is kept behind a wall. That admission is itself the argument. The industry concedes that performance differences arise from how data is handled, not from how large the model is. Whether 8.7 percent or 30 percent survives, and on what basis that slice is selected, determines baseline intelligence.

3. Second Gate โ€” The Reward Rule: How a Machine Is Given a "Temper"

What would be a student's study attitude is implemented in machines through reinforcement learning from human feedback (RLHF). People score the outputs, "this answer is logical" or "this should be avoided," and the model is retrained to maximize that score. How this whip-and-carrot logic is designed is what creates the personality of an AI.

As the table above showed, the effect of this gate is dramatic. A 1.3B InstructGPT beating a 175B GPT-3 in human preference was not because it knew more but because the attitude of answering was corrected. The opposite proof exists too: 65B LIMA used no reinforcement learning at all and still matched or beat GPT-4 in 43% of cases with only 1,000 curated examples. Alignment data is a question of quality and form, not quantity.

Company and approachWhat it doesResulting characterSource
OpenAI (GPT-4)Heavy investment in safety and alignment before launch6 months spent on alignment; refusals of disallowed requests down 82%, factual responses up 40%OpenAI, GPT-4 announcement (2023-03)
Meta (Llama 3)Revised post-training proceduresFalse refusal rate cut substantially, alignment and response diversity improvedMeta, Llama 3 announcement
Anthropic (Claude)Constitutional AIHarmlessness aligned through principle documents instead of human labelsBai et al., arXiv:2212.08073 (2022)
Meta (LIMA)Small amounts of high-quality instruction dataFormat learning and generalization with no RL (1,000 examples)arXiv:2305.11206

The same knowledge, in one company, becomes an over-refusing assistant; in another, an assistant that pushes through reasoning. The raw material is identical, the seasoning recipe differs.

4. Third Gate โ€” The Reasoning Logic: From a Machine That Recalls to a Machine That Thinks

This is the part I watched most closely. A high-performing AI does not reflexively blurt out an answer. Internally it follows a procedure: recognize the problem, break it into steps, check for logical contradiction, then derive the final answer. The technique for eliciting that procedure in a prompt is chain of thought; the version trained into the model itself is a reasoning model.

SubjectConditionResultSource
PaLM 540B (GSM8K)Standard prompting17.9%Wei et al., arXiv:2201.11903 (2022)
PaLM 540B (GSM8K)Chain-of-thought prompting56.9%Same paper
GPT-4o (IMO 2024 qualifying)Standard response13% (13.4%)Coverage of OpenAI's o1 announcement (2024-09)
o1 (IMO 2024 qualifying)Internally trained chain of thought plus long thinking time83% (83.3%)Same coverage
o1 (competitive programming)Codeforces89th percentileSame coverage

The numbers say one thing. The move from 17.9% to 56.9% did not come from adding four times more data. A single instruction in the prompt to think step by step more than tripled performance. OpenAI separated two axes in its own description: train-time compute, training the reasoning pattern with reinforcement learning, and test-time compute, letting the model think longer before answering. The second axis matters most. It means performance can improve by extending thinking time without increasing parameters. Extending parameters needs data; extending thinking does not. In front of the data wall, that difference is decisive.

One honest caveat: the IMO 83% and Codeforces 89th percentile are not official competition results but evaluations OpenAI ran and published on its own model, not third-party peer-reviewed results. Even so, set against GPT-4o's 13% on the same problems under the same conditions, the gap is clearly procedural rather than a matter of knowledge volume.

4-4. The Price of Thinking Time: Performance Rises, Throughput Falls

This approach has a cost. At the September 2024 launch, o1-preview was priced at 15 dollars per 1 million input tokens and 60 dollars per 1 million output tokens, 3 times and 4 times GPT-4o respectively, and usage was capped at 30 messages per week (o1-preview) or 50 (o1-mini). When thinking time is extended, answers improve but per-person throughput falls while cost rises. Knowledge did not grow; only the invoice did.

ItemValueVersus GPT-4oMeaning
o1-preview input price15 USD / 1M tokens3xThinking time is billed as computation
o1-preview output price60 USD / 1M tokens4xLong chains of thought are billed as output tokens
Usage limit30 / 50 messages per week-Throughput ceiling drops

This says two things at once. First, parameters require data, but thinking does not. In front of the data wall this is the most practical way around. Second, that detour only moved the bill from data cost to compute cost. Forget that thinking is not free and "a smarter AI" quickly reads as "a more expensive AI."

5. After the Data Wall: Synthetic Data and the Trap Inside It

The Epoch AI paper proposes three roads past the wall: synthetic data generation, transfer learning from data-rich domains, and data efficiency improvements. All three are alternatives to collecting more.

But synthetic data contains a trap. "AI models collapse when trained on recursively generated data," published in Nature 631 (2024), shows that when a model trains repeatedly on its own output it loses the original distribution and drifts into a biased state. The starting corpus held plenty of fact, yet fact disappeared at the speed of the circulation.

So my argument actually strengthens. Whether we use synthetic data or transfer learning, the same three questions come back: what belongs in the generation pool (filtering schema), on what basis the generator is scored (reward rule), and under what procedure the output is checked and discarded (reasoning logic). The source of the data changes; the three gates of design do not. Changing the source does not bring the recipe along.


6. So What About My Own Life โ€” Moving the Three Gates Into My Rules

This essay is about AI, but it ends where people are. The data of a life is finite. Time, the number of people you meet, and the number of days a body lasts all have a total. Just as Epoch AI computed the stock of the web, my stock is more or less fixed. That leaves one question: what to filter, what to reward, and in what order to think.

GateIts name in lifeWhat you actually do
1. Filtering schemaTasteChoose what to read, who to meet, which conversations to leave. Collect ten good pieces first rather than a hundred repeated ones
2. Reward ruleYour own scorecardPraise yourself not for outcomes but for what can be reused: reproducible rules, written records, organized experience
3. Reasoning logicThe order before answeringRecognize, decompose, write down what you do not know, then answer. A procedure, not an instinct

These three lines matter more with age. When young, data (experience) was scarce, so having more of it helped. Now that the total is roughly settled, leaving experience untouched just produces repeated data. Only by filtering, scoring what matters, and fixing the order do repeated events become personal knowledge. That is the answer on the human side of what the machine's intelligence gap is showing us.


7. Objections and Limits: Data Still Matters

It would be a problem if this argued that data is no longer needed. It does not. Four things should be clear.

First, volume still builds the basics. Scaling laws remain valid, and with the same recipe more data wins. Llama 3 using 15 trillion tokens was necessity, not showmanship. The claim is that once data passes a certain scale, quality of filtering and composition beats volume.

Second, the road past the data wall is still unvalidated. The Epoch AI paper itself leaves open whether synthetic data, transfer learning, and efficiency techniques can carry further progress. Research on model collapse shows a real hazard. I make no claim here.

Third, the three gates are hard to verify from outside. Filtering recipes and reward rules are trade secrets, and benchmark numbers are often self-reported, with known overfitting risk. The tables here are therefore not evidence that a number equals a model's intelligence; they are evidence that changing design variables changes outcomes even with the same data.

Fourth, the human analogy has limits. This essay closes by comparing AI to human wisdom, but a person's "will" and a machine's "reward function" are not the same thing. A machine has no will, only an objective function given to it. If anything this strengthens the claim: results split without will, so the difference is purely a difference in design.

8. Conclusion: The Race Has Left the Feeding Stage and Entered the Weaving Stage

To summarize: when everyone fishes in the same sea, the contest depends on who fillets the fish with more precise rules, in which schema it is cooked, and under which logic it is served. The knowledge-accumulation stage is over. The data wall is the closing bell.

GateIts name in AIWhat it doesWhat it corresponds to in a person
1. SelectionFiltering schemaDecides what to keep and discard in the raw materialTaste: choosing what counts as experience
2. AttitudeReward rule (RLHF, alignment)Decides which answers get pointsEducation and culture: what is praised and punished
3. ProcedureReasoning logic (CoT, thinking time)Decides the order of thinking before answeringThe order of reflection: recognize, decompose, verify, conclude

And I do not think this table ends with AI. Past fifty, having lived in the world, human wisdom looks the same. The volume of experience a person gathers evens out at a certain age. The difference between thirty and fifty is not the count of things that happened but which rules filtered them and in what order they were woven into insight. Some people repeat the same mistake with many experiences; others turn a single failure into a structure. The latter is the depth of an adult.

The fact that performance differences in cutting-edge machines come from the structure of how thinking is organized is a heavy question for those of us living in a flood of knowledge and information. What is needed now is not knowing more but designing the rules that knit what you know. In an age when data runs dry, the asset that survives for machines and people alike is not volume but structure.

Sources and References

  • Pablo Villalobos et al., "Will we run out of data? Limits of LLM scaling based on human-generated data", arXiv:2211.04325v2 (2024-06) - effective stock of public human text around 4 x 10^14 tokens; dataset size reaches the stock between 2026 and 2032, median 2028
  • Guilherme Penedo et al., "The FineWeb Datasets", arXiv:2406.17557 (2024-06) - FineWeb 15 trillion tokens from 96 Common Crawl snapshots; FineWeb-Edu 1.3 trillion tokens; curation strategies treated as trade secrets
  • Suriya Gunasekar et al., "Textbooks Are All You Need", arXiv:2306.11644 (2023-06) - phi-1 1.3B parameters, 6B+1B tokens, HumanEval pass@1 50.6%
  • Long Ouyang et al., arXiv:2203.02155 (2022-03) - outputs of 1.3B InstructGPT preferred over 175B GPT-3 despite 100x fewer parameters
  • Chunting Zhou et al., "LIMA: Less Is More for Alignment", arXiv:2305.11206 (2023-05) - 1,000 curated examples, 43% equivalent or better than GPT-4, 58% vs Bard, 65% vs DaVinci003
  • Jason Wei et al., "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models", arXiv:2201.11903 (2022-01) - GSM8K 17.9% standard to 56.9% with CoT
  • Yuntao Bai et al., "Constitutional AI: Harmlessness from AI Feedback", arXiv:2212.08073 (2022-12)
  • Meta, "Introducing Meta Llama 3" (2024-04-18) - over 15 trillion tokens, heuristic and quality-classifier filtering pipeline
  • OpenAI, GPT-4 announcement (2023-03) - 6 months of alignment work, 82% fewer refusals of disallowed requests, 40% more factual responses
  • Coverage of OpenAI o1 (2024-09) and IBL News - IMO 2024 qualifying 83.3% vs GPT-4o 13.4%, Codeforces 89th percentile, API prices of 15 and 60 USD per 1M tokens, weekly limits of 30 and 50 (figures include self-reported evaluations)
  • Shumailov et al., Nature 631 (2024) - model collapse under recursive training
  • Google Search Central spam policies - scaled content abuse demotion

Related posts


This is a free contribution by the operator; fact-checking and table construction were assisted by the Cline agent.