Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Speak Not of Data Inefficiency” by axel_sdq

6 min•1 oktober 2026

Om avsnittet

Humans are really bad at comparing themselves to models.

Take, for example, Leopold's preschooler graph:

Set aside, for a moment, concerns of scaling and RSI, and meditate: what characteristics does GPT-2 share with a preschooler?

  • Rudimentary control of human language
  • The ability to count the R's in "strawberry" incorrectly
  • ...

I can't come up with anything else, because these two entities are almost completely disjoint. Is GPT-2 capable of bipedal locomotion, recognizing its mother's voice, or naming people by face? Does a preschooler learn from eight million scraped web pages sourced from Reddit?

What exactly is a preschooler "trained" on? Thousands of hours of "multimodal data". Assuming that a preschooler sees at 720p, and is awake for 12,000 hours by the age of 3 (about 11 hours a day), they have consumed at least 27 terabytes of video data at streaming-quality compression, or about 3.6 petabytes uncompressed, not to mention audio and sensorimotor data.

In model-size terms, how large is a preschooler? Trillions of parameters, maybe. Beren Millidge's estimate, which assumes only ~1,000 synapses per neuron, puts the whole brain at an effective 10-30 trillion parameters.

So surely GPT-2 is much more data-efficient than humans, wielding "preschooler"-level control [...]

The original text contained 2 footnotes which were omitted from this narration.

---

First published:
October 1st, 2026

Source:
https://www.lesswrong.com/posts/Fyv5RNMWtESYK2DdL/speak-not-of-data-inefficiency

---

Narrated by TYPE III AUDIO.

---

Images from the article:

Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.