
“Speak Not of Data Inefficiency” by axel_sdq
Om avsnittet
Humans are really bad at comparing themselves to models.
Take, for example, Leopold's preschooler graph:
Set aside, for a moment, concerns of scaling and RSI, and meditate: what characteristics does GPT-2 share with a preschooler?
- Rudimentary control of human language
- The ability to count the R's in "strawberry" incorrectly
- ...
I can't come up with anything else, because these two entities are almost completely disjoint. Is GPT-2 capable of bipedal locomotion, recognizing its mother's voice, or naming people by face? Does a preschooler learn from eight million scraped web pages sourced from Reddit?
What exactly is a preschooler "trained" on? Thousands of hours of "multimodal data". Assuming that a preschooler sees at 720p, and is awake for 12,000 hours by the age of 3 (about 11 hours a day), they have consumed at least 27 terabytes of video data at streaming-quality compression, or about 3.6 petabytes uncompressed, not to mention audio and sensorimotor data.
In model-size terms, how large is a preschooler? Trillions of parameters, maybe. Beren Millidge's estimate, which assumes only ~1,000 synapses per neuron, puts the whole brain at an effective 10-30 trillion parameters.
So surely GPT-2 is much more data-efficient than humans, wielding "preschooler"-level control [...]
The original text contained 2 footnotes which were omitted from this narration.
---
First published:
October 1st, 2026
Source:
https://www.lesswrong.com/posts/Fyv5RNMWtESYK2DdL/speak-not-of-data-inefficiency
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.