Sveriges mest populära poddar
LessWrong (30+ Karma)

“Data filtering works a lot worse than you would expect” by Dohun Lee, J Rosser, Josh Engels, Neel Nanda

33 min7 juli 2026

This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase.

J Rosser and Dohun Lee are co-first authors for this post with equal contribution. Josh Engels and Neel Nanda supervised the project, and provided guidance and feedback throughout.

TLDR

  • Models can acquire undesirable traits from during supervised fine-tuning (SFT). A natural thing to try is to identify the data points with these traits and filter them out and retrain.
  • To our surprise, across most of our broad OLMo SFT behaviors, data filtering often has very little effect.
  • Most behavior targets like bold formatting, both-side framing, liberal-lean or tendency to say “your feelings are valid” are not affected much under targeted filtering.
  • We try many standard black-box/white-box training data attribution methods to find the data to filter, including LLM autoraters, probes, activation-based methods, and gradient-based methods like EKFAC. None of them outperform random baseline on most behaviors.
  • For example, despite less than 0.2% of documents both containing the words “feeling/concern” and “valid”, filtering out 10% of documents chosen across TDA methods does not lead to the model saying “Your feelings are valid” any less.
  • We test that our training data attribution methods work on a [...]

---

Outline:

(00:30) TLDR

(03:23) Introduction

(05:07) Set Up

(05:10) Speed Run SFT Model Organism

(06:36) Behavior Evaluations

(07:05) In the initial versions of our evaluation, the mid-train often got marked down for failing to stay on task/getting distracted - we edit the judge prompt to not mark down for slop/distractions, full prompt in the appendix.

(07:19) Training Data Attribution (TDA) Methods

(08:38) Data filtering on broad SFT behaviors work much worse than expected

(14:04) Potential Explanations and Limitations

(17:34) Did we actually find any differences between the mid-train base model and SFT?

(21:35) Appendix

(30:00) Toy Test Bed

---

First published:
July 7th, 2026

Source:
https://www.lesswrong.com/posts/aTybJ6CPQrxEY8rE2/data-filtering-works-a-lot-worse-than-you-would-expect

---

Narrated by TYPE III AUDIO.

---

Images from the article:




















Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.

Fler avsnitt av LessWrong (30+ Karma)

Visa alla avsnitt av LessWrong (30+ Karma)

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.