This work was largely done during Neel Nanda's MATS 10.0 Exploration Phase.
J Rosser and Dohun Lee are co-first authors for this post with equal contribution. Josh Engels and Neel Nanda supervised the project, and provided guidance and feedback throughout.
TLDR
- Models can acquire undesirable traits from during supervised fine-tuning (SFT). A natural thing to try is to identify the data points with these traits and filter them out and retrain.
- To our surprise, across most of our broad OLMo SFT behaviors, data filtering often has very little effect.
- Most behavior targets like bold formatting, both-side framing, liberal-lean or tendency to say “your feelings are valid” are not affected much under targeted filtering.
- We try many standard black-box/white-box training data attribution methods to find the data to filter, including LLM autoraters, probes, activation-based methods, and gradient-based methods like EKFAC. None of them outperform random baseline on most behaviors.
- For example, despite less than 0.2% of documents both containing the words “feeling/concern” and “valid”, filtering out 10% of documents chosen across TDA methods does not lead to the model saying “Your feelings are valid” any less.
- We test that our training data attribution methods work on a [...]
---
Outline:
(00:30) TLDR
(03:23) Introduction
(05:07) Set Up
(05:10) Speed Run SFT Model Organism
(06:36) Behavior Evaluations
(07:05) In the initial versions of our evaluation, the mid-train often got marked down for failing to stay on task/getting distracted - we edit the judge prompt to not mark down for slop/distractions, full prompt in the appendix.
(07:19) Training Data Attribution (TDA) Methods
(08:38) Data filtering on broad SFT behaviors work much worse than expected
(14:04) Potential Explanations and Limitations
(17:34) Did we actually find any differences between the mid-train base model and SFT?
(21:35) Appendix
(30:00) Toy Test Bed
---
First published:
July 7th, 2026
---
Narrated by TYPE III AUDIO.
---
Images from the article:
Apple Podcasts and Spotify do not show images in the episode description. Try Pocket Casts, or another podcast app.
Fler avsnitt av LessWrong (30+ Karma)
Visa alla avsnitt av LessWrong (30+ Karma)LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.
