Sveriges mest populära poddar
LessWrong (30+ Karma)
LessWrong (30+ Karma)

“Weight smuggling likely defeats attempts to cap FLOPs per training run” by Paul W, Pierre Peigné

3 min•21 september 2026

Om avsnittet

Epistemic status: >90% confidence in the principle, >70% confidence that mitigating these would be hard in practice, no full implementation yet.

tldr: it seems difficult for verification mechanisms to prevent chaining runs together or aggregating parallel ones; per-run FLOP caps could thus be covertly bypassed.


1. If governments want to regulate frontier AI training, one might want to cap individual training runs, e.g. putting bounds on the number of FLOPs per training run, and making sure each run starts either from scratch (i.e. random initialization at pre-training) or from a known model with accounted FLOPs (e.g. post-training, continuous pre-training, etc).

In the former we could have the initial weights generated by the verifier in some way, or ask for a proof that the weights were initialized with a verifier-controlled random seed. In the later we could ask for a proof that the weights corresponds to unaltered registered weights from a previous training run.

This approach would allow to enforce agreements or regulation targeting a specific model (or model generation) based on FLOPS thresholds: triggering specific evals beyond a specific threshold and possibly proving that a hard threshold has not been reached.


[...]

The original text contained 6 footnotes which were omitted from this narration.

---

First published:
September 21st, 2026

Source:
https://www.lesswrong.com/posts/neT4ncKSiJ6QLt4d7/weight-smuggling-likely-defeats-attempts-to-cap-flops-per

---

Narrated by TYPE III AUDIO.

LessWrong (30+ Karma) med LessWrong finns tillgänglig på flera plattformar. Informationen på denna sida kommer från offentliga podd-flöden.