I'm a co-author on a recent blog post from METR about the NanoGPT speed-run here [1]. I think it'd be of interest to anyone who enjoyed the original post. Appreciate the good beefy runs and spend here, it's a (from my experience) not super easy to do!
(Also: just to label this comment clearly: it's written hastily from a car, and based on lighter skim of the original blog post [2] than would be ideal. Please correct any mistakes or misinterpretations I have here!)
A few callouts:
1. If I understand the experiment correctly, they start the models at the original baseline. If this is true, I have some worries about contamination. Appendix C [3] has some notes on likely contamination we observed in recent models. This makes interpretation a bit harder.
2. If you look at the token scaling plots in the original post: not all models are hitting a performance plateau. This is an important point: we shouldn't treat these results as a full upper-bound on capabilities, but rather some bound on model performance @ cost (assuming good scaffolding, etc).
3. Our post is mostly about how to _interpret_ the results given here. Quoting from our post: "If we can estimate performance as a function of cost for both humans and agents, we can measure the “expenditure horizon” as the point at which those curves cross: the budget at which humans become more cost-effective than AIs. "
Feedback appreciated. I think you can see expenditure horizon as a sibling methodology (that is much less validated) to METR's time horizon work [4] - roughly, instead of baselining against the time it takes humans to complete tasks, you baseline against cost. This may be better suited to some types of problems similar to NanoGPT.
Comments
Cool post!
I'm a co-author on a recent blog post from METR about the NanoGPT speed-run here [1]. I think it'd be of interest to anyone who enjoyed the original post. Appreciate the good beefy runs and spend here, it's a (from my experience) not super easy to do!
(Also: just to label this comment clearly: it's written hastily from a car, and based on lighter skim of the original blog post [2] than would be ideal. Please correct any mistakes or misinterpretations I have here!)
A few callouts:
1. If I understand the experiment correctly, they start the models at the original baseline. If this is true, I have some worries about contamination. Appendix C [3] has some notes on likely contamination we observed in recent models. This makes interpretation a bit harder.
2. If you look at the token scaling plots in the original post: not all models are hitting a performance plateau. This is an important point: we shouldn't treat these results as a full upper-bound on capabilities, but rather some bound on model performance @ cost (assuming good scaffolding, etc).
3. Our post is mostly about how to _interpret_ the results given here. Quoting from our post: "If we can estimate performance as a function of cost for both humans and agents, we can measure the “expenditure horizon” as the point at which those curves cross: the budget at which humans become more cost-effective than AIs. "
Feedback appreciated. I think you can see expenditure horizon as a sibling methodology (that is much less validated) to METR's time horizon work [4] - roughly, instead of baselining against the time it takes humans to complete tasks, you baseline against cost. This may be better suited to some types of problems similar to NanoGPT.
[1] https://metr.org/blog/2026-07-21-expenditure-horizon/ (Most of this work was my coauthors listed on the post, not me. I'll claim credit for any mistakes though :) ).
[2] https://www.primeintellect.ai/blog/measuring-autonomous-rese...
[3] https://metr.org/blog/2026-07-21-expenditure-horizon/#append...
[4]https://metr.org/time-horizons/
(Edit: METR is hiring. Email is in bio if you're interested in helping AI companies and wider society understand the capabilities and risks of AI.)