I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development workflows.
It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.
Comments
I checked out the development benchmarks for agents, and was suprized to find that they're not really representative of many of my own day-to-day development workflows.
It appears that they're mostly testing the ability to make business tasks autonomous, with ~20% associated to development tasks (ssh here, install this, etc.), but not actual programming.
That's probably because the actual coding benchmarks were saturated several years ago.
Which is why you should perform your own benchmarks against your own software stack.
Agree. To this I would add, many things are saturated even for smaller models, which tend to be much cheaper and faster.
On many of my tests, there was no difference in the result between the smaller and bigger model, but there was a big difference in speed and price.