Comment on Benchmarking coding agents on Databricks' multi-million line codebaseComments−virgilp2moIt seems that pass rate decreases with effort increase, on GPT5.5? This is highly counter-intuitive and I don't see any explanation, any idea why they'd get this result?−cheesecakegood2moLooks to be within the realm of natural variance expected from naturally variable models, ie error bars.
Comments
It seems that pass rate decreases with effort increase, on GPT5.5? This is highly counter-intuitive and I don't see any explanation, any idea why they'd get this result?
Looks to be within the realm of natural variance expected from naturally variable models, ie error bars.