Comment on Browser Agent Benchmark: Comparing LLM models for web automationComments−pixel_popping7moIt's lacking the best model (Opus 4.5) on the benchmark tho.−djohnston7moYeah but then their own product might not score the highest.−pixel_popping7moExactly why I'm pointing it out, which feels a bit corrupt, but understandable.−djohnston7motbh i was a bit cranky yesterday - even if they are #2 on a legit benchmark that would be impressive
Comments
It's lacking the best model (Opus 4.5) on the benchmark tho.
Yeah but then their own product might not score the highest.
Exactly why I'm pointing it out, which feels a bit corrupt, but understandable.
tbh i was a bit cranky yesterday - even if they are #2 on a legit benchmark that would be impressive