Skip to content

Comment on GLM-5.3 Artificial Analysis Benchmarksparent

Comments

Muse Spark has a nice balance. not to mentions the Contribs version is old deepseek flash prices.

Tested muse spark 1.2 because it was rated so high on design arena, and I've missed a model that can do nice UI in the hands of an operator with no UI skills.

It produced worse UI mockups than GPT and GPT models are already the bottom of the barrel here. The only model that performed well was Kimi K3 - insanely good, but expensive.

It's hard to trust benchmarks these days.

If you just want it to generate UI out of nothing, the benchmarks aren't really for that.

If you want to generate a UI based on specific user input of some kind, then they are.

I'd suggest using one model for UI and another model for tacking onto that UI. LLMs are great at pattern matching, and benchmarks don't really capture one-shotting desirable UI.

That said, benchmaxxing is a thing and your experience with models is a thing. Benchmarks are fuzzy and should be taken with a grain of salt.

I found the sweetspot here: GPT-5.6 Sol (high) 57.3 $0.52 7,545

(Edit: TLDR; It gets on with it, makes the same mistakes you would, without overthinking and overengineering, most of the time)

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.