Skip to content

Comment on Ox-Alpha Is GLM?

Comments

If it’s not zhipu then why is it returning errors that zhipu does for other models? Who else would return the exact same errors even if they took a lot of core infra like tokenizer from z?

Someone could've trained model on top of GLM. Same way Cognition trained their SWE model on top of Kimi and Cursor did same with their Composer model.

While possible the amount of variation in serving infrastructure is unlikely to land with actually giving the exact same errors zhipu does.

It feels like glm flash, and there was a report zhipu had secured a huge new cluster suggesting they have the capacity. My guess anyway.

https://www.tomshardware.com/tech-industry/artificial-intell...

The reasoning levels are the same as GLM 5.3. GLM 5.3 is still not open...

I believe it's GLM 5.3 Flash or Air.

Reasoning levels are often just injected system prompts so not a great way to fingerprint models.

But it's an error, not a response.

Ziphu has that many resources to be able to serve capacity for 1 quadrillion tokens per day on Nous portal? My bet is that it's a Composer model from Cursor running on xAI cluster, they already did a Composer based on Kimi-K2.5

There's three options here:

- The provider has a massive amount of (unused) hardware. Google or Cursor seem most likely

- The model is extremely efficient, beyond anything we've seen so far

- Whomever made the model has improved the cache efficiency in such a way that it's very cheap to serve. See e.g Deepseeks or Xiaomi caching (pre-price increase)

Option 4: the claimed capacity is not true. Real world usage hasn’t reached anywhere close to it.

    > 1 quadrillion tokens per day on Nous portal
If you are referring to this number (https://xcancel.com/NousResearch/status/2090899914700054780), they are either mistaken, or they mean that they can route 1 quadrillion tokens per day, but the provider behind Ox Alpha certainly can't provide that. Almost all of my requests have hit a rate limit so far.

I was hitting 429 overloaded regularly with Ox on OpenRouter yesterday, but a lot of that turned out to be problems with my harness. I fixed some bugs, improved the back-off, and I haven't hit a 429 error since (touch wood).

OpenRouter says they're doing 6 Trillion tokens a day with Ox Alpha so far, and it has been their biggest launch of all time. OpenCode claimed they had capacity for 100T a day.

https://x.com/OpenRouter/status/2091912024922177562 https://x.com/opencode/status/2090544355824038300

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.