Skip to content

Hidden Rate Limits: How Providers Throttle LLM Throughput During Peak Demand

flyflow.dev
14 pointscarlcortright4 comments
On HN

Comments

Over the past few days we did an investigation of the main LLM providers, and have observed up to a 40% difference in average speed (tokens / second) from the leading LLM providers like GPT4.

I tried GPT4-turbo in the minutes following its announcement and it was blazingly fast.

They are quantifying it, trying to really measure might help your argument.

Looking ahead, I suspect as AI becomes even more ubiquitous/mainstream, AI service providers will offer various levels of analysis at different price points. E.g. the cheapest service will provide reliably accurate answers, but only to simple queries that consume little compute power.

Also envisioned is the all too common race-to-the-bottom scenario where services will simply tune their service to respond with the least compute power needed while harvesting and capitalizing on it users data.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.