Skip to content

Comment on Show HN: Self-host open-source LLMs on AWS with scale-to-zero

Comments

You got me interested, a rough table of model size to instance type to spot $/hr would help a lot. The 0.5B example is CPU only, so it doesn't say much about what a 7B or 70B actually costs.

You're right. A CPU only model is definitely is not a real prod workload where you'd need GPU machines.

Here is a rough table of model size to instance type to spot $/hr:

For up to 8B, you can use a c6i.xlearge which costs as spot about $0.4/hr.

For up to 70B, you can use a g5.12xlearge that costs you about $2/hr.

Besides the GPU machine be aware that you need a controller instance to route the requests and scale. That has a fixed cost of up to $0.08/hr. When not in use at all, just type veloxml down --all and it tears down the controller too for true $0/hr.

We're currently running larger LLM models benchmarks to add to the README this week.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.