And since you're not running those GPUs efficiently at 95% utilization rates or higher to serve tokens profitably, you paid far more than the people who simply pay per token from Together AI or Fireworks AI or something. They also get ZDR/session privacy. If you're really tin-foil hatted, you can go to a TEE cloud like Alpha Compute or something which is probably safer/more secure than your own computer.
You're very right to be wanting to use open sourced models. You're delusional for thinking that buying your GPUs directly saves any money. You are not a cloud service provider, don't act like one.
I have the one of the cheapest private and sovereign AI solutions money can buy, and I get a good laugh whenever the big providers are down.
Also, $1200 for a GPUs that can produce ~7M tokens a day of Qwen 3.8 27b at 80tps is faster and cheaper than any major provider can serve a model of that class as far as I am aware. Pays for itself pretty fast.
I mean I full source bootstrap deterministic operating systems for secure enclaves from zero. And by from zero, I mean from 180 bytes of human reviewable hex machine code all the way up to a llvm/rust/musl toolchain, custom rust init system, job manager, and a full cryptographic remote attestation stack. Also recently custom bootstrap compilers.
Truly I am not aware of many more complex problems in systems engineering than these, which is why I love working on them, though every new line written is exactly what I would have typed myself when I use LLMs. I mostly use LLMS to help me debug and surgically -delete- dead code and deps to get to results small enough to review in full.
And ~20M tokens a day results in about the max output I can keep up with and carefully review at key checkpoints.
My assumption is the fact my tokens are not unlimited and have ratcheted up one GPU at a time it has caused me to stay much more connected to every line written, but also I am a security engineer working on tech that cannot fail and must be reviewed by a minimum of two humans.
For someone working on video games, I imagine a more lax move fast and break things approach to LLMs might make more sense.
I cannot even comprehend what kind of project could possibly need 3B tokens in a day and still produce results a human could actually hope to review, but do share because I am curious!
Comments
And since you're not running those GPUs efficiently at 95% utilization rates or higher to serve tokens profitably, you paid far more than the people who simply pay per token from Together AI or Fireworks AI or something. They also get ZDR/session privacy. If you're really tin-foil hatted, you can go to a TEE cloud like Alpha Compute or something which is probably safer/more secure than your own computer.
You're very right to be wanting to use open sourced models. You're delusional for thinking that buying your GPUs directly saves any money. You are not a cloud service provider, don't act like one.
I have the one of the cheapest private and sovereign AI solutions money can buy, and I get a good laugh whenever the big providers are down.
Also, $1200 for a GPUs that can produce ~7M tokens a day of Qwen 3.8 27b at 80tps is faster and cheaper than any major provider can serve a model of that class as far as I am aware. Pays for itself pretty fast.
7M tokens in one day? My record is 3 Billion. So I guess it depends on the projects you are doing for how useful this is. Maybe one day...
I mean I full source bootstrap deterministic operating systems for secure enclaves from zero. And by from zero, I mean from 180 bytes of human reviewable hex machine code all the way up to a llvm/rust/musl toolchain, custom rust init system, job manager, and a full cryptographic remote attestation stack. Also recently custom bootstrap compilers.
Truly I am not aware of many more complex problems in systems engineering than these, which is why I love working on them, though every new line written is exactly what I would have typed myself when I use LLMs. I mostly use LLMS to help me debug and surgically -delete- dead code and deps to get to results small enough to review in full.
And ~20M tokens a day results in about the max output I can keep up with and carefully review at key checkpoints.
My assumption is the fact my tokens are not unlimited and have ratcheted up one GPU at a time it has caused me to stay much more connected to every line written, but also I am a security engineer working on tech that cannot fail and must be reviewed by a minimum of two humans.
For someone working on video games, I imagine a more lax move fast and break things approach to LLMs might make more sense.
I cannot even comprehend what kind of project could possibly need 3B tokens in a day and still produce results a human could actually hope to review, but do share because I am curious!