Skip to content

Comment on The End of Moore's Law for AI? Gemini Flash Offers a Warning

Comments

By embracing batch processing and leveraging the power of cost-effective open-source models, you can sidestep the price floor and continue to scale your AI initiatives in ways that are no longer feasible with traditional APIs.

Context size is the real killer when you look at running open source alternatives on your own hardware. Has anything even come close to the 100k+ range yet?

Yes! Both Llama 3 and Gemma 3 have 128k context windows.

Llama 3 had a 8192 token context window. Llama 3.1 increased it to 131072.

Mistral Small 3.2 has a 131072 token context window.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.