Skip to content

Comment on Ask HN: Is it feasible to run a model on device for complete privacy?parent

Comments

No open source model that’s any good?

the Gemma you tried is tiny, there are 31B and 26B (A4B) variants. there's also Qwen 3.6 with 27B and 35B (A3B) variants, reportedly pretty good. try them on open router or something. these require 30-40 Gb of memory to run between RAM and VRAM, less if quantized beyond near-lossless 8 bit.

there are near-SOTA open models, but they are 1T+ parameters, i.e. they require over a terabyte of memory to run.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.