Skip to content

Show HN: I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browser

aidekin.com
5 pointsstfurkan2 comments
On HN

Comments

Wow this is a cool prototype and looks amazing. I am surprised some level of baseline intelligence survives this kind of aggro quantization. Do you think 300mb initial download is ok for something like quick website where I want to ask quick support question? Are you planning to have hosted fallback to answer q while download is happening?

Thank you :) Currently I am not planning to have hosted fallback Q&A but it's a nice idea. It should only download the ~300mb initially once and then use the cached model.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.