Comment on Show HN: I wrote a 1-bit WebGPU runtime to run a 1.7B LLM in the browserparentComments−stfurkanOP2moThank you :) Currently I am not planning to have hosted fallback Q&A but it's a nice idea. It should only download the ~300mb initially once and then use the cached model.
Comments
Thank you :) Currently I am not planning to have hosted fallback Q&A but it's a nice idea. It should only download the ~300mb initially once and then use the cached model.