The energy/compute cost per performance is not a good ratio due to current optimization. Old hardware makes it worse.
Consider making friends with people who have a good desktop or laptop computer to see if you can use it for a little while when visiting and making them a meal or coffee.
If you give up on local, it reduces the cost by using servers.
Give up on performance and allow hallucination for an introduction to llm’s is my only option for budget and local. A very specific spellcheck or similar based llm would be possible on limited hardware.
Iirc, there is a publication on 1.3bit or 1.4bit quantization that someone implemented on GitHub.
Comments
The energy/compute cost per performance is not a good ratio due to current optimization. Old hardware makes it worse.
Consider making friends with people who have a good desktop or laptop computer to see if you can use it for a little while when visiting and making them a meal or coffee.
If you give up on local, it reduces the cost by using servers.
Give up on performance and allow hallucination for an introduction to llm’s is my only option for budget and local. A very specific spellcheck or similar based llm would be possible on limited hardware.
Iirc, there is a publication on 1.3bit or 1.4bit quantization that someone implemented on GitHub.
1.58 bits https://old.reddit.com/r/LocalLLaMA/comments/1bpa6ol/unoffic...