Comment on MicroGPT-C in pure C hits 10M TPS on Apple M5parentComments−alightsoul25dI think the bitter lesson only talks about task performance but not computational efficiency. Could tiny models improve efficiency? Maybe by just using a general architecture on specialized data, so the artichecture itself is not task specific?
Comments
I think the bitter lesson only talks about task performance but not computational efficiency. Could tiny models improve efficiency? Maybe by just using a general architecture on specialized data, so the artichecture itself is not task specific?