These seems very tiny models and as I understand it LLMs behave fairly differently at different scales.
The speed performance gain seems to only be on an M2 chip and I wonder if there's already much better non-GPU optimized attention approaches out there for those use cases.
Comments
These seems very tiny models and as I understand it LLMs behave fairly differently at different scales.
The speed performance gain seems to only be on an M2 chip and I wonder if there's already much better non-GPU optimized attention approaches out there for those use cases.