The ultimate guide to RL environments: building and scaling them in the LLM erahuggingface.co 7kashifr4modiscuss
The Smol Training Playbook: The Secrets to Building World-Class LLMshuggingface.co 265kashifr10mo19 comments
MaPO: A reference-free alignment technique for diffusion modelsmapo-t2i.github.io 2kashifr2y1 comment
OpenHermesPreferences: Dataset of ~1M AI preferences from teknium/OpenHermes-2.5huggingface.co 7kashifr2y1 comment
Diffusers: Modular Diffusion model library from HuggingFacegithub.com/huggingface 47kashifr4y5 comments