Comment on Show HN: LlamaGym – fine-tune LLM agents with online reinforcement learningparentComments−scribu2yActually, there have been attempts to do quantized backprop, but not sure how successfully.
Comments
Actually, there have been attempts to do quantized backprop, but not sure how successfully.