"Furthermore, we note that our model is only trained with cross-entropy loss (supervised learning) without relying on reinforcement learning from human feedback (RLHF). This calls for further investigations of the tradeoffs of simple cross-entropy loss and RLHF training. "
Does this mean RLHF is not really necessary for high quality chatbots?
Someone pointed out that some of the answers are using OpenAI's famous "As an AI...". Soooo you can roughly say that RLHF might still have had an impact here through the training data that came from an RLHF model.
But what we are seeing is a revolution in picking quality demonstration data that might make RLHF an optional last step for fine-tuned models.
I'm bullish on a new technique where the quality of instruction-tuned code model output are measured and evaluated automatically by a machine. I'm calling it RL<machine>F
Code Model outputs a test harness, an implementation, and how to run the implementation for a users inputs. Have RLMF evaluate implementations against the synthetic test harness and against hidden user inputs.
Comments
fantastic. Will keep my 3090 busy for a while!
"Furthermore, we note that our model is only trained with cross-entropy loss (supervised learning) without relying on reinforcement learning from human feedback (RLHF). This calls for further investigations of the tradeoffs of simple cross-entropy loss and RLHF training. "
Does this mean RLHF is not really necessary for high quality chatbots?
Yes and no.
Someone pointed out that some of the answers are using OpenAI's famous "As an AI...". Soooo you can roughly say that RLHF might still have had an impact here through the training data that came from an RLHF model.
But what we are seeing is a revolution in picking quality demonstration data that might make RLHF an optional last step for fine-tuned models.
I'm bullish on a new technique where the quality of instruction-tuned code model output are measured and evaluated automatically by a machine. I'm calling it RL<machine>F
Code Model outputs a test harness, an implementation, and how to run the implementation for a users inputs. Have RLMF evaluate implementations against the synthetic test harness and against hidden user inputs.
The RTX3090 is crackin' it. I am considering a dual set-up with nvlink.