Comment on A.I. Is Getting More Powerful, but Its Hallucinations Are Getting WorseparentComments−mountainriver1yThis is exactly it, it’s the result of RLVR, where we force the model to reason about how to get to an answer when that information isn’t in its base training.
Comments
This is exactly it, it’s the result of RLVR, where we force the model to reason about how to get to an answer when that information isn’t in its base training.