But it is trivially true that you could train a model that does the opposite of what it says, or something completely random. The interpretability is incidental.
That's like saying conversational question-answering is incidental to the RLHF post-training.
The objective of most large scale post training i am aware of is not to make think traces or to make them “more English and answer aligned”. It’s to achieve a result. That’s why I am calling it incidental
Comments
That's like saying conversational question-answering is incidental to the RLHF post-training.
The objective of most large scale post training i am aware of is not to make think traces or to make them “more English and answer aligned”. It’s to achieve a result. That’s why I am calling it incidental