Tangential, but why didn't the OpenAssistant team (lead by the author of the video) release the OpenAssistant dataset? As far as I know, the project was shut down, and only some initial highly filtered version of the data got released. This dataset could be very valuable for the community that created it.
The effort started at the ~beginning of February and ended at the ~end of October [1]. The dataset you link is from April and had "unsafe" content filtered.
Comments
Tangential, but why didn't the OpenAssistant team (lead by the author of the video) release the OpenAssistant dataset? As far as I know, the project was shut down, and only some initial highly filtered version of the data got released. This dataset could be very valuable for the community that created it.
They released the latest dump today :) https://huggingface.co/datasets/OpenAssistant/oasst2
It was fully released, no? https://huggingface.co/datasets/OpenAssistant/oasst1
The effort started at the ~beginning of February and ended at the ~end of October [1]. The dataset you link is from April and had "unsafe" content filtered.
[1] https://m.youtube.com/watch?v=gqtmUHhaplo&feature=youtu.be
My mistake - I guess I assumed that when the dataset was released back in April, that was the end of it, I didn't know collection was ongoing.
Looks like the "final" version was released yesterday:
https://huggingface.co/datasets/OpenAssistant/oasst2
As far as filtered/unfiltered goes, I have no idea.