This was my feeling too. Considering a bunch of new images models are coming out at once claiming they can all spell now implies to me it was likely just a training set issue - the caption generators just needed to be told to include any text in the images.
Comments
This was my feeling too. Considering a bunch of new images models are coming out at once claiming they can all spell now implies to me it was likely just a training set issue - the caption generators just needed to be told to include any text in the images.