Comment on Quantitative AI progress needs accurate and transparent evaluationparentComments−sebzim45001yIf every photo in streetview was included in the training data of a multimodal LLM it would be like 99.9999% of the training data/resource costs.It just isn't plausible that anyone has actually done that. I'm sure some people include a small sample of them, though.−bluefirebrand1yWhy would every photo in streetview be required in order to have Geoguessr's dataset in the training data?−bee_rider1yI’m pretty sure they are saying that Geoguessr's just pulls directly from Google Streetview. There isn’t a separate Geoguessr dataset, it just pulls from Google’s API (at least that’s what Wikipedia says).−bluefirebrand1yI suspect that Geoguessr's dataset is a subset of Google Streetview, but maybe it really is just pulling everything directly−bee_rider1yMy guess would be that they pull directly from street-view, maybe with some extra filtering for interesting locations.Why bother to create a copy, if it can be avoided, right?−clbrmbr1yYet.This is a good rebuttal when someone quips that we “are about to run out of data”. There’s oh so much more, just not in the form of books and blogs.
Comments
If every photo in streetview was included in the training data of a multimodal LLM it would be like 99.9999% of the training data/resource costs.
It just isn't plausible that anyone has actually done that. I'm sure some people include a small sample of them, though.
Why would every photo in streetview be required in order to have Geoguessr's dataset in the training data?
I’m pretty sure they are saying that Geoguessr's just pulls directly from Google Streetview. There isn’t a separate Geoguessr dataset, it just pulls from Google’s API (at least that’s what Wikipedia says).
I suspect that Geoguessr's dataset is a subset of Google Streetview, but maybe it really is just pulling everything directly
My guess would be that they pull directly from street-view, maybe with some extra filtering for interesting locations.
Why bother to create a copy, if it can be avoided, right?
Yet.
This is a good rebuttal when someone quips that we “are about to run out of data”. There’s oh so much more, just not in the form of books and blogs.