I wonder where they could have gotten their data from. Their 'about' page says, "If You Dig's results are calculated from the music preferences of a whole lotta real people. Like, a lot."
I'd really like to know where this data comes from. The results are accurate enough, but how can I know they'll become more accurate or improve over time?
To incentivize people to keep contributing, the plan is to build a notification service around your favorite artists, so you'll find out about new artists/albums/songs that are related to your tastes.
So the more popular the site gets, the more data there will be :)
Just picking some random favorites, I looked at "Symphony X", and then flipped over to "Least Likely". There are several artists there -- Adele, The Offspring, Johnny Cash -- that I quite enjoy.
IMHO, it doesn't seem correct to ask people for their "favorite bands". It would seem more accurate to ask them "when you're in a given mood, what bands will you most enjoy hearing?". At the right times, I'd be equally happy to hear Symphony X or Johnny Cash. But at any given time, I'm going to prefer one or the other. So why not just let the contributors spell out those similarity sets for you?
It seems to me that you're grouping things that are similar in some stronger sense than that they're the favorites of some people. And conversely, the way you're getting the data -- asking for favorites you'd be missing out on a bunch of minor bands that aren't anyone's favorite, but still enjoyed by people.
This site finds similar artists to the one you search for. It has no reason to expect you'll like The Offspring when you entered Symphony X.
It's like telling a friend you like Band X, and they say "Oh man, if you like Band X, you'll probably like Similar Band Y! or Similar Band Z!" It would be odd if you said you liked Jay-Z and someone recommended The Clash because they're so different, even though it's not uncommon to like both.
As the Contribute question stands, I might put Symphony X and Johnny Cash on my list of favorites, which would make the site more likely to recommend one in response to the other. But in real life, if someone asked me what I'd recommend to a Johnny Cash fan, there's no way I'd answer Symphony X.
And I propose that the way to address this is to not ask the user "what are your [unrelated] favorites", but rather, "pick a recent moment, and tell us what music you would have most enjoyed hearing at that time". That implies a stronger link than simply that the listed artists share a spot on my list of unrelated favorites.
Oh I know. :) That part's only semi-automated, though, I bring the new data in, in manual batches. If the overall artist popularities in the new data don't match the existing data to within a certain degree (and also watching IP addresses and submission time patterns), then that data won't get counted.
Comments
I wonder where they could have gotten their data from. Their 'about' page says, "If You Dig's results are calculated from the music preferences of a whole lotta real people. Like, a lot."
I'd really like to know where this data comes from. The results are accurate enough, but how can I know they'll become more accurate or improve over time?
New data comes in through http://ifyoudig.net/page/contribute
To incentivize people to keep contributing, the plan is to build a notification service around your favorite artists, so you'll find out about new artists/albums/songs that are related to your tastes.
So the more popular the site gets, the more data there will be :)
Just picking some random favorites, I looked at "Symphony X", and then flipped over to "Least Likely". There are several artists there -- Adele, The Offspring, Johnny Cash -- that I quite enjoy.
IMHO, it doesn't seem correct to ask people for their "favorite bands". It would seem more accurate to ask them "when you're in a given mood, what bands will you most enjoy hearing?". At the right times, I'd be equally happy to hear Symphony X or Johnny Cash. But at any given time, I'm going to prefer one or the other. So why not just let the contributors spell out those similarity sets for you?
It seems to me that you're grouping things that are similar in some stronger sense than that they're the favorites of some people. And conversely, the way you're getting the data -- asking for favorites you'd be missing out on a bunch of minor bands that aren't anyone's favorite, but still enjoyed by people.
This site finds similar artists to the one you search for. It has no reason to expect you'll like The Offspring when you entered Symphony X.
It's like telling a friend you like Band X, and they say "Oh man, if you like Band X, you'll probably like Similar Band Y! or Similar Band Z!" It would be odd if you said you liked Jay-Z and someone recommended The Clash because they're so different, even though it's not uncommon to like both.
That's exactly my point.
As the Contribute question stands, I might put Symphony X and Johnny Cash on my list of favorites, which would make the site more likely to recommend one in response to the other. But in real life, if someone asked me what I'd recommend to a Johnny Cash fan, there's no way I'd answer Symphony X.
And I propose that the way to address this is to not ask the user "what are your [unrelated] favorites", but rather, "pick a recent moment, and tell us what music you would have most enjoyed hearing at that time". That implies a stronger link than simply that the listed artists share a spot on my list of unrelated favorites.
Reading comprehension fail on my part!
This is how you've obtained all of your data? When did the site launch? There was no original set of data that you started with?
Has all the data gotten in that way? I couldn't find anything I listen to.
until Reddit thinks it'll be funny to try and get Rick Astley to the top of every single recommendation.
Oh I know. :) That part's only semi-automated, though, I bring the new data in, in manual batches. If the overall artist popularities in the new data don't match the existing data to within a certain degree (and also watching IP addresses and submission time patterns), then that data won't get counted.