If you think about wearables there's a more compelling case. If I walk past a car I like I can look at it and get specs, price, who has it for sale, etc. Or if I'm at the grocery store and wow, those tomatoes look amazing, I see recipes. While it seems weird to search by image right now, it will probably feel more natural once it's seamless; it certainly removes friction from the experience.
The thing is I don't think those are really image searches. If I take a picture of a car it's probably a car I can't afford so I'm not really looking for a dealer, I'm probably looking for either 1) what's the make/model, 2) where have I seen that before, 3) what year is?
For produce there are even more things I might be interested in. Is that a good tomato or what is that deformed looking red plant?
Essentially I'm saying that image searches have just never been useful for me without supplying textual context. Speech has the potential to help with that but just doesn't seem to work well in consumer devices yet. Speech will get there but image searches just aren't useful without speech or text.
While it seems weird to search by image right now, it will probably feel more natural once it's seamless;
Especially when we become so "plugged in" that I can search for what I see. Right now, if I want to see "tomato recipes" I'm probably not inclined to pull my phone out, take a picture (and who wants that on their stream anyway?) and then paste picture + recipes? I'd rather just type "tomato recipes". But if I can search what I see and speak, e.g. "oh look, tomatoes? Find recipes!" that's much easier.
Comments
If you think about wearables there's a more compelling case. If I walk past a car I like I can look at it and get specs, price, who has it for sale, etc. Or if I'm at the grocery store and wow, those tomatoes look amazing, I see recipes. While it seems weird to search by image right now, it will probably feel more natural once it's seamless; it certainly removes friction from the experience.
The thing is I don't think those are really image searches. If I take a picture of a car it's probably a car I can't afford so I'm not really looking for a dealer, I'm probably looking for either 1) what's the make/model, 2) where have I seen that before, 3) what year is?
For produce there are even more things I might be interested in. Is that a good tomato or what is that deformed looking red plant?
Essentially I'm saying that image searches have just never been useful for me without supplying textual context. Speech has the potential to help with that but just doesn't seem to work well in consumer devices yet. Speech will get there but image searches just aren't useful without speech or text.
Especially when we become so "plugged in" that I can search for what I see. Right now, if I want to see "tomato recipes" I'm probably not inclined to pull my phone out, take a picture (and who wants that on their stream anyway?) and then paste picture + recipes? I'd rather just type "tomato recipes". But if I can search what I see and speak, e.g. "oh look, tomatoes? Find recipes!" that's much easier.
or you could just type or say "honda civic" - is that really so hard?
A prediction is about what people will do, not what they should do. Taking that into account, is your objection relevant?