Skip to content

Comment on How Apple’s Siri Became One Autistic Boy's B.F.Fparent

Comments

When phones' personal assistants (every company seems to want to have theirs these days) become orders of magnitude more advanced, we may be able to tell them to "check on our vacation booking" and they will know to search through our emails and calendars and connect the dots, so in those tasks they will be more advanced. There will probably be more services tied more deeply into them- for example, you'll be able to say "get me an Uber" and they will reply with "A blue civic will be here in 3 minutes", or "fill up the fridge on Sunday morning" and an Instacart-like company will show up at your door Sunday morning with eggs, vegetables (but no milk because you didn't finish this week's).

However, we won't be able to have those deep meaningful conversations with our voice assistants, because they won't have the necessary life experience to follow meaningful conversations. The current products on the market have no parents, no age, no experiences in school, no previous job, no former lovers, etc (they will jump around those questions playfully, but the canned answers get old quite fast). Those trite details that we recount when we bond with people do not exist in current personal assistants and likely never will.

There are several reasons for this: the first being that being able to model a structure of "life experiences" that can be queried based on what the user is saying is an incredibly complex problem on which we have pretty much no angle of attack on.

12 year old child: "I got picked on at school today." Digital personal assistant: "You know, it happened to me too when I was your age. Let's talk to your mom about it."

Or even more complex:

24 year old student: "Hey, remember that problem on the Trigonometry 402 final from my senior year that I asked you to solve the equations for back in college?" Digital assistant: "Oh yes! Here it is."

This, happening for millions and millions of possible interactions, life experiences, contexts? That's the algorithmic equivalent of light speed travel. Who knows maybe we'll get here one day. But 10 years from now? 50 years from now? 100 years from now? Not a chance. Remember that 40 years ago, Minsky thought it'd take a bunch of grad students a summer to write a program that recognize objects in pictures.

The second reason that such systems are extremely unlikely to emerge is that even if it would be doable technically, it would be extremely expensive to a company. And there would be literally no demand for it, because the overwhelming majority of people don't care about talking to robots. They already don't have enough time to spend with their children, partners, friends, parents... why would anyone waste time with a fake person? No company would invest the billions, if not hundreds of billions of dollars, to solve this problem in the next few hundred years.

Those questions of sentience and whether powering off a digital voice is "killing it" are appealing to ask to those of us who grew up reading Isaac Asimov, because we want this future to exist so bad. But they are red herrings- those questions have no meaningful answers, because our society is not configured in a way in which those questions could actually arise. When you turn your phone off today, the personal assistant definitely doesn't "die"; and in 50 years, even if it can carry out tasks way more efficiently and give somewhat more "human sounding" answers to certain categories of questions, people will still have no problem turning it off.

The second reason that such systems are extremely unlikely to emerge is that even if it would be doable technically, it would be extremely expensive to a company. And there would be literally no demand for it, because the overwhelming majority of people don't care about talking to robots. They already don't have enough time to spend with their children, partners, friends, parents... why would anyone waste time with a fake person? No company would invest the billions, if not hundreds of billions of dollars, to solve this problem in the next few hundred years.

You've missed the point. Why would people watch a TV show and grow to love/hate characters, where there's no interaction between you and them, and all the lines, actions and scenes are scripted and rehearsed. They certainly don't have time to do it, and have children, partners, friends, parents. They would surely not binge watch, schedule time to watch new episodes come rain hail or shine or spend big bucks to visit filming locations etc.

Companies will invest stupid amounts of money in AI that has human qualities. Take the TV Show example: the AI is the "show" and each day they get into entertaining or emotional circumstances that you can "catch up" about, joke about etc. Their charisma is highly engineered to be incredibly engaging and fun to talk to, so you keep coming back for more. Over time they become a "good friend" that you love talking to. They never rebuff you, snub you, have no time for you, make fun of you (too much, just enough to joke around and have some banter). But why do they keep telling you how great Pepsi's new flavours are??

Because the TV show has no bugs, no "I'm sorry, I didn't understand why you said, rephrase" in the middle of your rant against whatever, and creates the same experience for everyone who watches it, so that people can discuss amongst themselves.

Slightly OT, but I'd be tempted to argue TV Shows can and do have bugs; plot holes, bad acting, continuity errors et al. They can be quite jarring.

That's a good point. At least they have less bugs than Siri.

Nevermind that for much of the history of TV, and still today, it randomly cuts out?

Heck, I sometimes say "I don't understand, can you explain?" Why would people be less tolerant of that from a robot than from me?

Not in the middle of a rant, or while you are professing your love for someone. In neutral discourse, sure that's expected.

You can even not speak the same language, humans can understand each other by other means like body language and cultural cues. It's not something you can replicate in a machine unless you really understand how it works. Unless you are saying that human behaviour is completely understood...

Not in the middle of a rant, or while you are professing your love for someone.

It's possible to detect mode of speech even without visual cues. Factors such as pitch dynamics, speed of talk, etc. can be accounted for. Software can be trained to be even more sensitive to these signals than we are.

Of course, the problem is that this varies somewhat among different people. Therefore, part of AI training would need to happen with actual customer after purchase. There're a lot of unknowns in this process from manufacturer's point of view, so I can understand why it's not happening yet.

Once a “special” mode of speech is detected, though, it's simple to avoid canned “Please rephrase” response in case of unclear voice input. Instead the software would change its own mode of speech appropriately, and then evade the direct reply—humans do this all the time in real-life communication.

And there would be literally no demand for it, because the overwhelming majority of people don't care about talking to robots. They already don't have enough time to spend with their children, partners, friends, parents... why would anyone waste time with a fake person?

People don't care about talking with robots because robots aren't able to hold conversations today. A "fake person" is an advantage - it doesn't have needs. It won't manipulate you, it won't think less of you, it won't get angry at you, it won't feel bad because you did something, it has perfect memory and unquestioning loyalty. It's always available. Best of all - if you want it to have the reverse of the above properties, you can always switch it on. Want an arrogant robotic friend? Just go to the settings.

I'm pretty sure most people would prefer robots to humans. Obviously, almost nobody would admit this today.

>24 year old student: "Hey, remember that problem on the Trigonometry 402 final from my senior year that I asked you to solve the equations for back in college?" Digital assistant: "Oh yes! Here it is."

Maybe I'm just being naive from my layman's perspective, but that did not seem that far ahead. It's fairly objective to pinpoint "Trigonometry 402 final from my senior year" in time, and, if the same AI was already there, there's no reason for it not to remember, unless we're systematically deleting stuff to save space.

I think the point is that differentiating between the first and the second is very difficult to do. And communicating to someone not technically inclined why it can respond to question A and not question B would be very frustrating.

Maybe they will be similar like the replicants in blade runner and each have a scripted past instilled into them. As my sibling comment states, its not terribly unlike TV shows.

AboutSource Built by g1lg1l

Hackerly is an independent reader for Hacker News, built on the public HN API. Not affiliated with Y Combinator.