IP Library Granted Patent US 10,311,856
Granted Patent B2
US 10,311,856 · App. 15/815,375 · Granted Jun 4, 2019

Synthesized voice selection for computational agents

Inventors: Valerie Nygaard (Mountain View, CA); Bogdan Caprita (Mountain View, CA); Robert Stets (Mountain View, CA); Saisuresh Krishnakumaran (Mountain View, CA); Jason Brant Douglas (San Francisco, CA)
Assignee: Google LLC
G10L13/043G06F3/167G06F16/951G10L13/04G10L15/26G10L13/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,311,856
App. No.
15/815,375
Granted
Jun 4, 2019
Kind
B2
Abstract

An example method includes receiving, by a computational assistant executing at one or more processors, a representation of an utterance spoken at a computing device; selecting, based on the utterance, an agent from a plurality of agents, wherein the plurality of agents includes one or more first party agents and a plurality of third-party agents; responsive to determining that the selected agent comprises a first party agent, selecting a reserved voice from a plurality of voices; and outputting synthesized audio data using the selected voice to satisfy the utterance.

Claims (78)

1. A method comprising:

receiving, by a computational assistant executing at one or more processors, a representation of an utterance spoken at a computing device;

selecting, based on the utterance, an agent from a plurality of agents, wherein the plurality of agents includes one or more first party agents and a plurality of third-party agents;

responsive to determining that the selected agent comprises a first party agent, selecting a reserved voice from a plurality of voices, wherein the reserved voice is associated with the first party agent;

obtaining, by the first party agent and based on the utterance, a plurality of search results, including a first sub-set of the search results and a second sub-set of the search results; and

outputting, for playback by one or more speakers of the computing device, to satisfy the utterance:

synthesized audio data, of the first party agent, that represents the first sub-set of the search results using the selected reserved voice, and

synthesized audio data, of the first party agent, that represents the second sub-set of the search results using an additional voice from the plurality of voices, wherein the additional voice is distinct from the selected reserved voice.

2. The method of claim 1 , wherein the utterance comprises a first utterance, the method further comprising:

receiving a representation of a second utterance spoken at the computing device;

selecting, based on the second utterance, a second agent from the plurality of agents;

responsive to determining that the selected second agent comprises a third-party agent, selecting a voice from the plurality of voices other than the reserved voice; and

outputting synthesized audio data using the selected voice to satisfy the second utterance.

3. The method of claim 2 , wherein selecting the second agent from the plurality of agents comprises:

determining the second utterance includes one or more tasks to be performed by at least one of the plurality of agents;

determining a capability level for each of the plurality of agents to perform one or more of the tasks to be performed by the at least one of the plurality of agents; and

responsive to determining that the capability level for each of the plurality agents, selecting the second agent from the plurality of agents based on the second agent having the highest capability level.

4. The method of claim 2 , further comprising:

determining the second utterance includes a multi-element task to be performed by at least one of the plurality of agents, wherein the multi-element task includes at least a first sub-set of elements and a second sub-set of elements;

performing, by the selected second agent, the first sub-set of elements of the multi-element task;

determining the selected second agent cannot perform the second sub-set of elements of the multi-element task;

selecting an additional agent to perform the second sub-set of elements of the multi-element task based on determining the additional agent is capable of performing the second sub-set of elements of the multi-element task; and

performing, by the additional agent, the second sub-set of elements of the multi-element task.

5. The method of claim 1 , wherein the one or more processors are included in the computing device.

6. The method of claim 1 , wherein the one or more processors are included in a computing system.

7. The method of claim 1 , further comprising:

subsequent to outputting both the synthesized audio data that represents the first sub-set of the search results and the synthesized audio data that represents the second sub-set of the search results:

outputting, by one or more of the speakers of the computing device, a request for feedback, from a user of the computing device, about the second sub-set of search results; and

in response to outputting the request for feedback, receiving a representation of a user sentiment toward the second sub-set of search results.

8. The method of claim 7 , further comprising:

based on the user sentiment toward the second sub-set of search results, adjusting a ranking of the second sub-set of search results.

9. The method of claim 1 , wherein the outputting to satisfy the utterance further comprises:

outputting, by one or more user interfaces of the computing device, an indication of the first sub-set of the search results in a first font; and

outputting, by one or more of the user interfaces of the computing device, an indication of the second sub-set of the search results in a second font.

10. A computing device comprising:

at least one processor; and

at least one memory comprising instructions that when executed, cause the at least one processor to execute an assistant configured to:

receive, from one or more microphones operably connected to the computing device, a representation of an utterance spoken at the computing device;

select, based on the utterance, an agent from a plurality of agents, wherein the plurality of agents includes one or more first party agents and a plurality of third-party agents, the memory further comprising instructions that when executed, cause the at least one processor to:

select, in response to determining that the selected agent comprises a first party agent, a reserved voice from a plurality of voices, wherein the reserved voice is associated with the first party agent;

obtain, by the first party agent and based on the utterance, a plurality of search results, including a first sub-set of the search results and a second sub-set of the search results; and

output, for playback by one or more speakers operably connected to the computing device, to satisfy the utterance:

synthesized audio data, of the first party agent, that represents the first sub-set of the search results using the selected reserved voice, and

synthesized audio data, of the first party agent, that represents the second sub-set of the search results using an additional voice from the plurality of voices, wherein the additional voice is distinct from the selected reserved voice.

11. The device of claim 10 , wherein the utterance comprises a first utterance, the assistant further configured to:

receive a representation of a second utterance spoken at the computing device;

select, based on the second utterance, a second agent from the plurality of agents;

select, responsive to determining that the selected second agent comprises a third-party agent, a voice from the plurality of voices other than the reserved voice; and

output synthesized audio data using the selected voice to satisfy the second utterance.

12. A computing system comprising:

one or more communication units;

at least one processor; and

at least one memory comprising instructions that when executed, cause the at least one processor to execute an assistant configured to:

receive, from a computing device and via the one or more communication units, a representation of an utterance spoken at the computing device; and

select, based on the utterance, an agent from a plurality of agents, wherein the plurality of agents includes one or more first party agents and a plurality of third-party agents, the memory further comprising instructions that when executed, cause the at least one processor to:

select, in response to determining that the selected agent comprises a first party agent, a reserved voice from a plurality of voices, wherein the reserved voice is associated with the first party agent;

obtain, by the first party agent and based on the utterance, a plurality of search results, including a first sub-set of the search results and a second sub-set of the search results; and

output, for playback, to satisfy the utterance:

synthesized audio data, of the first party agent, that represents the first sub-set of the search results using the selected reserved voice, and

synthesized audio data, of the first party agent, that represents the second sub-set of the search results using an additional voice from the plurality of voices, wherein the additional voice is distinct from the selected reserved voice.

13. The system of claim 12 , wherein the utterance comprises a first utterance, the assistant further configured to:

receive a representation of a second utterance spoken at the computing device;

select, based on the second utterance, a second agent from the plurality of agents;

select, responsive to determining that the selected second agent comprises a third-party agent, a voice from the plurality of voices other than the reserved voice; and

output synthesized audio data using the selected voice to satisfy the second utterance.

14. A non-transitory computer-readable storage medium storing instructions that, when executed, cause one or more processors to execute an assistant configured to:

receive a representation of an utterance spoken at a computing device;

select, based on the utterance, an agent from a plurality of agents, wherein the plurality of agents includes one or more first party agents and a plurality of third-party agents, the storage medium further comprising instructions that when executed, cause the one or more processors to:

select, in response to determining that the selected agent comprises a first party agent, a reserved voice from a plurality of voices, wherein the reserved voice is associated with the first party agent;

obtain, by the first agent and based on the utterance, a plurality of search results, including a first sub-set of the search results and a second sub-set of the search results; and

output, for playback, to satisfy the utterance:

synthesized audio data, of the first party agent, that represents the first sub-set of the search results using the selected reserved voice, and

synthesized audio data, of the first party agent, that represents the second sub-set of the search results using an additional voice from the plurality of voices, wherein the additional voice is distinct from the selected reserved voice.

15. The non-transitory computer-readable storage medium of claim 14 , wherein the utterance comprises a first utterance, the assistant further configured to:

receive a representation of a second utterance spoken at the computing device;

select, based on the second utterance, a second agent from the plurality of agents;

select, responsive to determining that the selected second agent comprises a third-party agent, a voice from the plurality of voices other than the reserved voice; and

output synthesized audio data using the selected voice to satisfy the second utterance.

Assignments (3)
CHANGE OF NAME Recorded May 24, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050255/0021 →
CHANGE OF NAME Recorded Apr 15, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 050268/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2017
From: NYGAARD, VALERIE; CAPRITA, BOGDAN; STETS, ROBERT; KRISHNAKUMARAN, SAISURESH; DOUGLAS, JASON BRANT
To: GOOGLE INC.
Reel/Frame 044273/0148 →
Continuity (3)
Continuation PCTUS2017054467 · Sep 29, 2017
Provisional Application 62403665 · Oct 3, 2016
Related Publication 20180096675A1 · Apr 5, 2018