IP Library Granted Patent US 11,521,618
Granted Patent B2
US 11,521,618 · App. 16/716,654 · Granted Dec 6, 2022

Collaborative voice controlled devices

Inventors: Victor Carbune (Winterthur, CH); Pedro Gonnet Anders (Zurich, CH); Thomas Deselaers (Zurich, CH); Sandro Feuz (Zurich, CH)
Assignee: GOOGLE LLC
G10L15/30G10L15/22G10L13/033G10L13/08G10L2015/088G10L2015/223G10L2015/228H04W4/80
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,521,618
App. No.
16/716,654
Granted
Dec 6, 2022
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for collaboration between multiple voice controlled devices are disclosed. In one aspect, a method includes the actions of identifying, by a first computing device, a second computing device that is configured to respond to a particular, predefined hotword; receiving audio data that corresponds to an utterance; receiving a transcription of additional audio data outputted by the second computing device in response to the utterance; based on the transcription of the additional audio data and based on the utterance, generating a transcription that corresponds to a response to the additional audio data; and providing, for output, the transcription that corresponds to the response.

Claims (42)

1. A computer-implemented method, comprising:

receiving, by a first computing device associated with a first user, data indicating that a second computing device, associated with a second user, is providing an audible presentation of an audible response to an utterance spoken by the second user; and

in response to receiving the data indicating that the second computing device associated with the second user is providing the audible presentation of the audible response to the utterance spoken by the second user:

generating, by the first computing device associated with the first user, an additional audible response based on a first transcription of the utterance spoken by the second user and a second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user; and

providing, by the first computing device and for audible presentation to the first user, the additional audible response generated based on the first transcription of the utterance spoken by the second user and the second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user.

2. The method of claim 1 , further comprising:

determining, by the first computing device, that the second computing device is co-located with the first computing device,

wherein generating the additional audible response is based further on determining that the second computing device is co-located with the first computing device.

3. The method of claim 1 , wherein the additional audible response is an additional response to the utterance spoken by the second user.

4. The method of claim 1 , wherein the additional audible response is an additional response to the audible response that was provided by the second computing device in response to the utterance spoken by the second user.

5. The method of claim 1 , wherein the additional audible response comprises synthesized speech and is part of a conversation that includes the utterance spoken by the second user.

6. The method of claim 1 , wherein:

receiving the data indicating that the second computing device is audibly presenting the audible response to the utterance spoken by the second user comprises receiving the second transcription of the audible response.

7. The method of claim 1 , wherein the audible response provided for audible presentation in response to the utterance spoken by the second user comprises synthesized speech and is part of a conversation that includes the utterance spoken by the second user.

8. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a first computing device associated with a first user, data indicating that a second computing device, associated with a second user, is providing an audible presentation of an audible response to an utterance spoken by the second user;

in response to receiving the data indicating that the second computing device associated with the second user is providing the audible presentation of the audible response to the utterance spoken by the second user:

generating, by the first computing device associated with the first user, an additional audible response based on a first transcription of the utterance spoken by the second user and a second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user; and

providing, by the first computing device and for audible presentation to the first user, the additional audible response generated based on the first transcription of the utterance spoken by the second user and the second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user.

9. The system of claim 8 , wherein the operations further comprise:

determining, by the first computing device, that the second computing device is co-located with the first computing device,

wherein generating the additional audible response is based further on determining that the second computing device is co-located with the first computing device.

10. The system of claim 8 , wherein the additional audible response is an additional response to the utterance spoken by the second user.

11. The system of claim 8 , wherein the additional audible response is an additional response to the audible response provided for audible presentation by the second computing device in response to the utterance spoken by the second user.

12. The system of claim 8 , wherein the additional audible response comprises synthesized speech and is part of a conversation that includes the utterance spoken by the second user.

13. The system of claim 8 , wherein:

receiving the data indicating that the second computing device is audibly presenting the audible response to the utterance spoken by the second user comprises receiving the second transcription of the audible response.

14. The system of claim 8 , wherein the audible response provided for audible presentation in response to the utterance spoken by the second user comprises synthesized speech and is part of a conversation that includes the utterance spoken by the second user.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, by a first computing device associated with a first user, data indicating that a second computing device, associated with a second user, is providing an audible presentation of an audible response to an utterance spoken by the second user;

in response to receiving the data indicating that the second computing device associated with the second user is providing the audible presentation of the audible response to the utterance spoken by the second user:

generating, by the first computing device associated with the first user, an additional audible response based on a first transcription of the utterance spoken by the second user and a second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user; and

providing, by the first computing device and for audible presentation to the first user, the additional audible response generated based on the first transcription of the utterance spoken by the second user and the second transcription of the audible response that was audibly presented by the second computing device in response to the utterance spoken by the second user.

16. The medium of claim 15 , wherein the operations further comprise:

determining, by the first computing device, that the second computing device is co-located with the first computing device,

wherein generating the additional audible response is based further on determining that the second computing device is co-located with the first computing device.

17. The medium of claim 15 , wherein the additional audible response is an additional response to the utterance spoken by the second user.

18. The medium of claim 15 , wherein the additional audible response is an additional response to the audible response provided for audible presentation to the second user.

19. The medium of claim 15 , wherein the additional audible response comprises synthesized speech and is part of a conversation that includes the utterance spoken by the second user.

20. The medium of claim 15 , wherein:

receiving the data indicating that the second computing device is audibly presenting the audible response to the utterance spoken by the second user comprises receiving the second transcription of the audible response.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2020
From: CARBUNE, VICTOR; ANDERS, PEDRO GONNET; DESELAERS, THOMAS; FEUZ, SANDRO
To: GOOGLE INC.
Reel/Frame 051459/0571 →
ENTITY CONVERSION Recorded Jan 9, 2020
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 051529/0782 →
Continuity (2)
Continuation 15387884 · Dec 22, 2016
Related Publication 20200126563A1 · Apr 23, 2020