IP Library › Granted Patent US 11,990,133
Granted Patent B2
US 11,990,133 · App. 18/219,480 · Granted May 21, 2024

Automated calling system

Inventors: Asaf Aharoni (Ramat Hasharon, IL); Arun Narayanan (Milpitas, CA); Nir Shabat (Geva, IL); Parisa Haghani (Jersey City, NJ); Galen Tsai Chuang (New York, NY); Yaniv Leviathan (New York, NY); Neeraj Gaur (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ); Rohit Prakash Prabhavalkar (Santa Clara, CA); Zhongdi Qu (New York, NY); Austin Severn Waters (Brooklyn, NY); Tomer Amiaz (Tel Aviv, IL); Michiel A. U. Bacchiani (Summit, NJ)
Assignee: GOOGLE LLC
G10L15/26G10L15/32H04M1/02H04M1/663H04M3/4286H04M3/5191
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,990,133
App. No.
18/219,480
Granted
May 21, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for an automated calling system are disclosed. In one aspect, a method includes the actions of receiving audio data of an utterance spoken by a user who is having a telephone conversation with a bot. The actions further include determining a context of the telephone conversation. The actions further include determining a user intent of a first previous portion of the telephone conversation spoken by the user and a bot intent of a second previous portion of the telephone conversation outputted by a speech synthesizer of the bot. The actions further include, based on the audio data of the utterance, the context of the telephone conversation, the user intent, and the bot intent, generating synthesized speech of a reply by the bot to the utterance. The actions further include, providing, for output, the synthesized speech.

Claims (59)

1. A method implemented by one or more processors, the method comprising:

receiving audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determining, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determining a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

determining a corresponding user intent for one or more of the previous portions of the ongoing conversation between the user and the bot;

causing, based on processing at the least (i) the representation of the utterance received during the ongoing conversation, (ii) the context of the ongoing conversation, and (iii) the corresponding user intent for one or more of the previous portions of the ongoing conversation, a reply by the bot, to the utterance, to be generated; and

causing synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

2. The method of claim 1 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

3. The method of claim 1 , wherein causing the synthesized speech to be generated comprises:

processing, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

4. The method of claim 1 , wherein the utterance includes a request to perform a task.

5. The method of claim 4 , further comprising:

based on the ongoing conversation:

determining whether the task has been completed; and

in response to determining that the task has been completed:

causing the bot to terminate the conversation.

6. The method of claim 5 , further comprising:

in response to determining that the task has not been completed:

causing the bot to continue the conversation.

7. A system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to:

receive audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determine, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determine a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

determine a corresponding user intent for one or more of the previous portions of the ongoing conversation between the user and the bot;

cause, based on processing at the least (i) the representation of the utterance received during the ongoing conversation, (ii) the context of the ongoing conversation, and (iii) the corresponding user intent for one or more of the previous portions of the ongoing conversation, a reply by the bot, to the utterance, to be generated; and

cause synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

8. The system of claim 7 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

9. The system of claim 8 , wherein the instructions to cause the synthesized speech to be generated comprise instructions to:

process, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

10. The system of claim 7 , wherein the utterance includes a request to perform a task.

11. The system of claim 10 , wherein the instructions further comprise instructions to:

based on the ongoing conversation:

determine whether the task has been completed; and

in response to determining that the task has been completed:

cause the bot to terminate the conversation.

12. The system of claim 11 , wherein the instructions further comprise instructions to:

in response to determining that the task has not been completed:

cause the bot to continue the conversation.

13. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

receiving audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determining, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determining a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

determining a corresponding user intent for one or more of the previous portions of the ongoing conversation between the user and the bot;

causing, based on processing at the least (i) the representation of the utterance received during the ongoing conversation, (ii) the context of the ongoing conversation, and (iii) the corresponding user intent for one or more of the previous portions of the ongoing conversation, a reply by the bot, to the utterance, to be generated; and

causing synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

15. The non-transitory computer-readable storage medium of claim 13 , wherein causing the synthesized speech to be generated comprises:

processing, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

16. The non-transitory computer-readable storage medium of claim 13 , wherein the utterance includes a request to perform a task.

17. The non-transitory computer-readable storage medium of claim 16 , further comprising:

based on the ongoing conversation:

determining whether the task has been completed; and

in response to determining that the task has been completed:

causing the bot to terminate the conversation.

18. The non-transitory computer-readable storage medium of claim 17 , further comprising:

in response to determining that the task has not been completed:

causing the bot to continue the conversation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: AHARONI, ASAF; NARAYANAN, ARUN; SHABAT, NIR; HAGHANI, PARISA; CHUANG, GALEN TSAI; LEVIATHAN, YANIV; GAUR, NEERAJ; MORENO MENGIBAR, PEDRO J.; PRABHAVALKAR, ROHIT PRAKASH; QU, ZHONGDI; WATERS, AUSTIN SEVERN; AMIAZ, TOMER; BACCHIANI, MICHIEL A.U.
To: GOOGLE LLC
Reel/Frame 064297/0321 →
Continuity (4)
Continuation 17964141 · Oct 12, 2022
Continuation 17505913 · Oct 20, 2021
Continuation 16580726 · Sep 24, 2019
Related Publication 20230352027A1 · Nov 2, 2023
Cited By (1)
US 12,254,883