IP Library Granted Patent US 11,468,893
Granted Patent B2
US 11,468,893 · App. 17/049,689 · Granted Oct 11, 2022

Automated calling system

Inventors: Asaf Aharoni (Ramat Hasharon, IL); Eyal Segalis (Tel Aviv, IL); Yaniv Leviathan (New York, NY)
Assignee: GOOGLE LLC
G10L15/222H04M3/5158H04M2201/39H04M2201/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,468,893
App. No.
17/049,689
Granted
Oct 11, 2022
Kind
B2
Abstract

Methods, systems, and apparatus for an automated calling system are disclosed. Some implementations are directed to using a bot to initiate telephone calls and conduct telephone conversations with a user. The bot may be interrupted while providing synthesized speech during the telephone call. The interruption can be classified into one of multiple disparate interruption types, and the bot can react to the interruption based on the interruption type. Some implementations are directed to determining that a first user is placed on hold by a second user during a telephone conversation, and maintaining the telephone call in an active state in response to determining the first user hung up the telephone call. The first user can be notified when the second user rejoins the call, and a bot associated with the first user can notify the first user that the second user has rejoined the telephone call.

Claims (86)

1. A method implemented by one or more processors, the method comprising:

initiating a telephone call with a user using a bot, the bot configured to initiate telephone calls and conduct telephone conversations;

providing, for output at a corresponding computing device of the user, synthesized speech of the bot; and

while providing the synthesized speech of the bot:

receiving, from the user, a user utterance that interrupts the synthesized speech the bot;

in response to receiving the user utterance that interrupts the synthesized speech, classifying the received user utterance as a given type of interruption, the given type of interruption being one of multiple disparate types of interruptions, the multiple disparate types of interruptions including at least: a non-meaningful interruption, a non-critical meaningful interruption, and a critical meaningful interruption; and

determining, based on the given type of interruption, whether to continue providing, for output at the corresponding computing device of the user, the synthesized speech of the bot that was being provided when the user utterance that interrupts the synthesized speech the bot is received.

2. The method of claim 1 , wherein the given type of interruption is the non-meaningful interruption, and wherein classifying the received user utterance as the non-meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes one or more of: background noise, affirmation words or phrases, or filler words or phrases; and

classifying the received user utterance as the non-meaningful interruption based on determining that the received user utterance includes one or more of: background noise, affirmation words or phrases, or filler words or phrases.

3. The method of claim 2 , wherein determining whether to continue providing the synthesized speech of the bot comprises:

determining to continue providing the synthesized speech of the bot based on classifying the received user utterance as the non-meaningful interruption.

4. The method of claim 1 , wherein the given type of interruption is the non-critical meaningful interruption, and wherein classifying the received user utterance as the non-critical meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes a request for information that is known by the bot, and that is yet to be provided; and

classifying the received user utterance as the non-critical meaningful interruption based on determining that the received user utterance includes the request for the information that is known by the bot, and that is yet to be provided.

5. The method of claim 4 , wherein determining whether to continue providing the synthesized speech of the bot comprises:

based on classifying the user utterance as the non-critical meaningful interruption, determining a temporal point in a remainder portion of the synthesized speech to cease providing, for output, the synthesized speech of the bot;

determining whether the remainder portion of the synthesized speech is responsive to the received utterance; and

in response to determining that the remainder portion is not responsive to the received user utterance:

providing, for output, an additional portion of the synthesized speech that is responsive to the received user utterance, and that is yet to be provided; and

after providing, for output, the additional portion of the synthesized speech, continuing providing, for output, the remainder portion of the synthesized speech of the bot from the temporal point.

6. The method of claim 5 , further comprising:

in response to determining that the remainder portion is responsive to the received user utterance:

continuing providing, for output, the remainder portion of the synthesized speech of the bot from the temporal point.

7. The method of claim 1 , wherein the given type of interruption is the critical meaningful interruption, and wherein classifying the received user utterance as the critical meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes a request for the bot to repeat the synthesized speech or a request to place the bot on hold; and

classifying the received user utterance as the non-critical meaningful interruption based on determining that the received user utterance includes the request for the bot to repeat the synthesized speech or the request to place the bot on hold.

8. The method of claim 7 , wherein determining whether to continue providing the synthesized speech of the bot comprises:

providing, for output, a remainder portion of a current word or term of the synthesized speech of the bot; and

after providing, for output, the remainder portion of the current word or term, cease providing, for output, the synthesized speech of the bot.

9. The method of claim 1 , wherein classifying the received user utterance as the given type of interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance using a machine learning model to determine the given type of interruption.

10. The method of claim 9 , further comprising:

training the machine learning model using a plurality of training instances, wherein each of the training instances include training instance input and corresponding training instance output,

wherein each training instance input includes training audio data corresponding to an interruption utterance or a transcription corresponding to the interruption utterance, and

wherein each corresponding training instance output includes a ground truth label corresponding the type of interruption included in the interruption utterance.

11. The method of claim 9 , wherein processing the audio data corresponding to the received user utterance or the transcription corresponding to the received user utterance using the machine learning model further comprises processing the synthesized speech being output when the user utterance was received along with the audio data or the transcription.

12. The method of claim 1 , wherein classifying the received user utterance as the given type of interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance using one or more rules that match tokens of the received user utterance to one or more terms associated with each of the multiple disparate interruption types.

13. The method of claim 1 , wherein initiating the telephone call with the user using the bot is responsive to receiving user input, from a given user associated with the bot, to initiate the telephone call.

14. The method of claim 13 , wherein the user input to initiate the telephone call includes information points that are to be included in the synthesized speech that is provided for output at the corresponding computing device of the user.

15. A system comprising:

one or more computers; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations, the operations comprising:

initiating a telephone call with a user using a bot, the bot configured to initiate telephone calls and conduct telephone conversations;

providing, for output at a corresponding computing device of the user, synthesized speech of the bot; and

while providing the synthesized speech of the bot:

receiving, from the user, a user utterance that interrupts the synthesized speech the bot;

in response to receiving the user utterance that interrupts the synthesized speech, classifying the received user utterance as a given type of interruption, the given type of interruption being one of multiple disparate types of interruptions, the multiple disparate types of interruptions including at least: a non-meaningful interruption, a non-critical meaningful interruption, and a critical meaningful interruption; and

determining, based on the given type of interruption, whether to continue providing, for output at the corresponding computing device of the user, the synthesized speech of the bot that was being provided when the user utterance that interrupts the synthesized speech the bot is received.

16. The system of claim 15 ,

wherein the given type of interruption is the critical meaningful interruption,

wherein classifying the received user utterance as the critical meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes a request for the bot to repeat the synthesized speech or a request to place the bot on hold; and

classifying the received user utterance as the non-critical meaningful interruption based on determining that the received user utterance includes the request for the bot to repeat the synthesized speech or the request to place the bot on hold; and

wherein determining whether to continue providing the synthesized speech of the bot comprises:

providing, for output, a remainder portion of a current word or term of the synthesized speech of the bot; and

after providing, for output, the remainder portion of the current word or term, cease providing, for output, the synthesized speech of the bot.

17. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations, the operations comprising:

initiating a telephone call with a user using a bot, the bot configured to initiate telephone calls and conduct telephone conversations;

providing, for output at a corresponding computing device of the user, synthesized speech of the bot; and

while providing the synthesized speech of the bot:

receiving, from the user, a user utterance that interrupts the synthesized speech the bot;

in response to receiving the user utterance that interrupts the synthesized speech, classifying the received user utterance as a given type of interruption, the given type of interruption being one of multiple disparate types of interruptions, the multiple disparate types of interruptions including at least: a non-meaningful interruption, a non-critical meaningful interruption, and a critical meaningful interruption; and

determining, based on the given type of interruption, whether to continue providing, for output at the corresponding computing device of the user, the synthesized speech of the bot that was being provided when the user utterance that interrupts the synthesized speech the bot is received.

18. The system of claim 15 ,

wherein the given type of interruption is the non-meaningful interruption,

wherein classifying the received user utterance as the non-meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes one or more of: background noise, affirmation words or phrases, or filler words or phrases; and

classifying the received user utterance as the non-meaningful interruption based on determining that the received user utterance includes one or more of: background noise, affirmation words or phrases, or filler words or phrases; and

wherein determining whether to continue providing the synthesized speech of the bot comprises:

determining to continue providing the synthesized speech of the bot based on classifying the received user utterance as the non-meaningful interruption.

19. The system of claim 15 ,

wherein the given type of interruption is the non-critical meaningful interruption,

wherein classifying the received user utterance as the non-critical meaningful interruption comprises:

processing audio data corresponding to the received user utterance or a transcription corresponding to the received user utterance to determine that the received user utterance includes a request for information that is known by the bot, and that is yet to be provided; and

classifying the received user utterance as the non-critical meaningful interruption based on determining that the received user utterance includes the request for the information that is known by the bot, and that is yet to be provided; and

wherein determining whether to continue providing the synthesized speech of the bot comprises:

based on classifying the user utterance as the non-critical meaningful interruption, determining a temporal point in a remainder portion of the synthesized speech to cease providing, for output, the synthesized speech of the bot;

determining whether the remainder portion of the synthesized speech is responsive to the received utterance; and

in response to determining that the remainder portion is not responsive to the received user utterance:

providing, for output, an additional portion of the synthesized speech that is responsive to the received user utterance, and that is yet to be provided; and

after providing, for output, the additional portion of the synthesized speech, continuing providing, for output, the remainder portion of the synthesized speech of the bot from the temporal point.

20. The system of claim 19 , the operations further comprising:

in response to determining that the remainder portion is responsive to the received user utterance:

continuing providing, for output, the remainder portion of the synthesized speech of the bot from the temporal point.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2020
From: AHARONI, ASAF; SEGALIS, EYAL; LEVIATHAN, YANIV
To: GOOGLE LLC
Reel/Frame 054287/0216 →
Continuity (2)
Provisional Application 62843660 · May 6, 2019
Related Publication 20210335365A1 · Oct 28, 2021
Cited By (1)
US 12,190,064