IP Library › Granted Patent US 12,254,883
Granted Patent B2
US 12,254,883 · App. 18/635,974 · Granted Mar 18, 2025

Automated calling system

Inventors: Asaf Aharoni (Ramat Hasharon, IL); Arun Narayanan (Milpitas, CA); Nir Shabat (Geva, IL); Parisa Haghani (Jersey City, NJ); Galen Tsai Chuang (New York, NY); Yaniv Leviathan (New York, NY); Neeraj Gaur (Jersey City, NJ); Pedro J. Moreno Mengibar (Jersey City, NJ); Rohit Prakash Prabhavalkar (Santa Clara, CA); Zhongdi Qu (New York, CA); Austin Severn Waters (Brooklyn, NY); Tomer Amiaz (Tel Aviv, IL); Michiel A. U. Bacchiani (Summit, NJ)
Assignee: GOOGLE LLC
G10L15/26G10L15/32H04M1/02H04M1/663H04M3/4286H04M3/5191
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,883
App. No.
18/635,974
Filed
Apr 15, 2024
Granted
Mar 18, 2025
Kind
B2
Art Unit
2693
USPC
704/235
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for an automated calling system are disclosed. In one aspect, a method includes the actions of receiving audio data of an utterance spoken by a user who is having a telephone conversation with a bot. The actions further include determining a context of the telephone conversation. The actions further include determining a user intent of a first previous portion of the telephone conversation spoken by the user and a bot intent of a second previous portion of the telephone conversation outputted by a speech synthesizer of the bot. The actions further include, based on the audio data of the utterance, the context of the telephone conversation, the user intent, and the bot intent, generating synthesized speech of a reply by the bot to the utterance. The actions further include, providing, for output, the synthesized speech.

Claims (62)

1. A method implemented by one or more processors, the method comprising:

receiving audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determining, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determining a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

causing at the least the representation of the utterance received during the ongoing conversation and the context of the ongoing conversation to be processed, using a sequence-to-sequence model, to generate a reply by the bot to the utterance; and

causing synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

2. The method of claim 1 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

3. The method of claim 1 , wherein causing the synthesized speech to be generated comprises:

processing, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

4. The method of claim 1 , wherein the utterance includes a request to perform a task.

5. The method of claim 4 , further comprising:

based on the ongoing conversation:

determining whether the task has been completed; and

in response to determining that the task has been completed:

causing the bot to terminate the conversation.

6. The method of claim 5 , further comprising:

in response to determining that the task has not been completed:

causing the bot to continue the conversation.

7. The method of claim 1 , further comprising:

determining one or more corresponding user intents for the ongoing conversation between the user and the bot,

wherein the one or more corresponding user intents are processed, using the sequence-to-sequence model and along with the representation of the utterance received during the ongoing conversation and the context of the ongoing conversation, to generate the reply by the bot to the utterance.

8. A system comprising:

at least one processor; and

memory storing instructions that, when executed by the at least one processor, cause the at least one processor to:

receive audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determine, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determine a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

cause at the least the representation of the utterance received during the ongoing conversation and the context of the ongoing conversation to be processed, using a sequence-to-sequence model, to generate a reply by the bot to the utterance; and

cause synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

9. The system of claim 8 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

10. The system of claim 9 , wherein the instructions to cause the synthesized speech to be generated comprise instructions to:

process, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

11. The system of claim 8 , wherein the utterance includes a request to perform a task.

12. The system of claim 11 , wherein the instructions further comprise instructions to:

based on the ongoing conversation:

determine whether the task has been completed; and

in response to determining that the task has been completed:

cause the bot to terminate the conversation.

13. The system of claim 12 , wherein the instructions further comprise instructions to:

in response to determining that the task has not been completed:

cause the bot to continue the conversation.

14. The system of claim 8 , wherein the instructions further comprise instructions to:

determine one or more corresponding user intents for the ongoing conversation between the user and the bot,

wherein the one or more corresponding user intents are processed, using the sequence-to-sequence model and along with the representation of the utterance received during the ongoing conversation and the context of the ongoing conversation, to generate the reply by the bot to the utterance.

15. A non-transitory computer-readable storage medium storing instructions that, when executed by at least one processor, cause the at least one processor to perform operations, the operations comprising:

receiving audio data of an utterance spoken by a user during a portion of an ongoing conversation between the user and a bot, the audio data being captured by one or more microphones of a computing device of the user;

determining, based on processing the audio data of the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot, a representation of the utterance received during the portion of the ongoing conversation;

determining a context of the ongoing conversation between the user and the bot, the context of the ongoing conversation between the user and the bot being based on one or more previous portions of the ongoing conversation between the user and the bot, and the one or more previous portions of the ongoing conversation between the user and the bot occurring prior to receiving the utterance spoken by the user during the portion of the ongoing conversation between the user and the bot;

causing at the least the representation of the utterance received during the ongoing conversation and the context of the ongoing conversation to be processed, using a sequence-to-sequence model, to generate a reply by the bot to the utterance; and

causing synthesized speech, that captures the reply by the bot to the utterance, to be provided for audible presentation to the user, the synthesized speech being provided for audible presentation to the user via one or more speakers of a computing device of the user.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the context of the ongoing conversation comprises one or more of: a task associated with the conversation, a time the conversation is initiated, or a location associated with the user.

17. The non-transitory computer-readable storage medium of claim 15 , wherein causing the synthesized speech to be generated comprises:

processing, using a speech synthesizer, the reply by the bot to generate the synthesized speech.

18. The non-transitory computer-readable storage medium of claim 15 , wherein the utterance includes a request to perform a task.

19. The non-transitory computer-readable storage medium of claim 18 , further comprising:

based on the ongoing conversation:

determining whether the task has been completed; and

in response to determining that the task has been completed:

causing the bot to terminate the conversation.

20. The non-transitory computer-readable storage medium of claim 19 , further comprising:

in response to determining that the task has not been completed:

causing the bot to continue the conversation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: AHARONI, ASAF; NARAYANAN, ARUN; SHABAT, NIR; HAGHANI, PARISA; CHUANG, GALEN TSAI; LEVIATHAN, YANIV; GAUR, NEERAJ; MORENO MENGIBAR, PEDRO J.; PRABHAVALKAR, ROHIT PRAKASH; QU, ZHONGDI; WATERS, AUSTIN SEVERN; AMIAZ, TOMER; BACCHIANI, MICHIEL A.U.
To: GOOGLE LLC
Reel/Frame 067362/0195 →
Continuity (5)
Continuation 18219480 · Jul 7, 2023
Continuation 17964141 · Oct 12, 2022
Continuation 17505913 · Oct 20, 2021
Continuation 16580726 · Sep 24, 2019
Related Publication 20240265923A1 · Aug 8, 2024
References Cited (146)
US 5815566A · Ramot et al. · 1998 [cited by applicant]
US 6304653B1 · O'Neil et al. · 2001 [cited by applicant]
US 6377567B1 · Leonard · 2002 [cited by applicant]
US 6731725B1 · Merwin et al. · 2004 [cited by applicant]
US 6922465B1 · Howe · 2005 [cited by applicant]
US 7084758B1 · Cole · 2006 [cited by applicant]
US 7337158B2 · Fratkina et al. · 2008 [cited by applicant]
US 7539656B2 · Fratkina et al. · 2009 [cited by applicant]
US 7792773B2 · McCord et al. · 2010 [cited by applicant]
US 7920678B2 · Cooper et al. · 2011 [cited by applicant]
US 8345835B1 · Or-Bach et al. · 2013 [cited by applicant]
US 8594308B2 · Soundar · 2013 [cited by applicant]
US 8938058B2 · Soundar · 2015 [cited by applicant]
US 8964963B2 · Soundar · 2015 [cited by applicant]
US 9232369B1 · Fujisaki · 2016 [cited by applicant]
US 9318108B2 · Gruber et al. · 2016 [cited by applicant]
US 9467566B2 · Soundar · 2016 [cited by applicant]
US 11158321B2 · Aharoni et al. · 2021 [cited by applicant]
US 11495233B2 · Aharoni · 2022 [cited by applicant]
US 11741966B2 · Aharoni · 2023 [cited by applicant]
US 11990133B2 · Aharoni · 2024 [cited by examiner]
US 20020051522A1 · Merrow et al. · 2002 [cited by applicant]
US 20020055975A1 · Petrovykh · 2002 [cited by applicant]
US 20030009530A1 · Philonenko · 2003 [cited by applicant]
US 20030063732A1 · Mcknight · 2003 [cited by applicant]
US 20040001575A1 · Tang · 2004 [cited by applicant]
US 20040083195A1 · McCord et al. · 2004 [cited by applicant]
US 20040213384A1 · Alles et al. · 2004 [cited by applicant]
US 20040240642A1 · Crandell et al. · 2004 [cited by applicant]
US 20050147227A1 · Chervirala et al. · 2005 [cited by applicant]
US 20050175168A1 · Summe et al. · 2005 [cited by applicant]
US 20050271250A1 · Vallone et al. · 2005 [cited by applicant]
US 20060039365A1 · Ravikumar et al. · 2006 [cited by applicant]
US 20060056600A1 · Merrow et al. · 2006 [cited by applicant]
US 20070036320A1 · Mandalia et al. · 2007 [cited by applicant]
US 20070201664A1 · Salafia · 2007 [cited by applicant]
US 20080209449A1 · Maehira · 2008 [cited by applicant]
US 20080309449A1 · Martin et al. · 2008 [cited by applicant]
US 20090022293A1 · Routt · 2009 [cited by applicant]
US 20090029674A1 · Brezina et al. · 2009 [cited by applicant]
US 20090089096A1 · Schoenberg · 2009 [cited by applicant]
US 20090089100A1 · Nenov et al. · 2009 [cited by applicant]
US 20090137278A1 · Haru et al. · 2009 [cited by applicant]
US 20090232295A1 · Ryskamp · 2009 [cited by applicant]
US 20100088613A1 · DeLuca et al. · 2010 [cited by applicant]
US 20100104087A1 · Byrd et al. · 2010 [cited by applicant]
US 20100228590A1 · Muller et al. · 2010 [cited by applicant]
US 20110092187A1 · Miller · 2011 [cited by applicant]
US 20110270687A1 · Bazaz · 2011 [cited by applicant]
US 20120016678A1 · Gruber et al. · 2012 [cited by applicant]
US 20120109759A1 · Oren et al. · 2012 [cited by applicant]
US 20120147762A1 · Hancock et al. · 2012 [cited by applicant]
US 20120157067A1 · Turner et al. · 2012 [cited by applicant]
US 20120173243A1 · Anand et al. · 2012 [cited by applicant]
US 20120271676A1 · Aravamudan et al. · 2012 [cited by applicant]
US 20130060587A1 · Bayrak et al. · 2013 [cited by applicant]
US 20130077772A1 · Lichorowic et al. · 2013 [cited by applicant]
US 20130090098A1 · Gidwani · 2013 [cited by applicant]
US 20130136248A1 · Kaiser-Nyman et al. · 2013 [cited by applicant]
US 20130163741A1 · Balasaygun et al. · 2013 [cited by applicant]
US 20130275164A1 · Gruber et al. · 2013 [cited by applicant]
US 20140024362A1 · Kang et al. · 2014 [cited by applicant]
US 20140029734A1 · Kim et al. · 2014 [cited by applicant]
US 20140037084A1 · Dutta · 2014 [cited by applicant]
US 20140107476A1 · Tung et al. · 2014 [cited by applicant]
US 20140200928A1 · Watanabe et al. · 2014 [cited by applicant]
US 20140247933A1 · Soundar · 2014 [cited by applicant]
US 20140310365A1 · Sample et al. · 2014 [cited by applicant]
US 20150139413A1 · Hewitt et al. · 2015 [cited by applicant]
US 20150142704A1 · London · 2015 [cited by applicant]
US 20150150019A1 · Sheaffer et al. · 2015 [cited by applicant]
US 20150237203A1 · Siminoff · 2015 [cited by applicant]
US 20150248817A1 · Steir et al. · 2015 [cited by applicant]
US 20150281446A1 · Milstein et al. · 2015 [cited by applicant]
US 20150339707A1 · Harrison et al. · 2015 [cited by applicant]
US 20150350331A1 · Kumar · 2015 [cited by applicant]
US 20150358790A1 · Nasserbakht · 2015 [cited by applicant]
US 20160105546A1 · Keys et al. · 2016 [cited by applicant]
US 20160139998A1 · Dunn et al. · 2016 [cited by applicant]
US 20160198045A1 · Kulkarni et al. · 2016 [cited by applicant]
US 20160227033A1 · Song · 2016 [cited by applicant]
US 20160227034A1 · Kulkarni et al. · 2016 [cited by applicant]
US 20160277569A1 · Shine et al. · 2016 [cited by applicant]
US 20160379230A1 · Chen · 2016 [cited by applicant]
US 20170037084A1 · Fasan · 2017 [cited by applicant]
US 20170039194A1 · Tschetter · 2017 [cited by applicant]
US 20170061091A1 · McElhinney et al. · 2017 [cited by applicant]
US 20170094052A1 · Zhang et al. · 2017 [cited by applicant]
US 20170177298A1 · Hardee et al. · 2017 [cited by applicant]
US 20170180499A1 · Gelfenbeyn et al. · 2017 [cited by applicant]
US 20170289332A1 · Lavian et al. · 2017 [cited by applicant]
US 20170358296A1 · Segalis et al. · 2017 [cited by applicant]
US 20170359463A1 · Segalis et al. · 2017 [cited by applicant]
US 20170359464A1 · Segalis et al. · 2017 [cited by applicant]
US 20170365277A1 · Park · 2017 [cited by applicant]
US 20180124255A1 · Kawamura et al. · 2018 [cited by applicant]
US 20180133900A1 · Breazeal et al. · 2018 [cited by applicant]
US 20180220000A1 · Segalis et al. · 2018 [cited by applicant]
US 20180227416A1 · Segalis et al. · 2018 [cited by applicant]
US 20180227417A1 · Segalis et al. · 2018 [cited by applicant]
US 20180227418A1 · Segalis et al. · 2018 [cited by applicant]
US 20180317064A1 · Kim · 2018 [cited by applicant]
US 20190189123A1 · Sugiyama et al. · 2019 [cited by applicant]
US 20190197396A1 · Rajkumar et al. · 2019 [cited by applicant]
US 20200035244A1 · Kim · 2020 [cited by applicant]
US 20210005191A1 · Chun et al. · 2021 [cited by applicant]
US 20210082410A1 · Teserra · 2021 [cited by applicant]
US 20210090570A1 · Aharoni et al. · 2021 [cited by applicant]
US 20220044684A1 · Aharoni et al. · 2022 [cited by applicant]
US 20230038343A1 · Aharoni et al. · 2023 [cited by applicant]
US 20230352027A1 · Aharoni et al. · 2023 [cited by applicant]
EP 1679693 · 2006 [cited by applicant]
WO 2007065193 · 2007 [cited by applicant]
Abadi et al, “Tensorflow: a system for large-scale machine learning,” 12th USENIX Symposium on Operating Systems Design and Implementation, 21 pages; dated Nov. 2016. [cited by applicant]
Bandanau et al, “Neural machine translation by jointly learning to align and translate,” arXiv; 15 pages; dated May 19, 2016. [cited by applicant]
Bengio et al, “Scheduled sampling for sequence prediction with recurrent neural networks,” arXiv, 9 pages; dated Sep. 23, 2015. [cited by applicant]
Caruana et al, “Multitask learning,” Springer, ; 35 pages; dated 1997. [cited by applicant]
Chan et al, “Listen, attend, and spell: a neural network for large vocabulary conversational speech recognition,” arXiv, 16 pages; dated Aug. 20, 2015. [cited by applicant]
Chen et al, “Spoken language understanding without speech recognition,” in IEEE Xplore, 5 pages; dated 2018. [cited by applicant]
Chiu et al, “State-of-the-art speech recognition with sequence-to-sequence models,” arXiv; 5 pages; dated Feb. 23, 2018. [cited by applicant]
Cho, K. et al, “Learning phrase representations using rnn encoder-decoder for statistical machine translation,” arXiv:1406.1078, 15 pages; dated Sep. 3, 2014. [cited by applicant]
Erik F., “Introduction to the conll-2003 shared tasks: language-independent named entity recognition,” CNTS—Language Technology Group, 4 pages; dated 2003. [cited by applicant]
googleblog.com [online], An AI system for accomplishing real-world tasks over the phone, May 8, 2018, retrieved from: URL<https://ai.googleblog.com/2018/05/duplex-ai-system-for-natural-conversation.html>, 6 pages; retri… [cited by applicant]
Graves et al, “Connectionist temporal classification: labelling unsegmented sequence data with recurrent neural networks,” ACM, 8 pages; dated 2006. [cited by applicant]
Hakkani-Tur, “Beyond asr 1-best: using word confusion networks in spoken language understanding,” Computer Speech & Language, dated 2006. [cited by applicant]
Hakkani-Tur, “Multi-domain joint semantic frame parsing using bi-directional rnn-lstm,” Interspeech, 5 pages; dated Sep. 2016. [cited by applicant]
Jouppi et al, “In-datacenter performance analysis of tensor processing unit,” in 44th International Symposium on Computer Architecture, 17 pages; dated Jun. 26, 2017. [cited by applicant]
Kim et al, “Onenet: joint domain, intent, slot prediction for spoken language understanding,” in Automatic Speech Recognition and Understanding Workshop (ASRU), ArXiv, 7 pages; dated Jan. 16, 2018. [cited by applicant]
Kingma et al, “Adam: A method for stochastic optimization,” arXiv, 15 pages; dated Jan. 30, 2017. [cited by applicant]
Ladhak et al, “LatticeRNN: recurrent neural network over lattices,” Interspeech, 5 pages; dated Sep. 2016. [cited by applicant]
Li et al, “Acoustic modeling for a google home,” Interspeech, 5 pages; dated 2017. [cited by applicant]
Liu et al, “Attention-based recurrent neural network models for joint intent detection and slot filling,” arXiv, 5 pages; dated Sep. 6, 2016. [cited by applicant]
Luong et al, “Multi-task sequence to sequence learning,” arXiv, 10 pages; dated Mar. 1, 2016. [cited by applicant]
Morbini et al, “A reranking approach fo recognition and classification of speech input in conversational dialogue systems,” IEEE Xplore, 6 pages; dated 2012. [cited by applicant]
Prabhavalkar et al, “Minimum word error rate training for attention-based sequence-to-sequence models,” arXiv, 5 pages; dated Dec. 5, 2017. [cited by applicant]
Sak et al, “Fast and accurate recurrent neural network acoustic models for speech recognition;” arXiv, 5 pages Jul. 24, 2015. [cited by applicant]
Saon et al, “English conversational telephone speech recognition by humans and machines,” arXiv, 7 pages; dated Mar. 6, 2017. [cited by applicant]
Schmidhuber et al, “Long short-term memory,” Neural Compute, 9(8): 1735-1780, 32 pages; dated 1997. [cited by applicant]
Schumann et al, “Incorporating asr errors with attention-based, jointly trained rnn for intent detection and slot filling,” Acoustics, Speech, and Signal Processing (ICASSP), 5 pages; dated 2018. [cited by applicant]
Schuster et al, “Bidirectional recurrent neural networks,” IEEE Xplore, 9 pages; dated Nov. 1997. [cited by applicant]
Serdyuk et al, “Towards end-to-end spoken language understanding,” arXiv, 5 pages; dated Feb. 23, 2018. [cited by applicant]
Shannon, “Optimizing expected word error rate via sampling for speech recognition,” arXiv, 5 pages; dated Jun. 8, 2017. [cited by applicant]
Stolcke et al, “Comparing human and machine errors in conversational speech transcription,” Interspeech, ArXiv, 5 pages; dated Aug. 29, 2017. [cited by applicant]
Sutskever et al, “Sequence to sequence learning with neural networks,” in Advances in neural information with neural networks, 9 pages; dated 2014. [cited by applicant]
Vaswani et al, “Attention is all you need,” arXiv, 15 pages; dated Dec. 6, 2017. [cited by applicant]
Wu et al, “Google's neural machine translation system: bridging the gap between human and machine translation,” arXiv, 23 pages; dated Oct. 8, 2016. [cited by applicant]