IP Library › Granted Patent US 12,640,154
Granted Patent B2
US 12,640,154 · App. 18/069,884 · Granted May 26, 2026

Stylizing text-to-speech (TTS) voice response for assistant systems

Inventors: Yang Gao (Menlo Park, CA); Weiyi Zheng (Mountain View, CA); Zhaojun Yang (Bellevue, WA); Thilo Wolfgang Koehler (Mountain View, CA); Christian Fuegen (Sunnyvale, CA); Qing He (Sunnyvale, CA)
Assignee: Meta Platforms Technologies, LLC
G10L15/26G10L15/02G10L15/22G10L2015/025
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,640,154
App. No.
18/069,884
Filed
Dec 21, 2022
Granted
May 26, 2026
Kind
B2
Art Unit
2653
USPC
704/231
Abstract

In one embodiment, a method includes receiving a voice input having first audio features at a client system, generating a text response corresponding to the voice input, wherein the text response is associated with style features, generating an output audio waveform of the text response by a text-to-speech model on the client system, wherein the output audio waveform is generated based on the first audio features and the style features, wherein the output audio waveform comprises second audio features, and rendering the output audio waveform at the client system in response to the voice input.

Claims (113)

1 . A method comprising:

receiving, at a head-wearable device, a first voice input from a user, the first voice input having one or more first audio features;

determining a first voice input emotion of the first voice input based on the one or more first audio features;

generating a first text response corresponding to the first voice input;

in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user, generating, by a text-to-speech model on the head-wearable device, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;

in response to the first voice input, presenting, at the head-wearable device, the first output audio waveform;

receiving, at the head-wearable device, a second voice input from the user, the second voice input having one or more second audio features distinct from the first audio features;

determining a background noise level;

determining a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;

generating a second text response corresponding to the second voice input;

in accordance with a determination that the backgrounds noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user;

generating, by the text-to-speech model on the head-wearable device, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and

in response to the second voice input, presenting, at the head-wearable device, the second output audio waveform;

receiving, at the head-wearable device, a third voice input from the user, the third voice input having the one or more second audio features distinct from the first audio features;

determining a second background noise level;

determining the second voice input emotion based on the one or more second audio features and the second background noise level;

generating a third text response corresponding to the third voice input;

in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user;

generating, by the text-to-speech model on the head-wearable device, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and

in response to the third voice input, presenting, at the head-wearable device, the third output audio waveform.

2 . The method of claim 1 , wherein:

the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;

the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and

the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.

3 . The method of claim 1 , further comprising:

before generating the second output audio waveform of the second text response, determining the one or more other audio features based on at least the second voice input emotion and the second voice input.

4 . The method of claim 3 , wherein the one or more other audio features are further based on one or more of:

environmental features associated with the user;

a time of receipt associated with the second voice input; and

one or more user settings associated with the user.

5 . The method of claim 3 , wherein the determining the first voice input emotion, the determining the second voice input emotion, and the determining the one or more other audio features are performed by a machine-learning model.

6 . The method of claim 1 , wherein determining the one or more other audio features based on at least the second voice input emotion and the second voice input is in accordance with the determination, based on the second voice input emotion, that the one or more other audio features are more appropriate than the one or more second audio features for responding to the user.

7 . The method of claim 1 , wherein:

the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion; and

the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and

the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.

8 . The method of claim 1 , the method further comprising:

after receiving the first voice input, automatically transcribing the first voice input into a first input text, wherein generating the first text response corresponding to the first voice input is further based on the first input text;

after receiving the second voice input, automatically transcribing the second voice input into a second input text, wherein generating the second text response corresponding to the second voice input is further based on the second input text; and

after receiving the third voice input, automatically transcribing the third voice input into a third input text, wherein generating the third text response corresponding to the third voice input is further based on the third input text.

9 . The method of claim 1 , wherein the text-to-speech model is based on a prosody model and an acoustic model.

10 . The method of claim 9 , wherein the prosody model determines:

a first duration and a first frequency of the first output audio waveform;

a second duration and a second frequency of the second output audio waveform; and

a third duration and a third frequency of the third output audio waveform.

11 . The method of claim 10 , wherein generating the second output audio waveform of the second text response comprises:

accessing, by the prosody model, a style embedding corresponding to the one or more other audio features; and

setting one or more of the second duration and the second frequency of the second output audio waveform based on the style embedding to generate response prosody features of the second output audio waveform.

12 . The method of claim 1 , wherein:

the head-wearable device includes one or more speakers;

presenting the first output audio waveform includes playing the first output audio waveform at the one or more speakers;

presenting the second output audio waveform includes playing the second output audio waveform at the one or more speakers; and

presenting the third output audio waveform includes playing the third output audio waveform at the one or more speakers.

13 . A system comprising one or more processors and a memory coupled to the one or more processors, the memory including executable instructions that, when executed by the one or more processors, cause the one or more processors to:

receive a first voice input from a user of a head-wearable device, the first voice input having one or more first audio features;

determine a first voice input emotion of the first voice input based on the one or more first audio features;

generate a first text response corresponding to the first voice input;

in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user of the head-wearable device, generate, by a text-to-speech model, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;

in response to the first voice input, cause the first output audio waveform to be presented to the user of the head-wearable device;

receive a second voice input from the user of the head-wearable device, the second voice input having one or more second audio features distinct from the first audio features;

determine a background noise level;

determine a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;

generate a second text response corresponding to the second voice input;

in accordance with a determination that the background noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user of the head-wearable device:

generate, by the text-to-speech model, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and

in response to the second voice input, cause the second output audio waveform to be presented to the user of the head-wearable device:

receive a third voice input from the user of the head-wearable device, the third voice input having the one or more second audio features distinct from the first audio features;

determine a second background noise level;

determine the second voice input emotion based on the one or more second audio features and the second background noise level;

generate a third text response corresponding to the third voice input;

in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user of the head-wearable device:

generate, by the text-to-speech model, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and

in response to the third voice input, cause the third output audio waveform to be presented to the user of the head-wearable device.

14 . The system of claim 13 , wherein:

the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;

the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and

the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.

15 . The system of claim 13 , wherein the executable instructions further cause the one or more processors to:

before generating the second output audio waveform of the second text response, determine the one or more other audio features based on at least the second voice input emotion and the second voice input.

16 . The system of claim 13 , wherein:

the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion;

the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and

the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.

17 . A computer-readable non-transitory storage medium including executable instructions that, when executed by one or more processors, cause the one or more processors to:

receive a first voice input from a user of a head-wearable device, the first voice input having one or more first audio features;

determine a first voice input emotion of the first voice input based on the one or more first audio features;

generate a first text response corresponding to the first voice input;

in accordance with a determination, based on the first voice input emotion, that the one or more first audio features are appropriate for responding to the user of the head-wearable device, generate, by a text-to-speech model, a first output audio waveform of the first text response, wherein the first output audio waveform has the one or more first audio features;

in response to the first voice input, cause the first output audio waveform to be presented to the user of the head-wearable device;

receive a second voice input from the user of the head-wearable device, the second voice input having one or more second audio features distinct from the first audio features;

determine a background noise level;

determine a second voice input emotion, distinct from the first voice input emotion, of the second voice input based on the one or more second audio features and the background noise level;

generate a second text response corresponding to the second voice input;

in accordance with a determination that the background noise level does not exceed a background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are not appropriate for responding to the user of the head-wearable device;

generate, by the text-to-speech model, a second output audio waveform of the second text response, wherein the second output audio waveform has one or more other audio features, distinct from the one or more second audio features; and

in response to the second voice input, cause the second output audio waveform to be presented to the user of the head-wearable device;

receive a third voice input from the user of the head-wearable device, the third voice input having the one or more second audio features distinct from the first audio features;

determine a second background noise level;

determine the second voice input emotion based on the one or more second audio features and the second background noise level;

generate a third text response corresponding to the third voice input;

in accordance with a determination that the second background noise level exceeds the background noise threshold and a determination, based on the second voice input emotion, that the one or more second audio features are appropriate for responding to the user of the head-wearable device:

generate, by the text-to-speech model, a third output audio waveform of the third text response, wherein the third output audio waveform has the one or more second audio features; and

in response to the third voice input, cause the third output audio waveform to be presented to the user of the head-wearable device.

18 . The computer-readable non-transitory storage medium of claim 17 , wherein:

the one or more first audio features include one or more of a first tone, a first speed, a first volume, a first emphasis, a first pronunciation, a first accent, a first frequency, and a first prosody;

the one or more second audio features include one or more of a second tone, a second speed, a second volume, a second emphasis, a second pronunciation, a second accent, a second frequency, and a second prosody; and

the one or more other audio features include one or more of another tone, another speed, another volume, another emphasis, another pronunciation, another accent, another frequency, and another prosody.

19 . The computer-readable non-transitory storage medium of claim 17 , wherein the executable instructions further cause the one or more processors to:

before generating the second output audio waveform of the second text response, determine the one or more other audio features based on at least the second voice input emotion and the second voice input.

20 . The computer-readable non-transitory storage medium of claim 17 , wherein:

the determination that the one or more first audio features are appropriate for responding to the user includes a determination that the first voice input emotion is one of a positive emotion and a neutral emotion;

the determination that the one or more second audio features are not appropriate for responding to the user includes a determination that the second voice input emotion is a negative emotion; and

the determination that the one or more second audio features are appropriate for responding to the user includes a determination that the second voice input emotion is the negative emotion.

Continuity (2)
Continuation 16790497 · Feb 13, 2020
Related Publication 20230118412A1 · Apr 20, 2023
References Cited (319)
US 7024348B1 · Scholz et al. · 2006 [cited by applicant]
US 7124123B1 · Roskind et al. · 2006 [cited by applicant]
US 7158678B2 · Nagel et al. · 2007 [cited by applicant]
US 7397912B2 · Aasman et al. · 2008 [cited by applicant]
US 7415413B2 · Eide et al. · 2008 [cited by applicant]
US 8027451B2 · Arendsen et al. · 2011 [cited by applicant]
US 8560564B1 · Hoelzle et al. · 2013 [cited by applicant]
US 8677377B2 · Cheyer et al. · 2014 [cited by applicant]
US 8935192B1 · Ventilla et al. · 2015 [cited by applicant]
US 8983383B1 · Haskin · 2015 [cited by applicant]
US 9154739B1 · Nicolaou et al. · 2015 [cited by applicant]
US 9299059B1 · Marra et al. · 2016 [cited by applicant]
US 9304736B1 · Whiteley et al. · 2016 [cited by applicant]
US 9338242B1 · Suchland et al. · 2016 [cited by applicant]
US 9338493B2 · Van Os et al. · 2016 [cited by applicant]
US 9390724B2 · List · 2016 [cited by applicant]
US 9418658B1 · David et al. · 2016 [cited by applicant]
US 9472206B2 · Ady · 2016 [cited by applicant]
US 9479931B2 · Ortiz, Jr. et al. · 2016 [cited by applicant]
US 9576574B2 · Van Os · 2017 [cited by applicant]
US 9659577B1 · Langhammer · 2017 [cited by applicant]
US 9747895B1 · Jansche et al. · 2017 [cited by applicant]
US 9792281B2 · Sarikaya · 2017 [cited by applicant]
US 9830924B1 · Degges, Jr. · 2017 [cited by examiner]
US 9858925B2 · Gruber et al. · 2018 [cited by applicant]
US 9865260B1 · Vuskovic et al. · 2018 [cited by applicant]
US 9875233B1 · Tomkins et al. · 2018 [cited by applicant]
US 9875741B2 · Gelfenbeyn et al. · 2018 [cited by applicant]
US 9886953B2 · Lemay et al. · 2018 [cited by applicant]
US 9990591B2 · Gelfenbeyn et al. · 2018 [cited by applicant]
US 10042032B2 · Scott et al. · 2018 [cited by applicant]
US 10134395B2 · Typrin · 2018 [cited by applicant]
US 10199051B2 · Binder et al. · 2019 [cited by applicant]
US 10241752B2 · Lemay et al. · 2019 [cited by applicant]
US 10276170B2 · Gruber et al. · 2019 [cited by applicant]
US 10382866B2 · Min · 2019 [cited by examiner]
US 10706837B1 · Chicote et al. · 2020 [cited by applicant]
US 10891947B1 · Le Chevalier · 2021 [cited by examiner]
US 20020184030A1 · Brittan · 2002 [cited by examiner]
US 20020184031A1 · Brittan · 2002 [cited by examiner]
US 20050141997A1 · Rast · 2005 [cited by applicant]
US 20050169440A1 · Agapi et al. · 2005 [cited by applicant]
US 20070112570A1 · Kaneyasu · 2007 [cited by applicant]
US 20080240379A1 · Maislos et al. · 2008 [cited by applicant]
US 20100161327A1 · Chandra · 2010 [cited by examiner]
US 20120246191A1 · Xiong · 2012 [cited by applicant]
US 20120265528A1 · Gruber et al. · 2012 [cited by applicant]
US 20130185066A1 · Tzirkel-Hancock · 2013 [cited by examiner]
US 20130268839A1 · Lefebvre et al. · 2013 [cited by applicant]
US 20130275138A1 · Gruber et al. · 2013 [cited by applicant]
US 20130275164A1 · Gruber et al. · 2013 [cited by applicant]
US 20140052446A1 · Mori et al. · 2014 [cited by applicant]
US 20140164506A1 · Tesch et al. · 2014 [cited by applicant]
US 20150179168A1 · Hakkani-Tur et al. · 2015 [cited by applicant]
US 20150186359A1 · Fructuoso · 2015 [cited by examiner]
US 20160092160A1 · Graff et al. · 2016 [cited by applicant]
US 20160093289A1 · Pollet · 2016 [cited by applicant]
US 20160225370A1 · Kannan et al. · 2016 [cited by applicant]
US 20160255082A1 · Rathod · 2016 [cited by applicant]
US 20160328096A1 · Tran et al. · 2016 [cited by applicant]
US 20160343368A1 · Yassa et al. · 2016 [cited by applicant]
US 20160378849A1 · Myslinski · 2016 [cited by applicant]
US 20160378861A1 · Eledath et al. · 2016 [cited by applicant]
US 20170091168A1 · Bellegarda et al. · 2017 [cited by applicant]
US 20170132019A1 · Karashchuk et al. · 2017 [cited by applicant]
US 20170353469A1 · Selekman et al. · 2017 [cited by applicant]
US 20170359707A1 · Diaconu et al. · 2017 [cited by applicant]
US 20170365277A1 · Park · 2017 [cited by examiner]
US 20180018562A1 · Jung · 2018 [cited by applicant]
US 20180018987A1 · Zass · 2018 [cited by applicant]
US 20180061400A1 · Carbune et al. · 2018 [cited by applicant]
US 20180096071A1 · Green · 2018 [cited by applicant]
US 20180096072A1 · He et al. · 2018 [cited by applicant]
US 20180107917A1 · Hewavitharana et al. · 2018 [cited by applicant]
US 20180122372A1 · Wanderlust · 2018 [cited by examiner]
US 20180144761A1 · Amini · 2018 [cited by examiner]
US 20180184203A1 · Piechowiak · 2018 [cited by examiner]
US 20180189629A1 · Yatziv et al. · 2018 [cited by applicant]
US 20180254034A1 · Li · 2018 [cited by applicant]
US 20180277145A1 · Yamaya · 2018 [cited by examiner]
US 20190005021A1 · Miller et al. · 2019 [cited by applicant]
US 20190080698A1 · Miller · 2019 [cited by applicant]
US 20190147849A1 · Talwar et al. · 2019 [cited by applicant]
US 20190172443A1 · Shechtman · 2019 [cited by examiner]
US 20190172454A1 · Kitajima · 2019 [cited by examiner]
US 20190279629A1 · Okamoto · 2019 [cited by examiner]
US 20190287512A1 · Zoller · 2019 [cited by examiner]
US 20190295533A1 · Wang · 2019 [cited by examiner]
US 20190311718A1 · Huber · 2019 [cited by examiner]
US 20190348020A1 · Clark et al. · 2019 [cited by applicant]
US 20190355354A1 · Geng · 2019 [cited by examiner]
US 20190392851A1 · Kim · 2019 [cited by examiner]
US 20200005763A1 · Chae · 2020 [cited by examiner]
US 20200005764A1 · Chae · 2020 [cited by applicant]
US 20200041803A1 · Reyes · 2020 [cited by examiner]
US 20200066251A1 · Kumano et al. · 2020 [cited by applicant]
US 20200075036A1 · Shin · 2020 [cited by examiner]
US 20200098358A1 · Rakshit · 2020 [cited by examiner]
US 20200152194A1 · Jeong · 2020 [cited by examiner]
US 20200152197A1 · Penilla · 2020 [cited by examiner]
US 20200160833A1 · Nakagawa · 2020 [cited by examiner]
US 20200184967A1 · Gupta et al. · 2020 [cited by applicant]
US 20200233958A1 · Yao et al. · 2020 [cited by applicant]
US 20200234693A1 · Sung · 2020 [cited by examiner]
US 20200294502A1 · Kurihara et al. · 2020 [cited by applicant]
US 20200302927A1 · Andruszkiewicz et al. · 2020 [cited by applicant]
US 20200302946A1 · Jhawar · 2020 [cited by examiner]
US 20200314577A1 · Asfaw · 2020 [cited by examiner]
US 20200342852A1 · Kim et al. · 2020 [cited by applicant]
US 20200402497A1 · Semenov et al. · 2020 [cited by applicant]
US 20210035551A1 · Stanton et al. · 2021 [cited by applicant]
US 20210049996A1 · Chae · 2021 [cited by applicant]
US 20210049997A1 · Yun · 2021 [cited by examiner]
US 20210056972A1 · Dube · 2021 [cited by examiner]
US 20210074261A1 · Yang et al. · 2021 [cited by applicant]
US 20210082304A1 · Daley · 2021 [cited by examiner]
US 20210082427A1 · Fujita · 2021 [cited by examiner]
US 20210090551A1 · Jang et al. · 2021 [cited by applicant]
US 20210104220A1 · Mennicken · 2021 [cited by examiner]
US 20210117553A1 · Shpurov et al. · 2021 [cited by applicant]
US 20210117712A1 · Huang et al. · 2021 [cited by applicant]
US 20210125610A1 · Cheung · 2021 [cited by examiner]
US 20210142796A1 · Saito · 2021 [cited by applicant]
US 20210151029A1 · Gururani et al. · 2021 [cited by applicant]
US 20210192332A1 · Gangotri · 2021 [cited by examiner]
US 20210225357A1 · Zhao et al. · 2021 [cited by applicant]
US 20210329389A1 · Pandey · 2021 [cited by examiner]
US 20220215184A1 · Freitag · 2022 [cited by examiner]
US 20220229524A1 · Mckenzie et al. · 2022 [cited by applicant]
AU 2017203668A1 · 2018 [cited by applicant]
CN 111028827A · 2020 [cited by examiner]
EP 2530870A1 · 2012 [cited by applicant]
EP 3122001A1 · 2017 [cited by applicant]
JP 2003271194A · 2003 [cited by applicant]
RU 2338248C1 · 2008 [cited by applicant]
WO WO2011055830A1 · 2011 [cited by examiner]
WO 2012116241A2 · 2012 [cited by applicant]
WO 2016195739A1 · 2016 [cited by applicant]
WO 2017053208A1 · 2017 [cited by applicant]
WO 2017116488A1 · 2017 [cited by applicant]
WO WO2020143652A1 · 2020 [cited by examiner]
WO WO2020230926A1 · 2020 [cited by examiner]
WO 2021003810A1 · 2021 [cited by applicant]
Tachibana, Makoto, et al. “Speech synthesis with various emotional expressions and speaking styles by style interpolation and morphing.” IEICE transactions on information and systems 88.11 (2005): 2484-2491. (Year: 2005… [cited by examiner]
Vaswani A., et al., “Attention Is All You Need,” Advances in Neural Information Processing Systems, December (NIPS), Dec. 4, 2017, pp. 5998-6008. [cited by applicant]
Vries H.D., et al., “Talk The Walk: Navigating New York City Through Grounded Dialogue,” arXiv prepnnt arXiv: 1807.03367, Dec. 23, 2018, 23 Pages. [cited by applicant]
Wang P., et al., “FVQA: Fact-Based Visual Question Answering,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Aug. 8, 2017, pp. 2413-2427. [cited by applicant]
Wang Y., et al., “Style Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis,” ICML 2018, Mar. 23, 2018, 11 Pages. [cited by applicant]
Wang Y., et al., “Uncovering Latent Style Factors for Expressive Speech Synthesis,” arXiv pre print arXiv: 1711.00520, Nov. 1, 2017, 5 Pages. [cited by applicant]
Wang Z., et al., “Knowledge Graph Embedding by Translating on Hyperplanes,” 28th AAAI Conference on Artificial Intelligence, Jun. 21, 2014, 8 Pages. [cited by applicant]
Wei W., et al., “Airdialogue: An Environment for Goal-Oriented Dialogue Research,” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Nov. 4, 2018, 11 Pages. [cited by applicant]
Welbl J., et al., “Constructing Datasets for Multi-Hop Reading Comprehension Across Documents,” Transactions of the Association for Computational Linguistics, Jun. 11, 2018, vol. 6, pp. 287-302. [cited by applicant]
Weston J., et al., “Memory Networks,” arXiv preprint arXiv: 1410.3916, Nov. 29, 2015, 15 Pages. [cited by applicant]
“What is Conversational AI?,” Glia Blog [Online], Sep. 13, 2017 [Retrieved on Jun. 14, 2019], 3 Pages, Retrieved from Internet: URL: https://blog.salemove.com/what-is-conversational-ai/. [cited by applicant]
Wikipedia, “Dialog Manager,” Dec. 16, 2006-Mar. 13, 2018 [Retrieved on Jun. 14, 2019], 8 Pages, Retrieved from Internet: https://en.wikipedia.org/wiki/Dialog_manager. [cited by applicant]
Wikipedia, “Natural-Language Understanding,” Nov. 11, 2018-Apr. 3, 2019 [Retrieved on Jun. 14, 2019], 5 Pages, Retrieved from Internet: URL: https://en.wikipedia.org/wiki/Natural-language_understanding. [cited by applicant]
Wikipedia, “Speech Synthesis,” Jan. 24, 2004-Jun. 5, 2019 [Retrieved on Jun. 14, 2019], 19 Pages, Retrieved from Internet: URL: https://en.wikipedia.org/wiki/Speech_synthesis. [cited by applicant]
Wikipedia, “Question Answering,” Feb. 15, 2018 [Retrieved on Jun. 14, 2019], 6 Pages, Retrieved from Internet: URL: https://en.wikipedia.org/wiki/Question_answering. [cited by applicant]
Williams J.D., et al., “Hybrid Code Networks: Practical and Efficient End-to-End Dialog Control with Supervised and Reinforcement Learning,” arXiv preprint arXiv: 1702.03274, Apr. 24, 2017, 13 Pages. [cited by applicant]
Wu Q., et al., “Image Captioning and Visual Question Answering Based on Attributes and External Knowledge,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Dec. 16, 2016, vol. 40 (6), 14 Pages. [cited by applicant]
Xu K., et al., “Question Answering on Freebase via Relation Extraction and Textual Evidence,” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics, Jun. 9, 2016, 11 Pages. [cited by applicant]
Yamagishi J., et al., “Modeling of Various Speaking Styles and Emotions for HMM-based Speech Synthesis,” 8th European Conference on Speech Communication and Technology, Jan. 2003, 4 pages. [cited by applicant]
Yamagishi J., et al., “Speaking Style Adaptation using Context Clustering Decision Tree for HMM-based Speech Synthesis,” 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, May 17, 2004, vol.… [cited by applicant]
Yang Z., et al., “Hierarchical Attention Networks for Document Classification,” Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics, Jun. 12-17, 2016, 10 Pag… [cited by applicant]
Yang Z., et al., “HOTPOTQA: A Dataset for Diverse, Explainable Multi-Hop Question Answering,” Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Sep. 25, 2018, 12 Pages. [cited by applicant]
Yin W., et al., “Simple Question Answering by Attentive Convolutional Neural Network,” arXiv preprint arXiv: 1606.03391, Oct. 11, 2016, 11 Pages. [cited by applicant]
Yoon S., et al., “Multimodal Speech Emotion Recognition Using Audio and Text,” 2018 IEEE Spoken Language Technology Workshop (SET), Dec. 18, 2018, 7 Pages. [cited by applicant]
Young T., et al., “Augmenting End-to-End Dialogue Systems with Commonsense Knowledge,” The Thirty-Second AAAI Conference on Artificial Intelligence, Apr. 26, 2018, 8 Pages. [cited by applicant]
Zhang S., et al., “Personalizing Dialogue Agents: I Have a Dog, Do You Have Pets Too?,” Facebook AI Research, Sep. 25, 2018, 16 Pages. [cited by applicant]
Jiang L., et al., “MemexQA: Visual Memex Question Answering,” arXiv preprint arXiv: 1708.01336, Aug. 4, 2017, 10 Pages. [cited by applicant]
Jung H., et al., “Learning What to Remember: Long-term Episodic Memory Networks for Learning from Streaming Data,” arXiv preprint arXiv: 1812.04227, Dec. 11, 2018, 10 Pages. [cited by applicant]
Kartsaklis D., et al., “Mapping Text to Knowledge Graph Entities using Multi-Sense LSTMs,” Department of Theoretical and Applied Linguistics, Aug. 23, 2018, 12 Pages. [cited by applicant]
King S., “The Blizzard Challenge 2017,” In Proceedings Blizzard Challenge, Aug. 2017, pp. 1-17. [cited by applicant]
Kottur S., et al., “Visual Coreference Resolution in Visual Dialog using Neural Module Networks,” In Proceedings of the European Conference on Computer Vision (ECCV), Sep. 6, 2018, pp. 153-169. [cited by applicant]
Kumar A., et al., “Ask Me Anything: Dynamic Memory Networks for Natural Language Processing,” International Conference on Machine Learning, Jan. 6, 2016, 10 Pages. [cited by applicant]
Lao N., et al., “Random Walk Inference and Learning in a Large Scale Knowledge Base,” Proceedings of the Conference on Empirical Methods in Natural Language Processing Association for Computational Linguistics, Jul. 27,… [cited by applicant]
Lee C.C., et al., “Emotion Recognition using a Hierarchical Binary Decision Tree Approach,” Speech Communication, Nov. 1, 2011, vol. 53, pp. 1162-1171. [cited by applicant]
Li J., et al., “A Persona-based Neural Conversation Model,” arXiv preprint arXiv: 1603.06155, Jun. 8, 2016, 10 Pages. [cited by applicant]
Li Y., et al., “Adaptive Batch Normalization for Practical Domain Adaptation,” Pattern Recognition, Aug. 2018, vol. 80, pp. 109-117. [cited by applicant]
Locatello F., et al., “Challenging Common Assumptions in the Unsupervised Learning of Disentangled Representations,” ICLR 2019, Nov. 29, 2018, 37 Pages. [cited by applicant]
Long Y., et al., “A Knowledge Enhanced Generative Conversational Service Agent,” DSTC6 Workshop, Dec. 2017, 6 Pages. [cited by applicant]
Mccross T., “Dialog Management,” Feb. 15, 2018 [Retrieved on Jun. 14, 2019], 12 Pages, Retrieved from Internet: https://tutorials.botsfloor.com/dialog-management-799c20a39aad. [cited by applicant]
Miller A.H., et al., “PARLAI: A Dialog Research Software Platform,” Facebook AI Research, May 18, 2017, 7 Pages. [cited by applicant]
Moon S., et al., “Completely Heterogeneous Transfer Learning with Attention—What and What Not to Transfer,” Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI, Aug. 19, 2017… [cited by applicant]
Moon S., et al., “Multimodal Named Entity Recognition for Short Social Media Posts,” Feb. 2018, 9 Pages. [cited by applicant]
Moon S., et al., “Multimodal Transfer Deep Learning with Applications in Audio-Visual Recognition,” Dec. 9, 2014, 6 Pages. [cited by applicant]
Moon S., et al., “Zeroshot Multimodal Named Entity Disambiguation for Noisy Social Media Posts,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 15-20, 2018, vol. 1, 9 Pages. [cited by applicant]
Mower E., et al., “A Framework for Automatic Human Emotion Classification using Emotion Profiles,” IEEE Transactions on Audio, Speech and Language Processing, Jul. 2011, vol. 19.5, pp. 1057-1070. [cited by applicant]
Nickel M., et al., “Holographic Embeddings of Knowledge Graphs,” Proceedings of Thirtieth AAAI Conference on Artificial Intelligence, Mar. 2, 2016, 7 Pages. [cited by applicant]
Oord A.V.D., et al., “WaveNet: A Generative Model for Raw Audio,” arXiv preprint, arXiv: 1609.03499, Sep. 19, 2016, pp. 1-15. [cited by applicant]
Ostendorf M., et al., “Human Language Technology: Opportunities and Challenges,” IEEE International Conference on Acoustics Speech and Signal Processing, Mar. 23, 2005, 5 Pages. [cited by applicant]
Ott M., et al., “New Advances in Natural Language Processing to Better Connect People,” [online], Aug. 14, 2019, 12 Pages, Retrieved from the Internet: URL: https://ai.facebook.com/blog/new-advances-in-natural-language-… [cited by applicant]
“Overview of Language Technology,” Feb. 15, 2018, 1 Page, Retrieved from Internet: URL: https://www.dfki.de/lt/lt-general.php. [cited by applicant]
Parthasarathi P., et al., “Extending Neural Generative Conversational Model using External Knowledge Sources,” Computational and Language, Sep. 14, 2018, 6 Pages. [cited by applicant]
Pennington J., et al., “GLOVE: Global Vectors for Word Representation,” Proceedings of the 2014 Conference on Empirical Methods in Natural Language Processing (EMNLP), Oct. 24-29, 2014, 12 Pages. [cited by applicant]
Poliak A., et al., “Efficient, Compositional, Order-Sensitive n-gram Embeddings,” Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics, Apr. 3-7, 2017, vol. 2, pp. … [cited by applicant]
Rajpurkar P., et al., “Know What You Don't Know: Unanswerable Questions for SQuAD,” arXiv preprint arXiv: 1806.03822, Jun. 11, 2018, 9 Pages. [cited by applicant]
Rajpurkar P., et al., “SQuAD: 100,000+ Questions for Machine Comprehension of Text,” arXiv preprint arXiv: 1606.05250, Oct. 11, 2016, 10 Pages. [cited by applicant]
Reddy S., et al., “COQA: A Conversational Question Answering Challenge,” Transactions of the Association for Computational Linguistics, May 29, 2019, vol. 7, pp. 249-266. [cited by applicant]
Salem Y., et al., “History-Guided Conversational Recommendation,” Proceedings of the 23rd International Conference on World Wide Web, ACM, Apr. 7-11, 2014, 6 Pages. [cited by applicant]
Scherer K.R., et al., “Vocal Cues in Emotion Encoding and Decoding,” Motivation and emotion, Jun. 1, 1991, vol. 15 (2), 27 Pages. [cited by applicant]
Seo M., et al., “Bidirectional Attention Flow for Machine Comprehension,” arXiv preprint arXiv:1611.01603, Jun. 21, 2018, 13 Pages. [cited by applicant]
Shah P., et al., “Bootstrapping a Neural Conversational Agent with Dialogue Selfplay Crowdsourcing and On-line Reinforcement Learning,” Proceedings of the 2018 Conference of the North American Chapter of the Association… [cited by applicant]
Shen J., et al., “Natural TTS Synthesis by Conditioning Wavenet on Mel Spectrogram Predictions,” 2018 IEEE International Conference on Acoustics Speech and Signal Processing, Apr. 15, 2018, 5 Pages. [cited by applicant]
Skerry-Ryan R.J., et al., “Towards End-to-End Prosody Transfer for Expressive Speech Synthesis with Tacotron,” Proceeding s of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, Mar. 2018… [cited by applicant]
Sukhbaatar S., et al., “End-to-End Memory Networks,” Advances in Neural Information Processing Systems, Nov. 24, 2015, 9 Pages. [cited by applicant]
Sun Y., et al., “Conversational Recommender System,” The 41st International ACM SIGIR Conference on Research & Development in Information Retrieval, ACM, Jul. 8-12, 2018, 10 Pages. [cited by applicant]
Sutskever I., et al., “Sequence to Sequence Learning with Neural Networks,” Advances in Neural Information Processing Systems, Sep. 10, 2014, 9 Pages. [cited by applicant]
Tachibana M., et al., “HMM-based Speech Synthesis with Various Speaking Styles using Model Interpolation,” Speech Prosody, Mar. 2004, 4 pages. [cited by applicant]
Tits N., et al., “Visualization and Interpretation of Latent Spaces for Controlling Expressive Speech Synthesis Through Audio Analysis,” Interspeech 2019, Mar. 27, 2019, 5 Pages. [cited by applicant]
Tran K., et al., “Recurrent Memory Networks for Language Modeling,” arXiv preprint arXiv: 1601.01272, Apr. 22, 2016, 11 Pages. [cited by applicant]
Trigeorgis G., et al., “Adieu Features? End-to-End Speech Emotion Recognition Using a Deep Convolutional Recurrent Network,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Mar. 2… [cited by applicant]
Turniski F., et al., “Analysis of 3G and 4G Download Throughput in Pedestrian Zones,” Proceedings Elmar—2013, Sep. 12, 2016, pp. 9-12, XP032993723. [cited by applicant]
U.S. Appl. No. 62/660,876, inventors Kumar; Anuj et al., filed Apr. 20, 2018. [cited by applicant]
U.S. Appl. No. 62/675,090, inventors Hanson; Michael Robert et al., filed May 22, 2018. [cited by applicant]
U.S. Appl. No. 62/747,628, inventors Liu; Honglei et al., filed Oct. 18, 2018. [cited by applicant]
U.S. Appl. No. 62/749,608, inventors Challa; Ashwini et al., filed Oct. 23, 2018. [cited by applicant]
U.S. Appl. No. 62/750,746, inventors Liu; Honglei et al., filed Oct. 25, 2018. [cited by applicant]
U.S. Appl. No. 62/923,342, inventors Hanson; Michael Robert et al., filed Oct. 18, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/182,542, inventors Michael; Hanson et al., filed Nov. 6, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/183,650, inventors Xiaohu; Liu et al., filed Nov. 7, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/192,538, inventor Koukoumidis; Emmanouil, filed Nov. 15, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/222,923, inventors Jason; Schissel et al., filed Dec. 17, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/222,957, inventors Emmanouil; Koukoumidis et al., filed Dec. 17, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/229,828, inventors Xiaohu; Liu et al., filed Dec. 21, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/247,439, inventors Xiaohu; Liu et al., filed Jan. 14, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/264,173, inventors Ashwini; Challa et al., filed Jan. 31, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/376,832, inventors Liu; Honglei et al., filed Apr. 5, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/388,130, inventors Xiaohu; Liu et al., filed Apr. 18, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/389,634, inventors Crook; Paul Anthony et al., filed Apr. 19, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/389,708, inventors William; Crosby Presant et al., filed Apr. 19, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/389,728, inventors William; Presant et al., filed Apr. 19, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/389,738, inventors Peng; Fuchun et al., filed Apr. 19, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/389,769, inventors Honglei; Liu et al., filed Apr. 19, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/434,010, inventors Dogaru; Sergiu et al., filed Jun. 6, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/552,559, inventors Seungwhan; Moon et al., filed Aug. 27, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/557,055, inventors Moon; Seungwhan et al., filed Aug. 30, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/659,070, inventors Lisa; Xiaoyi Huang et al., filed Oct. 21, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/659,203, inventors Huang; Lisa Xiaoyi et al., filed Oct. 21, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/659,419, inventor Huang; Lisa Xiaoyi, filed Oct. 21, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/703,700, inventors Ahmed; Aly et al., filed Dec. 4, 2019. [cited by applicant]
Co-pending U.S. Appl. No. 16/733,044, inventors Francislav; P. Penov et al., filed Jan. 2, 2020. [cited by applicant]
Co-pending U.S. Appl. No. 16/741,630, inventors Crook; Paul Anthony et al., filed Jan. 13, 2020. [cited by applicant]
Co-pending U.S. Appl. No. 16/741,642, inventors Fuchun; Peng et al., filed Jan. 13, 2020. [cited by applicant]
Co-pending U.S. Appl. No. 16/742,668, inventors Xiaohu; Liu et al., filed Jan. 14, 2020. [cited by applicant]
Co-pending U.S. Appl. No. 16/742,769, inventors Liu; Xiaohu et al., filed Jan. 14, 2020. [cited by applicant]
Co-Pending U.S. Appl. No. 15/953,957, filed Apr. 16, 2018, 117 pages. [cited by applicant]
Dalton J., et al., “Vote Goat: Conversational Movie Recommendation,” The 41st International ACM SIGIR Conference on Research Development in Information Retrieval, ACM, Jul. 2018, 5 Pages. [cited by applicant]
Dettmers T., et al., “Convolutional 2D Knowledge Graph Embeddings,” Thirty-Second AAAI Conference on Artificial Intelligence, Apr. 25, 2018, 8 Pages. [cited by applicant]
Dinan E., et al., “Advances in Conversational AI,” [online], Aug. 2, 2019, 12 Pages, Retrieved from the Internet: URL: https://ai.facebook.com/blog/advances-in-conversational-ai/. [cited by applicant]
Dubey M., et al., “EARL: Joint Entity and Relation Linking for Question Answering Over Knowledge Graphs,” International Semantic Web Conference, Springer, Cham, Jun. 25, 2018, 16 Pages. [cited by applicant]
Dubin R., et al., “Adaptation Logic for HTTP Dynamic Adaptive Streaming using Geo-Predictive Crowdsourcing,” Springer, Feb. 5, 2016, 11 Pages. [cited by applicant]
Duchi J., et al., “Adaptive Subgradient Methods for Online Learning and Stochastic Optimization,” Journal of Machine Learning Research, Jul. 2011, vol. 12, pp. 2121-2159. [cited by applicant]
Dyer C., et al., “Recurrent Neural Network Grammars,” Proceedings of NAACL-HLT, San Diego, California, Jun. 12-17, 2016, pp. 199-209. [cited by applicant]
Ganin Y., et al., “Domain-Adversarial Training of Neural Networks,” Journal of Machine Learning Research, Jan. 2016, vol. 17, pp. 2096-2030. [cited by applicant]
Gao Y., “Demo for Interactive Text-to-Speech via Semi-Supervised Style Transfer Learning,” Interspeech [Online], 2019 [Retrieved on Oct. 21, 2019], 3 Pages, Retrieved from Internet: URL: https://yolanda-gao.github.io/In… [cited by applicant]
Ghazvininejad M., et al., “A Knowledge-Grounded Neural Conversation Model,” Thirty-Second AAAI Conference on Artificial Intelligence, Apr. 27, 2018, 8 Pages. [cited by applicant]
Glass J., “A Brief Introduction to Automatic Speech Recognition,” Nov. 13, 2007, 22 Pages, Retrieved from Internet: http://www.cs.columbia.edu/˜mcollins/6864/slides/asr.pdf. [cited by applicant]
Goetz J., et al., “Active Federated Learning,” Machine Learning, Sep. 27, 2019, 5 Pages. [cited by applicant]
Google Allo Makes Conversations Eeasier, Productive, and more Expressive, May 19, 2016 [Retrieved on Jul. 11, 2019], 13 Pages, Retrieved from Internet: URL: https://www.trickyways.com/2016/05/google-allo-makes-conversat… [cited by applicant]
Han K., et al., “Speech Emotion Recognition using Deep Neural Network and Extreme Learning Machine,” Fifteenth Annual Conference of the International Speech Communication Association, Sep. 2014, 5 pages. [cited by applicant]
He H., et al., “Learning Symmetric Collaborative Dialogue Agents with Dynamic Knowledge Graph Embeddings,” Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics, Apr. 24, 2017, 18 Pages. [cited by applicant]
Henderson M., et al., “The Second Dialog State Tracking Challenge,” Proceedings of the 15th Annual Meeting of the Special Interest Group on Discourse and Dialogue (SJGDL4L), Jun. 18-20, 2014, 10 Pages. [cited by applicant]
Henter G.E., et al., “Principles for Learning Controllable TTS from Annotated and Latent Variation,” Interspeech, Aug. 20-24, 2017, 5 Pages. [cited by applicant]
Hodari Z., et al., “Learning Interpretable Control Dimensions for Speech Synthesis by Using External Data,” Interspeech, Sep. 2-6, 2018, 5 Pages. [cited by applicant]
Honglei L., et al., “Explore-Exploit: A Framework for Interactive and Online Learning,” Dec. 1, 2018, 7 pages. [cited by applicant]
Hsiao W., et al., “Fashion++: Minimal Edits for Outfit Improvement,” Computer Vision and Pattern Recognition, Apr. 19, 2019, 17 Pages. [cited by applicant]
Huang S., “Word2Vec and FastText Word Embedding with Gensim,” Towards Data Science [Online], Feb. 4, 2018 [Retrieved on Jun. 14, 2019], 11 Pages, Retrieved from Internet: URL: https://towardsdatascience.com/word-embeddi… [cited by applicant]
Hudson D.A., et al., “GQA: A New Dataset for Real-world Visual Reasoning and Compositional Question Answering,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, May 10, 2019, 18 Pages. [cited by applicant]
Agrawal A., et al., “VQA: Visual Question Answering,” International Journal of Computer Vision, Oct. 2016, pp. 1-25. [cited by applicant]
Bahdanau D., et al., “Neural Machine Translation by Jointly Learning to Align and Translate,” Conference paper at International Conference on Learning Representations, arXiv preprint arXiv:1409.0473v7, Mar. 2015, 15 pag… [cited by applicant]
Bailey K., “Conversational AI and the Road Ahead,” TechCrunch [Online], Feb. 15, 2018 [Retrieved on Jun. 14, 2019], 13 Pages, Retrieved from Internet: URL: https://techcrunch.com/2017/02/25/conversational-ai-and-the-roa… [cited by applicant]
Banse R., et al., “Acoustic Profiles in Vocal Emotion Expression,” Journal of Personality and Social Psychology, Mar. 1996, vol. 70 (3), pp. 614-636. [cited by applicant]
Bast H., et al., “Easy Access to the Freebase Dataset,” Proceedings of the 23rd International Conference on World Wide Web, ACM, Apr. 7-11, 2014, 4 Pages. [cited by applicant]
Bauer L., et al., “Commonsense for Generative Multi-Hop Question Answering Tasks,” 2018 Empirical Methods in Natural Language Processing, Jun. 1, 2019, 22 Pages. [cited by applicant]
Bordes A., et al., “Learning End-to-End Goal-Oriented Dialog,” Facebook AI, New York, USA, Mar. 30, 2017, 15 Pages. [cited by applicant]
Bordes A., et al., “Translating Embeddings for Modeling Multi-Relational Data,” Advances in Neural Information Processing Systems, Dec. 5, 2013, 9 Pages. [cited by applicant]
Bordes A., “Large-Scale Simple Question Answering with Memory Networks,” arXiv preprint arXiv:1506.02075, Jun. 5, 2015, 10 Pages. [cited by applicant]
Bui D., et al., “Federated User Representation Learning,” Machine Leaning, Sep. 27, 2019, 10 Pages. [cited by applicant]
Busso C., et al., “IEMOCAP: Interactive Emotional Dyadic Motion Capture Database,” Language Resources and Evaluation, Dec. 1, 2008, vol. 42 (4), 30 Pages. [cited by applicant]
Calrke P., et al., “Think You have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge,” Allen Instituite for Artificial Intelligence, Seattle, WA, USA, Mar. 14, 2018, 10 Pages. [cited by applicant]
Carlson A., et al., “Toward an Architecture for Never-Ending Language Learning,” Twenty-Fourth AAAI Conference on Artificial intelligence, Jul. 5, 2010, 8 Pages. [cited by applicant]
Challa A., et al., “Generate, Filter, and Rank: Grammaticality Classification for Production-Ready NLG Systems,” Facebook Conversational AI, Apr. 9, 2019, 10 Pages. [cited by applicant]
“Chat Extensions,” [online], Apr. 18, 2017, 8 Pages, Retrieved from the Internet: URL: https://developers.facebook.com/docs/messenger-platform/guides/chat-extensions/. [cited by applicant]
Chen C.Y., et al., “Gunrock: Building a Human-Like Social Bot by Leveraging Large Scale Real User Data,” 2nd Proceedings of Alexa Price, 2018, 19 Pages. [cited by applicant]
Chen Y., et al., “Jointly Modeling Inter-Slot Relations by Random Walk on Knowledge Graphs for Unsupervised Spoken Language Understanding,” Proceedings of the 2015 Conference of the North American Chapter of the Associa… [cited by applicant]
Choi E., et al., “QuAC: Question Answering in Context,” Allen Instituite for Artificial Intelligence, Aug. 28, 2018, 11 Pages. [cited by applicant]
Conneau A., et al., “Supervised Learning of Universal Sentence Representations from Natural Language Inference Data,” Conference: Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing, J… [cited by applicant]
Co-pending U.S. Appl. No. 14/593,723, inventors Colin; Patrick Treseler et al., filed Jan. 9, 2015. [cited by applicant]
Co-pending U.S. Appl. No. 15/808,638, inventors Ryan; Brownhill et al., filed Nov. 9, 2017. [cited by applicant]
Co-pending U.S. Appl. No. 15/966,455, inventor Scott; Martin, filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 15/967,193, inventors Testuggine; Davide et al., filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 15/967,239, inventors Vivek; Natarajan et al., filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 15/967,279, inventors Fuchun; Peng et al., filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 15/967,290, inventors Fuchun; Peng et al., filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 15/967,342, inventors Vivek; Natarajan et al., filed Apr. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/011,062, inventors Jinsong; Yu et al., filed Jun. 18, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/025,317, inventors Gupta; Sonal et al., filed Jul. 2, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/036,827, inventors Emmanouil; Koukoumidis et al., filed Jul. 16, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/038,120, inventors Schissel; Jason et al., filed Jul. 17, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/048,049, inventor Markku; Salkola, filed Jul. 27, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/048,072, inventor Salkola; Markku, filed Jul. 27, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/048,101, inventor Salkola; Markku , filed Jul. 27, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/053,600, inventors Vivek; Natarajan et al., filed Aug. 2, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/057,414, inventors Kahn; Jeremy Gillmor et al., filed Aug. 7, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/103,775, inventors Zheng; Zhou et al., filed Aug. 14, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/107,601, inventor Rajesh; Krishna Shenoy, filed Aug. 21, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/107,847, inventors Rajesh; Krishna Shenoy et al., filed Aug. 21, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/118,169, inventors Baiyang; Liu et al., filed Aug. 30, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/121,393, inventors Zhou; Zheng et al., filed Sep. 4, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/127,173, inventors Zheng; Zhou et al., filed Sep. 10, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/129,638, inventors Vivek; Natarajan et al., filed Sep. 12, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/135,752, inventors Liu; Xiaohu et al., filed Sep. 19, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/150,069, inventors Jiedan; Zhu et al., filed Oct. 2, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/150,184, inventors Francislav; P. Penov et al., filed Oct. 2, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/151,040, inventors Brian; Nelson et al., filed Oct. 3, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/168,536, inventors Dumoulin; Benoit F et al., filed Oct. 23, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/176,081, inventors Anusha; Balakrishnan et al., filed Oct. 31, 2018. [cited by applicant]
Co-pending U.S. Appl. No. 16/176,312, inventors Koukoumidis; Emmanouil et al., filed Oct. 31, 2018. [cited by applicant]