IP Library Granted Patent US 11,354,521
Granted Patent B2
US 11,354,521 · App. 16/792,572 · Granted Jun 7, 2022

Facilitating communications with automated assistants in multiple languages

Inventors: James Kuczmarski (San Francisco, CA); Vibhor Jain (Sunnyvale, CA); Amarnag Subramanya (Menlo Park, CA); Nimesh Ranjan (San Francisco, CA); Melvin Jose Johnson Premkumar (Sunnyvale, CA); Vladimir Vuskovic (Zollikerberg, CH); Luna Dai (San Francisco, CA); Daisuke Ikeda (Sunnyvale, CA); Nihal Sandeep Balani (Sunnyvale, CA); Jinna Lei (San Francisco, CA); Mengmeng Niu (San Jose, CA); Hongjie Chai (Palo Alto, CA); Wangqing Yuan (Wilmington, MA)
Assignee: GOOGLE LLC
G06F40/58G06F16/3329G06F16/3337G06F40/47G06K9/6215G06N20/00H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,354,521
App. No.
16/792,572
Granted
Jun 7, 2022
Kind
B2
Abstract

Techniques described herein relate to facilitating end-to-end multilingual communications with automated assistants. In various implementations, speech recognition output may be generated based on voice input in a first language. A first language intent may be identified based on the speech recognition output and fulfilled in order to generate a first natural language output candidate in the first language. At least part of the speech recognition output may be translated to a second language to generate an at least partial translation, which may then be used to identify a second language intent that is fulfilled to generate a second natural language output candidate in the second language. Scores may be determined for the first and second natural language output candidates, and based on the scores, a natural language output may be selected for presentation.

Claims (39)

1. A method for generating training data for training a machine translation model to translate from a first language to a second language, the method implemented by one or more processors and comprising:

performing natural language understanding processing in the first language to identify a first language intent of a user based on a textual query in the first language that is obtained from input provided by the user;

using the machine translation model, translating the textual query in the first language to generate a translation of the textual query in the second language;

performing natural language understanding processing in the second language to identify a second language intent of the user based on the translation of the textual query in the second language;

comparing the first and second language intents;

in response to determining, based on the comparing, that the first and second language intents match, generating and storing a training example of the training data using the textual query in the first language and the translation of the textual query in the second language; and

updating the machine translation model based on the training data.

2. The method of claim 1 , further comprising:

receiving voice input provided by a user at an input component of a client device in the first language; and

performing speech recognition on the voice input to generate the textual query in the first language.

3. The method of claim 1 , wherein the updating comprises training the machine translation model using the training data.

4. The method of claim 1 , wherein the machine translation model comprises a neural machine translation model.

5. The method of claim 1 , wherein the comparing includes comparing one or more arguments associated with the first language intent with one or more arguments associated with the second language intent.

6. At least one non-transitory computer-readable medium comprising instructions that, in response to execution of the instructions by one or more processors, cause the one or more processors to perform the following operations to generate training data for training a machine translation model to translate from a first language to a second language:

perform natural language understanding processing in the first language to identify a first language intent of a user based on a textual query in the first language that is obtained from input provided by the user;

using the machine translation model, translate the textual query in the first language to generate a translation of the textual query in the second language;

perform natural language understanding processing in the second language to identify a second language intent of the user based on the translation of the textual query in the second language;

compare the first and second language intents;

in response to a determination that the first and second language intents match, generate and store a training example of the training data using the textual query in the first language and the translation of the textual query in the second language; and

update the machine translation model based on the training data.

7. The at least one non-transitory computer-readable medium of claim 6 , further comprising instructions to:

receive voice input provided by a user at an input component of a client device in the first language; and

perform speech recognition on the voice input to generate the textual query in the first language.

8. The at least one non-transitory computer-readable medium of claim 6 , wherein the instructions to update include instructions to train the machine translation model using the training data.

9. The at least one non-transitory computer-readable medium of claim 6 , wherein the machine translation model comprises a neural machine translation model.

10. The at least one non-transitory computer-readable medium of claim 6 , wherein the comparison includes a comparison of one or more arguments associated with the first language intent to one or more arguments associated with the second language intent.

11. A system comprising one or more processors and memory storing instructions that, in response to execution of the instructions by one or more of the processors, cause the one or more processors to perform the following operations to generate training data for training a machine translation model to translate from a first language to a second language:

perform natural language understanding processing in the first language to identify a first language intent of a user based on a textual query in the first language that is obtained from input provided by the user;

using the machine translation model, translate the textual query in the first language to generate a translation of the textual query in the second language;

perform natural language understanding processing in the second language to identify a second language intent of the user based on the translation of the textual query in the second language;

compare the first and second language intents;

in response to a determination that the first and second language intents match, generate and store a training example of the training data using the textual query in the first language and the translation of the textual query in the second language; and

update the machine translation model based on the training data.

12. The system of claim 11 , further comprising instructions to:

receive voice input provided by a user at an input component of a client device in the first language; and

perform speech recognition on the voice input to generate the textual query in the first language.

13. The system of claim 11 , wherein the instructions to compare comprise instructions to train the machine translation model using the training data.

14. The system of claim 11 , wherein the machine translation model comprises a neural machine translation model.

15. The system of claim 11 , wherein the comparison includes a comparison of one or more arguments associated with the first language intent to one or more arguments associated with the second language intent.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2020
From: KUCZMARSKI, JAMES; JAIN, VIBHOR; SUBRAMANYA, AMARNAG; RANJAN, NIMESH; PREMKUMAR, MELVIN JOSE JOHNSON; VUSKOVIC, VLADIMIR; DAI, LUNA; IKEDA, DAISUKE; BALANI, NIHAL SANDEEP; LEI, JINNA; NIU, MENGMENG; CHAI, HONGJIE; YUAN, WANGQING
To: GOOGLE LLC
Reel/Frame 052232/0040 →
Continuity (3)
Continuation In Part 16082175
Provisional Application 62639740 · Mar 7, 2018
Related Publication 20200184158A1 · Jun 11, 2020
Cited By (1)
US 12,475,333