IP Library Granted Patent US 10,679,148
Granted Patent B2
US 10,679,148 · App. 16/402,787 · Granted Jun 9, 2020

Implicit bridging of machine learning tasks

Inventors: Zhifeng Chen (Sunnyvale, CA); Michael Schuster (Saratoga, CA); Melvin Jose Johnson Premkumar (Sunnyvale, CA); Yonghui Wu (Fremont, CA); Quoc V. Le (Sunnyvale, CA); Maxim Krikun (Castro Valley, CA); Thorsten Brants (Palo Alto, CA)
Assignee: Google LLC
G06N20/00G06F40/44G06F40/47G06N3/0445G06N3/0454G06N3/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,679,148
App. No.
16/402,787
Granted
Jun 9, 2020
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media for performing machine learning tasks. One method includes receiving (i) a model input, and (ii) data identifying a first machine learning task to be performed on the model input to generate a first type of model output for the model input; augmenting the model input with an identifier for the first machine learning task to generate an augmented model input; and processing the augmented model input using a machine learning model. An exemplary system applying implicit bridging for machine learning tasks, as described in this specification, trains a machine learning model to perform certain types of machine learning tasks without requiring explicit training data for the certain types of machine learning tasks to be used during training.

Claims (33)

1. A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving (i) a source language text segment, and (ii) data identifying a target language that the source language text segment is to be translated into by a machine learning model, wherein:

the machine learning model has been trained on a plurality of different paired data sets,

each paired data set includes (iii) first language text segments in a respective first language and (iv) second language text segments in a respective second language that are translations of the first language text segments into the second language, and

wherein the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language;

augmenting the source language text segment with an identifier that identifies at least the target language to generate an augmented model input; and

processing the augmented model input using the machine learning model to generate a target language text segment that is a translation of the source language text segment into the target language even though the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language.

2. The system of claim 1 , wherein the machine learning model comprises (i) an encoder subsystem configured to receive an augmented model input, and (ii) a decoder subsystem configured to generate the target language text segment.

3. The system of claim 2 , wherein the encoder subsystem and decoder subsystem comprise respective recurrent neural networks.

4. The system of claim 2 , wherein the decoder subsystem comprises an attention mechanism.

5. The system of claim 1 , wherein the augmented model input comprises the source language text segment with a prepended token identifier for at least the target language.

6. A computer-implemented method comprising:

receiving (i) a source language text segment, and (ii) data identifying a target language that the source language text segment is to be translated into by a machine learning model, wherein:

the machine learning model has been trained on a plurality of different paired data sets,

each paired data set includes (iii) first language text segments in a respective first language and (iv) second language text segments in a respective second language that are translations of the first language text segments into the second language, and

wherein the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language;

augmenting the source language text segment with an identifier that identifies at least the target language to generate an augmented model input; and

processing the augmented model input using the machine learning model to generate a target language text segment that is a translation of the source language text segment into the target language even though the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language.

7. The method of claim 6 , wherein the machine learning model comprises (i) an encoder subsystem configured to receive an augmented model input, and (ii) a decoder subsystem configured to generate the target language text segment.

8. The method of claim 7 , wherein the encoder subsystem and decoder subsystem comprise respective recurrent neural networks.

9. The method of claim 2 , wherein the decoder subsystem comprises an attention mechanism.

10. The method of claim 6 , wherein the augmented model input comprises the source language text segment with a prepended token identifier for at least the target language.

11. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

receiving (i) a source language text segment, and (ii) data identifying a target language that the source language text segment is to be translated into by a machine learning model, wherein:

the machine learning model has been trained on a plurality of different paired data sets,

each paired data set includes (iii) first language text segments in a respective first language and (iv) second language text segments in a respective second language that are translations of the first language text segments into the second language, and

wherein the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language;

augmenting the source language text segment with an identifier that identifies at least the target language to generate an augmented model input; and

processing the augmented model input using the machine learning model to generate a target language text segment that is a translation of the source language text segment into the target language even though the machine learning model has not been trained on any paired data sets that include translations of text segments from the source language into the target language.

12. The computer-readable storage media of claim 11 , wherein the machine learning model comprises (i) an encoder subsystem configured to receive an augmented model input, and (ii) a decoder subsystem configured to generate the target language text segment.

13. The computer-readable storage media of claim 12 , wherein the encoder subsystem and decoder subsystem comprise respective recurrent neural networks.

14. The computer-readable storage media of claim 12 , wherein the decoder subsystem comprises an attention mechanism.

15. The computer-readable storage media of claim 11 , wherein the augmented model input comprises the source language text segment with a prepended token identifier for at least the target language.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2019
From: CHEN, ZHIFENG; SCHUSTER, MICHAEL; PREMKUMAR, MELVIN JOSE; WU, YONGHUI; LE, QUOC V.; KRIKUN, MAXIM; BRANTS, THORSTEN
To: GOOGLE INC.
Reel/Frame 049187/0879 →
CHANGE OF NAME Recorded May 15, 2019
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 049193/0924 →
Continuity (4)
Continuation PCTUS2017059776 · Nov 2, 2017
Continuation 15394708 · Dec 29, 2016
Provisional Application 62418098 · Nov 4, 2016
Related Publication 20190258961A1 · Aug 22, 2019
Cited By (1)
US 12,462,113