IP Library Granted Patent US 11,455,555
Granted Patent B1
US 11,455,555 · App. 16/731,321 · Granted Sep 27, 2022

Methods, mediums, and systems for training a model

Inventors: Prince Gill (Menlo Park, CA); Honglei Liu (Menlo Park, CA); Wenhai Yang (Menlo Park, CA); Kshitiz Malik (Menlo Park, CA); Nanshu Wang (Menlo Park, CA); David Reiss (Menlo Park, CA)
Assignee: META PLATFORMS, INC.
G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,455,555
App. No.
16/731,321
Granted
Sep 27, 2022
Kind
B1
Abstract

Exemplary embodiments relate to methods, mediums, and systems for moving language models from a server to the client device. Such embodiments may be deployed in an environment where the server is not able to provide modeling services to the clients, such as an end-to-end encrypted (E2EE) environment. Several different techniques are described to address issues of size and complexity reduction, model architecture optimization, model training, battery power reduction, and latency reduction.

Claims (33)

1. A method comprising:

accessing a local model on an end-user device and a machine-learning-trained teacher model remote from the end-user device, the machine-learning-trained teacher model configured to teach a first task to the local model, the first task to comprise detecting when to trigger a suggestion, a recommendation, a reminder, or calling an agent, the local model being smaller in size than the teacher model;

receiving a baseline output from the teacher model; and

training the local model so that an output of the local model approximates the baseline output of the teacher model;

wherein the local model is configured to be applied in an end-to-end encrypted environment.

2. The method of claim 1 , wherein the teacher model is a first teacher model, and further comprising training the local model using a second teacher model distinct from the first teacher model.

3. The method of claim 2 , wherein the first teacher model and the second teacher model generate different baseline outputs after receiving a common input, and the local model is trained based on a combination of the different baseline outputs.

4. The method of claim 2 , wherein the first teacher model is configured to teach the local model the first task and the second teacher model is configured to teach the local model a second task different from the first task.

5. The method of claim 4 , wherein the second task comprises one or more of detecting when to trigger a suggestion or recommendation, when to trigger a reminder, or when to trigger a calling of an agent.

6. The method of claim 1 , wherein the teacher model is configured to be trained based on training data and the local model is configured to be trained without access to the training data.

7. The method of claim 1 , wherein the local model is configured to be applied in the end-to-end encrypted environment for a communication system, the communication system to exchange messages between the end-user device and another end-user device, wherein message content for the messages is visible only to end-user devices of the communication system.

8. A non-transitory computer-readable medium storing instructions configured to be executed by a processor to cause the processor to:

access a local model on an end-user device and a machine-learning-trained teacher model remote from the end-user device, the machine-learning-trained teacher model configured to teach a first task to the local model, the first task to comprise detecting when to trigger a suggestion, a recommendation, a reminder, or calling an agent, the local model being smaller in size than the teacher model;

receive a baseline output from the teacher model; and

train the local model so that an output of the local model approximates the baseline output of the teacher model;

wherein the local model is configured to be applied in an end-to-end encrypted environment.

9. The medium of claim 8 , wherein the teacher model is a first teacher model, and further comprising training the local model using a second teacher model distinct from the first teacher model.

10. The medium of claim 9 , wherein the first teacher model and the second teacher model generate different baseline outputs after receiving a common input, and the local model is trained based on a combination of the different baseline outputs.

11. The medium of claim 9 , wherein the first teacher model is configured to teach the local model the first task and the second teacher model is configured to teach the local model a second task different from the first task.

12. The medium of claim 11 , wherein the second task comprises one or more of detecting when to trigger a suggestion or recommendation, when to trigger a reminder, or when to trigger a calling of an agent.

13. The medium of claim 8 , wherein the teacher model is configured to be trained based on training data and the local model is configured to be trained without access to the training data.

14. The medium of claim 8 , wherein the local model is configured to be applied in the end-to-end encrypted environment for a communication system, the communication system to exchange messages between the end-user device and at least one other end-user device, wherein message content for the messages is visible only to end-user devices of the communication system.

15. An apparatus comprising:

a network interface configured to:

access a local model on an end-user device and a machine-learning-trained teacher model remote from the end-user device, the machine-learning-trained teacher model configured to teach a first task to the local model, the first task to comprise detecting when to trigger a suggestion, a recommendation, a reminder, or calling an agent, the local model being smaller in size than the teacher model, and wherein the local model is configured to be applied in an end-to-end encrypted environment, and

receive a baseline output from the teacher model;

a non-transitory computer-readable medium storing instructions configured to train the local model so that an output of the local model approximates the baseline output of the teacher model; and

a processor configured to execute the instructions to provide a trainer.

16. The apparatus of claim 15 , wherein the teacher model is a first teacher model, and further comprising training the local model using a second teacher model distinct from the first teacher model.

17. The apparatus of claim 16 , wherein the first teacher model and the second teacher model generate different baseline outputs after receiving a common input, and the local model is trained based on a combination of the different baseline outputs.

18. The apparatus of claim 16 , wherein the first teacher model is configured to teach the local model the first task and the second teacher model is configured to teach the local model a second task different from the first task.

19. The apparatus of claim 18 , wherein the second task comprises one or more of detecting when to trigger a suggestion or recommendation, when to trigger a reminder, or when to trigger a calling of an agent.

20. The apparatus of claim 15 , wherein the teacher model is configured to be trained based on training data and the local model is configured to be trained without access to the training data.

Assignments (2)
CHANGE OF NAME Recorded Feb 9, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058981/0383 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 24, 2021
From: GILL, PRINCE; LIU, HONGLEI; YANG, WENHAI; MALIK, KSHITIZ; WANG, NANSHU; REISS, DAVID
To: FACEBOOK, INC.
Reel/Frame 057277/0517 →
Continuity (1)
Continuation 16731304 · Dec 31, 2019
Cited By (2)
US 12,530,040 US 12,641,047