IP Library Granted Patent US 11,227,122
Granted Patent B1
US 11,227,122 · App. 16/731,345 · Granted Jan 18, 2022

Methods, mediums, and systems for representing a model in a memory of device

Inventors: Prince Gill (Menlo Park, CA); Honglei Liu (Menlo Park, CA); Wenhai Yang (Menlo Park, CA); Kshitiz Malik (Menlo Park, CA); Nanshu Wang (Menlo Park, CA); David Reiss (Menlo Park, CA)
Assignee: FACEBOOK, INC.
G06F40/30G06F40/20G06N7/00G06N7/005G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,122
App. No.
16/731,345
Granted
Jan 18, 2022
Kind
B1
Abstract

Exemplary embodiments relate to methods, mediums, and systems for moving language models from a server to the client device. Such embodiments may be deployed in an environment where the server is not able to provide modeling services to the clients, such as an end-to-end encrypted (E2EE) environment. Several different techniques are described to address issues of size and complexity reduction, model architecture optimization, model training, battery power reduction, and latency reduction.

Claims (40)

1. A method comprising:

accessing a model configured to be executed as a plurality of submodels, the submodels comprising a first submodel and a second submodel;

loading the first submodel into a memory of a device;

generating a first output from the first submodel;

determining to proceed with processing the model based on the first output;

in response to the determining, removing the first submodel from the memory and loading the second submodel into the memory after the first submodel is removed to continue executing the model; and

generating a second output from the second submodel.

2. The method of claim 1 , wherein the first output is executed before the second submodel based on the first submodel screening out more options than the second submodel or based on the first submodel being more likely to terminate processing than the second submodel.

3. The method of claim 2 , wherein the first submodel is a personalization model that determines whether or not a given user is likely to use a recommendation output by the model.

4. The method of claim 1 , wherein the first submodel is independent of calculations performed by the second submodel.

5. The method of claim 1 , wherein the model is a natural language understanding model.

6. The method of claim 1 , wherein the device is an end-user device in an end-to-end encrypted environment.

7. The method of claim 1 , wherein a device executing the model exhibits reduced power usage as compared to the same device loading an entirety of the model into memory at once.

8. A non-transitory computer-readable medium storing instructions configured to be executed by a processor to cause the processor to:

access a model configured to be executed as a plurality of submodels, the submodels comprising a first submodel and a second submodel;

load the first submodel into a memory of a device;

generate a first output from the first submodel;

determine to proceed with processing the model based on the first output;

in response to the determining, remove the first submodel from the memory and loading the second submodel into the memory after the first submodel is removed to continue executing the model; and

generate a second output from the second submodel.

9. The medium of claim 8 , wherein the first output is executed before the second submodel based on the first submodel screening out more options than the second submodel or based on the first submodel being more likely to terminate processing than the second submodel.

10. The medium of claim 9 , wherein the first submodel is a personalization model that determines whether or not a given user is likely to use a recommendation output by the model.

11. The medium of claim 8 , wherein the first submodel is independent of calculations performed by the second submodel.

12. The medium of claim 8 , wherein the model is a natural language understanding model.

13. The medium of claim 8 , wherein the device is an end-user device in an end-to-end encrypted environment.

14. The medium of claim 8 , wherein a device executing the model exhibits reduced power usage as compared to the same device loading an entirety of the model into memory at once.

15. An apparatus comprising:

a non-transitory computer-readable medium storing a model configured to be executed as a plurality of submodels, the submodels comprising a first submodel and a second submodel;

a memory configured to support the plurality of submodels during execution; and

a hardware processor configured to:

load the first submodel into the memory;

generate a first output from the first submodel;

determine to proceed with processing the model based on the first output;

in response to the determining, remove the first submodel from the memory and loading the second submodel into the memory after the first submodel is removed to continue executing the model; and

generate a second output from the second submodel.

16. The apparatus of claim 15 , wherein the first output is executed before the second submodel based on the first submodel screening out more options than the second submodel or based on the first submodel being more likely to terminate processing than the second submodel.

17. The apparatus of claim 16 , wherein the first submodel is a personalization model that determines whether or not a given user is likely to use a recommendation output by the model.

18. The apparatus of claim 15 , wherein the first submodel is independent of calculations performed by the second submodel.

19. The apparatus of claim 15 , wherein the model is a natural language understanding model.

20. The apparatus of claim 15 , wherein the apparatus is an end-user device in an end-to-end encrypted environment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2024
From: GILL, PRINCE; LIU, HONGLEI; YANG, WENHAI; MALIK, KSHITIZ; WANG, NANSHU; REISS, DAVID
To: FACEBOOK, INC.
Reel/Frame 067204/0664 →
CHANGE OF NAME Recorded Feb 9, 2022
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058981/0383 →
Continuity (1)
Continuation 16731304 · Dec 31, 2019
Cited By (2)
US 12,321,428 US 12,626,690