IP Library Granted Patent US 11,568,309
Granted Patent B1
US 11,568,309 · App. 16/460,185 · Granted Jan 31, 2023

Systems and methods for resource-efficient data collection for multi-stage ranking systems

Inventor: Nikita Igorevych Lytkin (Sunnyvale, CA)
Assignee: Meta Platforms, Inc.
G06N20/00G06F17/18G06K9/623G06K9/6259
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,309
App. No.
16/460,185
Granted
Jan 31, 2023
Kind
B1
Abstract

Systems, methods, and non-transitory computer-readable media can receive a set of candidate training items for training an early stage model in a multi-stage recall optimization model, wherein the multi-stage recall optimization model comprises the early stage model and a target model. A random subset of the candidate training items is selected from the set of candidate training items. For each training item in the subset of candidate training items, a score is determined based on the target model. Each training item in the subset of candidate training items is labeled with a label based on a probability of the training item being a top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

Claims (39)

1. A computer-implemented method comprising:

receiving, by a computing system, a set of candidate training items for training an early stage model in a multi-stage recall optimization model, wherein the multi-stage recall optimization model comprises the early stage model and a target model;

selecting, by the computing system, a random subset of candidate training items from the set of candidate training items to train the early stage model;

determining, by the computing system, for each training item in the subset of candidate training items, a score based on the target model; and

labeling, by the computing system, each training item in the subset of candidate training items with a label based on a probability of the training item being in a top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

2. The computer-implemented method of claim 1 , wherein the label is a non-binary label.

3. The computer-implemented method of claim 2 , wherein the label is a value between 0 and 1.

4. The computer-implemented method of claim 1 , wherein the probability of each training item being in the top-K of the set of candidate training items is determined based on a hypergeometric distribution.

5. The computer-implemented method of claim 4 , wherein the label for a training item is equal to the probability of the training item being in the top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

6. The computer-implemented method of claim 1 , further comprising training the early stage model based on the labels for the subset of candidate training items.

7. The computer-implemented method of claim 1 , wherein the target model is trained based on a labeled set of training data, wherein the labeled set of training data is labeled based on user feedback information.

8. The computer-implemented method of claim 1 , further comprising ranking the subset of candidate training items based on the scores.

9. The computer-implemented method of claim 1 , further comprising:

providing a set of candidate items to the multi-stage recall optimization model; and

identifying a set of the top-K items from the set of candidate items based on the multi-stage recall optimization model.

10. The computer-implemented method of claim 9 , wherein identifying the set of the top-K items from the set of candidate items based on the multi-stage recall optimization model comprises:

identifying a first subset of the set of candidate items based on the early stage model;

providing the first subset of the set of candidate items to the target model; and

identifying the set of top-K items from the set of candidate items based on the target model.

11. A system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, cause the system to perform a method comprising:

receiving a set of candidate training items for training an early stage model in a multi-stage recall optimization model, wherein the multi-stage recall optimization model comprises the early stage model and a target model;

selecting a random subset of candidate training items from the set of candidate training items to train the early stage model;

determining, for each training item in the subset of candidate training items, a score based on the target model; and

labeling each training item in the subset of candidate training items with a label based on a probability of the training item being in a top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

12. The system of claim 11 , wherein the label is a non-binary label.

13. The system of claim 12 , wherein the label is a value between 0 and 1.

14. The system of claim 11 , wherein the probability of each training item being in the top-K of the set of candidate training items is determined based on a hypergeometric distribution.

15. The system of claim 14 , wherein the label for a training item is equal to the probability of the training item being in the top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

16. A non-transitory computer-readable storage medium including instructions that, when executed by at least one processor of a computing system, cause the computing system to perform a method comprising:

receiving a set of candidate training items for training an early stage model in a multi-stage recall optimization model, wherein the multi-stage recall optimization model comprises the early stage model and a target model;

selecting a random subset of candidate training items from the set of candidate training items to train the early stage model;

determining, for each training item in the subset of candidate training items, a score based on the target model; and

labeling each training item in the subset of candidate training items with a label based on a probability of the training item being in a top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the label is a non-binary label.

18. The non-transitory computer-readable storage medium of claim 17 , wherein the label is a value between 0 and 1.

19. The non-transitory computer-readable storage medium of claim 16 , wherein the probability of each training item being in the top-K of the set of candidate training items is determined based on a hypergeometric distribution.

20. The non-transitory computer-readable storage medium of claim 19 , wherein the label for a training item is equal to the probability of the training item being in the top-K of the set of candidate training items had the set of candidate training items been scored based on the target model.

Assignments (2)
CHANGE OF NAME Recorded Nov 23, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058238/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2019
From: LYTKIN, NIKITA IGOREVYCH
To: FACEBOOK, INC.
Reel/Frame 049891/0520 →