IP Library › Granted Patent US 11,544,553
Granted Patent B1
US 11,544,553 · App. 16/448,749 · Granted Jan 3, 2023

Data retrieval using reinforced co-learning for semi-supervised ranking

Inventors: Shibi He (Urbana, IL); Yanen Li (Los Angeles, CA); Ning Xu (Irvine, CA)
Assignee: Snap Inc.
G06N3/08G06K9/6267G06K9/6297G06N3/0454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,544,553
App. No.
16/448,749
Granted
Jan 3, 2023
Kind
B1
Abstract

A computer-implement method comprises: training a classifier with labeled data from a dataset; classifying, by the trained classifier, unlabeled data from the dataset; providing, by the classifier to a policy gradient, a reward signal for each data/query pair; transferring, by the classifier to a ranker, learning; training, by the policy gradient, the ranker; ranking data from the dataset based on a query; and retrieving data from the ranked data in response to the query.

Claims (40)

1. A computer-implemented method, comprising:

training a classifier with labeled data from a dataset;

classifying, by the trained classifier, unlabeled data from the dataset;

providing iteratively, by the classifier to a policy gradient function, a reward signal for data/query pairs, wherein the reward signal is a combination of a normalized discounted cumulative gain and a discriminative score output by the classifier;

transferring, by the classifier to a ranker, learning from the classifying;

training, by the policy gradient function, the ranker;

ranking, by the trained ranker, data from the dataset based on a query; and

retrieving data from the ranked data in response to the query.

2. The method of claim 1 , wherein the classifier and ranker are neural networks.

3. The method of claim 1 , wherein the classifier is a neural network with feature sharing.

4. The method of claim 3 , wherein the transferred learning includes an intermediate feature.

5. The method of claim 1 , wherein the training the classifier is iterative.

6. The method of claim 1 , wherein the training the classifier comprises using both positive and negative samples from the dataset.

7. The method of claim 6 , wherein the negative samples further include negative classified data.

8. The method of claim 1 , wherein the ranking the data from the dataset based on the query comprises a Markov Decision Process.

9. A system, comprising:

one or more processors of a machine;

a memory storing instruction that, when executed by the one or more processors, cause the machine to perform operations comprising:

training a classifier with labeled data from a dataset;

classifying, by the trained classifier, unlabeled data from the dataset;

providing iteratively, by the classifier to a policy gradient function, a reward signal for data/query pairs, wherein the reward signal is a combination of a normalized discounted cumulative gain and a discriminative score output by the classifier;

transferring, by the classifier to a ranker, learning from the classifying;

training, by the policy gradient function, the ranker;

ranking, by the trained ranker, data from the dataset based on a query; and

retrieving data from the ranked data in response to the query.

10. A machine-readable storage device embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:

training a classifier with labeled data from a dataset;

classifying, by the trained classifier, unlabeled data from the dataset;

providing iteratively, by the classifier to a policy gradient function, a reward signal for data/query pairs, wherein the reward signal is a combination of a normalized discounted cumulative gain and a discriminative score output by the classifier;

transferring, by the classifier to a ranker, learning from the classifying;

training, by the policy gradient function, the ranker;

ranking, by the trained ranker, data from the dataset based on a query; and

retrieving data from the ranked data in response to the query.

11. The device of claim 10 , wherein the classifier and ranker are neural networks.

12. The device of claim 10 , wherein the classifier is a neural network with feature sharing.

13. The device of claim 12 , wherein the transferred learning includes an intermediate feature.

14. The device of claim 10 , wherein the training the classifier is iterative.

15. The device of claim 10 , wherein the training the classifier comprises using both positive and negative samples from the dataset.

16. The device of claim 15 , wherein the negative samples include negative classified data.

17. The device of claim 10 , wherein the ranking the data from the dataset based on the query comprises a Markov Decision Process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: HE, SHIBI; LI, YANEN; XU, NING
To: SNAP INC.
Reel/Frame 061686/0217 →
Cited By (1)
US 12,361,287