IP Library Granted Patent US 11,436,235
Granted Patent B2
US 11,436,235 · App. 16/578,780 · Granted Sep 6, 2022

Pipeline for document scoring

Inventors: Ricardo Baeza-Yates (Palo Alto, CA); Berkant Baria Cambazoglu (Melbourne, AU); Darshan Mallenahalli Shankaralingappa (Berlin, DE); Matteo Catena (Barcelona, ES)
Assignee: NTENT
G06F16/24578G06F7/14G06F16/22G06F16/248G06F16/93G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,436,235
App. No.
16/578,780
Granted
Sep 6, 2022
Kind
B2
Abstract

One or more techniques and/or systems are provided for implementing a pipeline used to generate, train, test, and implement a document scoring model for assigning document scores to documents. Features from various sources are combined to create a joined page level feature set, a joined domain level feature set, and a host level feature set. Numerical features and content features are extracted from ground truth documents and random documents. The numerical features are joined with the joined feature sets to create a set of joined features. The document scoring model is trained using the set of joined features and a training technique. A document is scored with a document score using the document scoring model based upon the content features and the set of joined features with document scores obtained during training.

Claims (45)

1. A method comprising:

combining feature types within one or more levels of feature sets into a joined host level feature set, wherein the one or more levels of feature sets comprise page level features joined into a joined page level feature set, domain level features joined into a joined domain level feature set, and host level features joined into a joined host level feature set, wherein a domain level feature corresponds to a feature of a domain associated with a target document, and wherein a host level feature corresponds to a feature of a host associated with a target document;

extracting numerical features and content features from ground truth documents and random documents, wherein a numerical feature corresponds to a ratio of an amount of a first type of content within a target document to an amount of a second type of content within the target document;

joining the numerical features with the one or more levels of feature sets to create a set of joined features for the ground truth documents and the random documents;

training a document scoring model utilizing machine learning to score documents using the set of joined features;

scoring documents with document scores using the document scoring model based upon the content features and the set of joined features with document scores obtained during training; and

selectively indexing a subset of the documents based upon the document scores of the documents.

2. The method of claim 1 , wherein a set of search results for a query comprises a document, and the method comprising:

assigning a rank to a document within the set of search results based upon a document score assigned to the document.

3. The method of claim 2 , comprising:

displaying the set of search results in response to receiving the query, wherein a document is populated within the set of search results based upon the rank.

4. The method of claim 1 , comprising:

indexing a document based upon a document score exceeding a threshold.

5. The method of claim 1 , comprising:

refraining from indexing a document based upon a document score not exceeding a threshold.

6. The method of claim 1 , wherein the document score is indicative of an importance of the document.

7. The method of claim 1 , wherein the document score is indicative of at least one of an importance or quality of the document.

8. The method of claim 1 , wherein the document score is indicative of a relevancy of the document.

9. The method of claim 1 , wherein a numerical feature corresponds to a numerical statistic of a target document.

10. The method of claim 1 , wherein a document comprises a webpage.

11. The method of claim 1 , wherein a document comprises a text document.

12. The method of claim 1 , wherein a numerical feature corresponds to a number of times a target document is linked to.

13. The method of claim 1 , wherein the document score is indicative of a quality of the document.

14. The method of claim 2 , comprising displaying the set of search results.

15. The method of claim 1 , wherein the machine learning comprises a gradient boosted decision tree regression technique.

16. The method of claim 1 , comprising:

merging the numerical features with the content features for scoring a document using the document scoring model.

17. The method of claim 1 , wherein the document score is indicative of an importance and a quality of the document.

18. The method of claim 1 , wherein the content features comprise textual features of a target document.

19. A non-transitory machine readable medium comprising instructions for performing a method, which when executed by a machine, causes the machine to:

combine feature types within one or more levels of feature sets into a joined host level feature set, wherein the one or more levels of feature sets comprise page level features joined into a joined page level feature set, domain level features joined into a joined domain level feature set, and host level features joined into a joined host level feature set, wherein a domain level feature corresponds to a feature of a domain associated with a target document, and wherein a host level feature corresponds to a feature of a host associated with a target document;

extract numerical features and content features from ground truth documents and random documents, wherein a numerical feature corresponds to a ratio of an amount of a first type of content within a target document to an amount of a second type of content within the target document;

join the numerical features with the one or more levels of feature sets to create a set of joined features for the ground truth documents and the random documents;

train a document scoring model utilizing machine learning to score documents using the set of joined features;

score documents with document scores using the document scoring model based upon the content features and the set of joined features with document scores obtained during training; and

selectively index a subset of the documents based upon the document scores of the documents.

20. A computing device comprising:

a memory comprising instructions; and

a processor coupled to the memory, the processor configured to execute the instructions to cause the processor to:

combine feature types within one or more levels of feature sets into a joined host level feature set, wherein the one or more levels of feature sets comprise page level features joined into a joined page level feature set, domain level features joined into a joined domain level feature set, and host level features joined into a joined host level feature set, wherein a domain level feature corresponds to a feature of a domain associated with a target document, and wherein a host level feature corresponds to a feature of a host associated with a target document;

extract numerical features and content features from ground truth documents and random documents, wherein a numerical feature corresponds to a ratio of an amount of a first type of content within a target document to an amount of a second type of content within the target document;

join the numerical features with the one or more levels of feature sets to create a set of joined features for the ground truth documents and the random documents;

train a document scoring model utilizing machine learning to score documents using the set of joined features;

score documents with document scores using the document scoring model based upon the content features and the set of joined features with document scores obtained during training; and

selectively index a subset of the documents based upon the document scores of the documents.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 12, 2022
From: NTENT, INC.
To: SEEKR TECHNOLOGIES INC.
Reel/Frame 061394/0525 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S NAME AND ENTITY TYPE INSIDE THE ASSIGNMENT DOCUMENT AND ON THE COVER SHEET PREVIOUSLY RECORDED AT REEL: 050459 FRAME: 0460. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 6, 2022
From: BAEZA-YATES, RICARDO; CAMBAZOGLU, BERKANT BARLA; SHANKARALINGAPPA, DARSHAN MALLENAHALLI; CATENA, MATTEO
To: NTENT, INC.
Reel/Frame 061697/0277 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2019
From: BAEZA-YATES, RICARDO; CAMBAZOGLU, BERKANT BARLA; SHANKARALINGAPPA, DARSHAN MALLENAHALLI; CATENA, MATTEO
To: NTENT
Reel/Frame 050459/0458 →
Continuity (1)
Related Publication 20210089543A1 · Mar 25, 2021