IP Library Granted Patent US 11,194,878
Granted Patent B2
US 11,194,878 · App. 16/571,870 · Granted Dec 7, 2021

Method of and system for generating feature for ranking document

Inventors: Aleksandr Valerievich Safronov (Moscow, RU); Vasily Vladimirovich Zavyalov (Tyumen, RU)
Assignee: YANDEX EUROPE AG
G06F16/9538G06F16/9532G06F16/9536G06K9/623G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,194,878
App. No.
16/571,870
Granted
Dec 7, 2021
Kind
B2
Abstract

A method and a system for ranking a document in response to a query, the document having no value for a given feature with respect to the query. A set of documents relevant to the query is generated. The document is selected, and a set of past queries having presented the document as a search result are retrieved. Respective values for the given feature for the document with respect to the set of past queries are retrieved. A respective similarity parameter is determined between the query and each of the set of past queries. The value of the given feature for the document is generated based at least in part on the respective similarity parameter and the respective value for the given feature of at least one past query. The set of documents including the document is ranked based in part on the given feature.

Claims (77)

1. A computer-implemented method for ranking at least one document in response to a given query using a machine learning algorithm (MLA) executed by a server, the method executable by the server, the server being connected to a search log database, the server being connected to an electronic device over a communication network, the method comprising:

receiving, by the server, the given query;

generating, by the server, a set of documents relevant to the given query, the set of documents having a plurality of features;

selecting, by the server, the at least one document from the set of documents, the at least one document having no respective value for a given feature of the plurality of features;

retrieving, from the search log database, a set of past queries having been submitted on the server, each past query of the set of past queries having presented the at least one document as a respective search result in a respective search engine results page (SERP);

retrieving, from the search log database, for each respective past query, a respective value of the given feature for the at least one document;

retrieving, from the search log database, a respective set of past documents for each respective past query of the set of past queries, the respective set of past documents having been presented as respective search results in response to the respective past query;

determining, by the server, a respective similarity parameter between the given query and each respective past query of the set of past queries, the determining the respective similarity parameter between the given query and each respective past query of the set of past queries being based on a degree of an overlap between:

the set of documents relevant to the given query, and

the respective set of documents of the respective past query;

generating, by the server, the respective value of the given feature for the at least one document based at least in part on:

the respective similarity parameter of at least one past query of the set of past queries, and

the respective value for the given feature of the at least one past query of the set of past queries;

ranking, by the MLA, the set of documents to obtain ranked list of documents, the ranking being based on the plurality of features, the at least one given document being ranked based at least in part on the respective value of the given feature; and

transmitting, to the electronic device, the ranked list of documents to be presented as a SERP.

2. The method of claim 1 , wherein

each respective document of the set of documents relevant to the given query has a respective annotation, the respective annotation including:

at least one respective past search query having been used to access the respective document on the search engine server; and wherein

the retrieving the set of past queries is based on the respective annotation of the at least one document.

3. The method of claim 1 , wherein

at least a subset of the set of documents is associated with respective user interaction parameters; and wherein

each respective document of the respective set of past documents for each respective query is associated with respective past user interaction parameters; and

wherein

the determining the respective similarity parameter is further based on:

the respective user interaction parameters of the respective query of the subset of documents, and

the respective user interaction parameters of the respective set of past documents.

4. The method of claim 1 , wherein the method further comprises, prior to the generating the respective value of the given feature:

selecting, by the server, the at least one respective past query based on the respective similarity parameter being above a predetermined threshold.

5. The method of claim 1 , wherein

the retrieving the respective value of the given feature of the at least one document for each respective past query further comprises:

retrieving a respective relevance score of the at least one document to the respective past query; and wherein

the generating the respective value of the given feature is further based on the respective relevance score.

6. The method of claim 5 , wherein the generating the respective value of the given feature for the at least one document is further based on:

a respective value of at least one other feature of the plurality of features for the given document.

7. The method of claim 6 , wherein the given feature is one of:

a query-dependent feature, and

a user interaction parameter.

8. A system for ranking at least one document in response to a given query using a machine learning algorithm (MLA) executed by the system, the system being connected to a search log database, the system being connected to an electronic device, the system comprising:

a processor;

a non-transitory computer-readable medium comprising instructions;

the processor, upon executing the instructions, being configured to:

receive the given query;

generate a set of documents relevant to the given query, the set of documents having a plurality of features;

select the at least one document from the set of documents, the at least one document having no respective value for a given feature of the plurality of features;

retrieve, from the search log database, a set of past queries having been submitted on the server, each past query of the set of past queries having presented the at least one document as a respective search result in a respective search engine results page (SERP);

retrieve, from the search log database, for each respective past query, a respective value of the given feature for the at least one document;

retrieve, from the search log database, a respective set of past documents for each respective past query of the set of past queries, the respective set of past documents having been presented as respective search results in response to the respective past query;

determine a respective similarity parameter between the given query and each respective past query of the set of past queries, the determining the respective similarity parameter between the given query and each respective past query of the set of past queries is based on a degree of an overlap between:

the set of documents relevant to the given query, and

the respective set of documents of the respective past query;

generate the respective value of the given feature for the at least one document based at least in part on:

the respective similarity parameter of at least one past query of the set of past queries, and

the respective value for the given feature of the at least one past query of the set of past queries;

rank, via the MLA, the set of documents to obtain ranked list of documents, the ranking being based on the plurality of features, the at least one given document being ranked based at least in part on the respective value of the given feature; and

transmit, to the electronic device, the ranked list of documents to be presented as a SERP.

9. The system of claim 8 , wherein

each respective document of the set of documents relevant to the given query has a respective annotation, the respective annotation including:

at least one respective past search query having been used to access the respective document on the search engine server; and

wherein the retrieving the set of past queries is based on the respective annotation of the at least one document.

10. The system of claim 8 , wherein

at least a subset of the set of documents is associated with respective user interaction parameters; and wherein

each respective document of the respective set of past documents for each respective query is associated with respective past user interaction parameters; and

wherein

the determining the respective similarity parameter is further based on:

the respective user interaction parameters of the respective query of the subset of documents, and

the respective user interaction parameters of the respective set of past documents.

11. The system of claim 8 , wherein the processor is further configured to, prior to the generating the respective value of the given feature:

select the at least one respective past query based on the respective similarity parameter being above a predetermined threshold.

12. The system of claim 8 , wherein

to retrieve the respective value of the given feature of the at least one document for each respective past query, the processor is further configured to:

retrieve a respective relevance score of the at least one document to the respective past query; and wherein

the generating the respective value of the given feature is further based on the respective relevance score.

13. The system of claim 12 , wherein the generating the respective value of the given feature for the at least one document is further based on:

a respective value of at least one other feature of the plurality of features for the given document.

14. The system of claim 13 , wherein the given feature is one of:

a query-dependent feature, and

a user interaction parameter.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0537 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: YANDEX.TECHNOLOGIES LLC
To: YANDEX LLC
Reel/Frame 051602/0382 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 23, 2020
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 051602/0485 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 21, 2020
From: SAFRONOV, ALEKSANDR VALERIEVICH; ZAVYALOV, VASILY VLADIMIROVICH
To: YANDEX.TECHNOLOGIES LLC
Reel/Frame 051572/0285 →
Cited By (1)
US 12,670,153