IP Library Granted Patent US 10,713,293
Granted Patent B2
US 10,713,293 · App. 16/200,165 · Granted Jul 14, 2020

Method and system of computer-processing one or more quotations in digital texts to determine author associated therewith

Inventor: Yaroslav Victorovich Akulov (Moscow, RU)
Assignee: YANDEX EUROPE AG
G06F16/38G06F16/35G06F16/9535G06F40/205
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,713,293
App. No.
16/200,165
Granted
Jul 14, 2020
Kind
B2
Abstract

A method and system for computer-processing quotations in digital text to determine an author associated therewith is disclosed. The method comprises receiving a plurality of digital texts. The plurality of digital texts are parsed to extract one or more quotations. At least one candidate authors are identified for each of the digital text. A quotation similarity value for a given quotation with respect to each of the remaining one or more quotations is assigned. A quotation cluster is generated, which comprises one or more similar quotations and a set of candidate authors. The set of candidate authors is analyzed to identify a given candidate author meeting a condition. The candidate author meeting the condition is stored as the author of the one or more similar quotations.

Claims (66)

1. A computer implemented method for processing one or more quotations in digital texts to determine an author associated therewith, the method executable by a server configured to execute a news aggregator service, the server being coupled to a plurality of digital news services via a communication network, the method comprising:

receiving a plurality of digital texts from a database;

parsing each of the plurality of digital texts to extract one or more quotations therefrom, the parsing being executed by applying one or more parsing rules;

identifying at least one associated candidate author for each of the one or more quotations, the identifying being executed by applying one or more identification rules;

assigning, by a first classifier, a quotation similarity value for a given quotation with respect to each of a remaining one or more quotations, the quotation similarity value being representative of a likelihood of the given quotation originating from a same quotation with respect to each of the remaining one or more quotations;

generating a quotation cluster, the quotation cluster comprising:

one or more similar quotations, the one or more similar quotations comprising the given quotation and a subset of the remaining one or more quotations each having the similarity value above a threshold;

a set of candidate authors, the set of candidate authors comprising at least one candidate author associated with each of the one or more similar quotations;

analyzing the set of candidate authors to identify a given candidate author meeting a condition; and

storing the candidate author meeting the condition as the author of the one or more similar quotations.

2. The method of claim 1 , wherein the digital texts correspond to news articles representative of a same topic, wherein the plurality of news articles are received from the plurality of digital news services.

3. The method of claim 1 , wherein the one or more parsing rules comprise extracting one or more portion of digital texts interposed between a given set of quotation marks.

4. The method of claim 3 , wherein the one or more identification rules comprise identifying at least one capitalized word within a predetermined distance from the given set of quotation marks.

5. The method of claim 1 , wherein:

the given quotation is a first quotation;

the threshold is a first threshold; and

assigning the quotation similarity value to the first quotation in respect to a second quotation comprises:

determining a shortest common consecutive string of words between the first quotation and the second quotation; and

determining if a length of the shortest common consecutive string of words is above a second threshold.

6. The method of claim 5 , wherein the quotation similarity value comprises a binary value.

7. The method of claim 1 , wherein analyzing the set of candidate authors to identify the given candidate author meeting the condition comprises:

determining a frequency of occurrence of the given candidate author within the set of candidate authors; and

determining that the frequency of occurrence of the given candidate author is a highest frequency within the set of candidate authors.

8. The method of claim 7 , wherein:

the server is further coupled to an image database, the image database comprising a plurality of images associated with at least one of the one or more candidate authors; and

the method further comprises:

prior to transmitting the data packet, retrieving an image associated with the author; and

wherein the data packet further comprises the image.

9. The method of claim 1 , further comprising:

in response to a client device accessing the news aggregator service, transmitting to the client device a data packet, the data packet comprising:

a best quotation corresponding to one of the one or more similar quotations; and

the author of the best quotation.

10. The method of claim 9 , wherein the best quotation corresponds to one of the one or more similar quotations having a longest string of consecutive words.

11. A server for processing one or more quotations in digital texts to determine an author associated therewith, the server being coupled to a plurality of digital news services via a communication network, the server comprising a processor configured to:

receive a plurality of digital texts from a database;

parse each of the plurality of digital texts to extract one or more quotations therefrom, the parsing being executed by applying one or more parsing rules;

identify at least one associated candidate author for each of the one or more quotations, the identifying being executed by applying one or more identification rules;

assign, by a first classifier, a quotation similarity value for a given quotation with respect to each of a remaining one or more quotations, the quotation similarity value being representative of a likelihood of the given quotation originating from a same quotation with respect to each of the remaining one or more quotations;

generate a quotation cluster, the quotation cluster comprising:

one or more similar quotations, the one or more similar quotations comprising the given quotation and a subset of the remaining one or more quotations each having the similarity value above a threshold;

a set of candidate authors, the set of candidate authors comprising at least one candidate author associated with each of the one or more similar quotations;

analyze the set of candidate authors to identify a given candidate author meeting a condition; and

store the candidate author meeting the condition as the author of the one or more similar quotations.

12. The server of claim 11 , wherein the digital texts correspond to news articles representative of a same topic, wherein the plurality of news articles are received from the plurality of digital news services.

13. The server of claim 11 , wherein the one or more parsing rules comprises extracting one or more portion of digital texts interposed between a given set of quotation marks.

14. The server of claim 13 , wherein the one or more identification rules comprise identifying at least one capitalized word within a predetermined distance from the given set of quotation marks.

15. The server of claim 11 , wherein:

the given quotation is a first quotation;

the threshold is a first threshold; and

to assign the quotation similarity value to the first quotation in respect to a second quotation, the processor is configured to:

determine a shortest common consecutive string of words between the first quotation and the second quotation; and

determine if a length of the shortest common consecutive string of words is above a second threshold.

16. The server of claim 15 , wherein the quotation similarity value comprises a binary value.

17. The server of claim 11 , wherein to analyze the set of candidate authors to identify the given candidate author meeting the condition, the processor is configured to:

determine a frequency of occurrence of the given candidate author within the set of candidate authors; and

determine that the frequency of occurrence of the given candidate author is a highest frequency within the set of candidate authors.

18. The server of claim 17 , wherein:

the server is further coupled to an image database, the image database comprising a plurality of images associated with at least one of the one or more candidate authors; and

the processor is further configured to:

prior to transmitting the data packet, retrieve an image associated with the author; and

wherein the data packet further comprises the image.

19. The server of claim 11 , wherein the processor is further configured to:

in response to a client device accessing the news aggregator service, transmit to the client device a data packet, the data packet comprising:

a best quotation corresponding to one of the one or more similar quotations; and

the author of the best quotation.

20. The server of claim 19 , wherein the best quotation corresponds to one of the one or more similar quotations having a longest string of consecutive words.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0384 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2019
From: AKULOV, YAROSLAV VICTOROVICH
To: YANDEX LLC
Reel/Frame 048522/0838 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 6, 2019
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 048522/0884 →