IP Library Granted Patent US 9,760,570
Granted Patent B2
US 9,760,570 · App. 14/300,148 · Granted Sep 12, 2017

Finding and disambiguating references to entities on web pages

Inventors: Leonardo A. Laroco, Jr. (Philadelphia, PA); Nikola Jevtic (Newark, NJ); Nikolai V. Yakovenko (New York, NY); Jeffrey Reynar (New York, NY)
Assignee: Google Inc.
G06F17/30011G06F17/30876G06N5/04G06N99/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,760,570
App. No.
14/300,148
Granted
Sep 12, 2017
Kind
B2
Abstract

A system and method for disambiguating references to entities in a document. In one embodiment, an iterative process is used to disambiguate references to entities in documents. An initial model is used to identify documents referring to an entity based on features contained in those documents. The occurrence of various features in these documents is measured. From the number occurrences of features in these documents, a second model is constructed. The second model is used to identify documents referring to the entity based on features contained in the documents. The process can be repeated, iteratively identifying documents referring to the entity and improving subsequent models based on those identifications. Additional features of the entity can be extracted from documents identified as referring to the entity.

Claims (58)

1. A method for identifying texts referring to an entity, the method comprising:

at a computer having one or more processors and memory storing programs for execution by the one or more processors:

storing an object representing the entity;

storing a plurality of facts, wherein at least one of the plurality of facts is associated with the object;

determining a first set of features from the stored plurality of facts that are associated with the object, wherein the first set of features are sufficient for identifying a document referring to the entity;

determining a second set of features from the stored plurality of facts that are associated with the object, wherein

the second set of features are sufficient for identifying a document referring to the entity, and

the second set of features are distinct from the first set of features;

identifying a first text from one of the stored plurality of facts associated with the first set of features;

identifying a second text from one of the stored plurality of facts associated with the second set of features;

identifying a representative document as associated with the entity, wherein the first text and the second text are within the representative document; and

associating a fact selected from the representative document with the object.

2. The method of claim 1 , further comprising identifying, as associated with the entity, a third text distinct from the first text and the second text.

3. The method of claim 2 , further comprising extracting facts from the third text and identifying the facts as associated with the entity.

4. The method of claim 1 , wherein the first set of features is stored as a set of facts in a fact repository.

5. The method of claim 1 , wherein

the first text is identified using a first model; and

the second text is identified using a second model distinct from the first model.

6. The method of claim 5 , wherein the second model is selected in accordance with a number of occurrences of the first set of features in a document.

7. The method of claim 1 , wherein the second set of features includes at least one feature not included in the first set of features.

8. The method of claim 1 , wherein the first set of features includes at least one feature not included in the second set of features.

9. The method of claim 1 , further comprising storing at least one feature of the second set of features as a fact in a fact repository.

10. The method of claim 1 , further comprising estimating importance of the entity.

11. The method of claim 1 , further comprising estimating importance of the entity based on an estimated importance of a portion of the second text.

12. The method of claim 1 , further comprising associating at least one document with the entity.

13. The method of claim 1 , wherein the identifying of the second text comprises estimating a probability that a portion of the second text refers to the entity.

14. The method of claim 1 ,

wherein the first set of features comprises at least a first feature and a second feature, and

wherein a second model specifies that an occurrence of the first feature is sufficient to identify a document referring to the entity.

15. The method of claim 14 , wherein the second model specifies that an occurrence of the second feature is not sufficient to identify a document referring to the entity.

16. The method of claim 1 , wherein the representative document is an audio file.

17. A system for identifying texts referring to an entity, the system comprising one or more instructions for:

storing an object representing the entity;

storing a plurality of facts, wherein at least one of the plurality of facts is associated with the object;

determining a first set of features from the stored plurality of facts that are associated with the object, wherein the first set of features are sufficient for identifying a document referring to the entity;

determining a second set of features from the stored plurality of facts that are associated with the object, wherein

the second set of features are sufficient for identifying a document referring to the entity, and

the second set of features are distinct from the first set of features;

identifying a first text from one of the stored plurality of facts associated with the first set of features;

identifying a second text from one of the stored plurality of facts associated with the second set of features;

identifying a representative document as associated with the entity, wherein the first text and the second text are within the representative document; and

associating a fact selected from the representative document with the object.

18. The system of claim 17 , wherein the one or more instructions further comprise instructions for identifying, as associated with the entity, a third text distinct from the first text and the second text.

19. The system of claim 17 , wherein the one or more instructions further comprise instructions for extracting facts from a third text and identifying the facts as associated with the entity.

20. The system of claim 17 , wherein the representative document is an audio file.

21. A non-transitory computer readable storage medium storing one or more programs configured for execution by a computer, the one or more programs comprising instructions for:

storing an object representing the entity;

storing a plurality of facts, wherein at least one of the plurality of facts is associated with the object;

determining a first set of features from the stored plurality of facts that are associated with the object, wherein the first set of features are sufficient for identifying a document referring to the entity;

determining a second set of features from the stored plurality of facts that are associated with the object, wherein

the second set of features are sufficient for identifying a document referring to the entity, and

the second set of features are distinct from the first set of features;

identifying a first text from one of the stored plurality of facts associated with the first set of features;

identifying a second text from one of the stored plurality of facts associated with the second set of features;

identifying a representative document as associated with the entity, wherein the first text and the second text are within the representative document; and

associating a fact selected from the representative document with the object.

22. The non-transitory computer readable storage medium of claim 21 , wherein the one or more programs further comprise instructions for identifying, as associated with the entity, a third text distinct from the first text and the second text.

23. The non-transitory computer readable storage medium of claim 21 , wherein the representative document is an audio file.

Assignments (1)
CHANGE OF NAME Recorded Dec 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044695/0115 →
Continuity (3)
Continuation 13364244 · Feb 1, 2012
Continuation 11551657 · Oct 20, 2006
Related Publication 20140289177A1 · Sep 25, 2014