IP Library Granted Patent US 9,613,024
Granted Patent B1
US 9,613,024 · App. 14/935,404 · Granted Apr 4, 2017

System and methods for creating datasets representing words and objects

Inventor: Guangsheng Zhang (Palo Alto, CA)
G06F17/2785G06F17/274G06F17/2705G06F17/277
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,613,024
App. No.
14/935,404
Granted
Apr 4, 2017
Kind
B1
Abstract

Systems and methods are presented for producing datasets as symbolic or associative representations of terms or objects or entities. A term can be a word or a phrase, which can also be the name of an object or a topic or concept. A dataset is produced for a specific term or object. The dataset contains a plurality of other terms or values associated with the specific term, and can serve as a representation of the specific term by other terms or values obtained using machine-based algorithms from text contents. Terms or values in the datasets also represent information about an object, or information about properties associated with the object. Methods for obtaining the datasets include discovering the relationships between terms in a plurality of text contents, based on occurrence, location, and attributes associated with the terms in the text contents.

Claims (41)

1. A computer system for producing a dataset for representing a term or an object, the system comprising:

one or more processors operable to

receive a first group of text contents comprising a plurality of text units;

receive, or identify from the text contents, a first term comprising a word or a phrase;

identify a text unit comprising a sentence or a phrase containing the first term and one or more second terms each comprising a word or a phrase;

identify a relation between the first term and one or more second terms in the text unit using a machine-based algorithm based on occurrence, or location, or attributes associated with the first term or the one or more second terms;

determine one or more numerical values to represent the relation or the strength of the relation between the first term and the corresponding one or more second terms;

collect one or more of the one or more numerical values into a group of numerical values;

associate the group of numerical values to the first term to form a dataset;

output the dataset as a representation of the first term or an object represented by the first term based on relations between the first term and terms other than the first term.

2. The system of claim 1 , wherein the one or more processors are further operable to

collect, based on the relation or based on the numerical values, one or more of the one or more second terms into a group of second terms;

associate the group of second terms to the first term to form the dataset.

3. The system of claim 2 , wherein the dataset is further used for providing a representation of the first term by other terms, or providing a representation of an object represented by the first term, wherein the object comprises a physical object or a conceptual object, wherein the group of second terms represent properties associated with the object.

4. The system of claim 2 , wherein at least one of the one or more second terms in the dataset is associated with one of the numerical values.

5. The system of claim 4 , wherein the at least one of the one or more second terms is collected based on the one of the numerical values.

6. The system of claim 4 , wherein the function of the one of the numerical values includes representing the strength of association between the at least one of the one or more second terms and the first term, or between a property or attribute represented by the at least one of the one or more second terms and the object represented by the first term.

7. The system of claim 1 , wherein the one or more numerical values are determined based on the number of text units that contain the first term or the one or more second terms, or the number of occurrences of the first term or the one or more second terms in the text units.

8. The system of claim 7 , wherein the one or more numerical values are determined further by dividing the one or more numerical values by the total number of text units in the first group.

9. The system of claim 1 , wherein the one or more numerical values are determined based on the location of the first term or the one or more second terms in the text units.

10. The system of claim 1 , wherein the one or more numerical values are determined based on whether the text unit is a phrase, a sentence, a paragraph, or a document containing a plurality of sentences or paragraphs.

11. The system of claim 1 , wherein the one or more numerical values are determined based on a grammatical attribute associated with the first term, wherein the grammatical attribute includes at least a subject or a predicate of a sentence, or a head or a modifier of a multi-word phrase, or a sub-component of a multi-word phrase.

12. A computer system for producing a dataset for representing a term or information related to an object, the system comprising:

one or more processors operable to

receive a first group of text contents comprising a plurality of text units;

receive, or identify from the text contents, a first term comprising a word or a phrase;

identify a text unit comprising a sentence or a phrase containing the first term and one or more second terms each comprising a word or a phrase;

identify a relation between the first term and one or more second terms in the text unit using a machine-based algorithm based on occurrence, or location, or attributes associated with the first term or the one or more second terms;

determine a strength measure of the relation between the first term and the corresponding one or more second terms;

collect, based on the relation and the strength measure, one or more of the one or more second terms into a group of second terms;

associate the group of second terms to the first term to form a dataset; and

output the dataset as a representation of the first term by other terms associated with the first term, or information associated with an object represented by the first term, wherein the object comprises a physical or conceptual entity, wherein the group of second terms represent properties associated with the object.

13. The system of claim 12 , wherein the one or more processors are further operable to produce

a first score to represent the strength measure based on the occurrence, location, or attributes associated with the first term or the one or more second terms, wherein the group of second terms are collected based on the first score.

14. The system of claim 13 , wherein the first score is produced based on the number of text units that contain the first term or the one or more second terms, or the number of occurrences of the first term or the one or more second terms in the text units.

15. The system of claim 14 , wherein the first score is produced further by dividing the first score by the total number of text units in the first group.

16. The system of claim 13 , wherein the first score is produced based on the location of the first term or the one or more second terms in the text units.

17. The system of claim 13 , wherein the first score is produced based on whether the text unit is a phrase, a sentence, a paragraph, or a document containing a plurality of sentences or paragraphs.

18. The system of claim 13 , wherein the first score is produced based on a grammatical attribute associated with the first term, wherein the grammatical attribute includes at least a subject or a predicate of a sentence, or a head or a modifier of a multi-word phrase, or a sub-component of a multi-word phrase.

19. The system of claim 13 , wherein the function of the first score includes representing the strength of association between the at least one second term and the first term, or between a property or attribute represented by the at least one second term and the object represented by the first term.

20. The system of claim 13 , wherein the first score is produced based on the occurrence or attributes associated with the one or more second terms in text units that do not contain the first term, or based on the number of text units that contain the one or more second terms but do not contain the first term.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2021
From: ZHANG, CHIZONG; ZHANG, GUANGSHENG
To: LINFO IP LLC
Reel/Frame 057128/0500 →
Continuity (3)
Continuation 13763716 · Feb 10, 2013
Continuation 12631829 · Dec 5, 2009
Provisional Application 61151729 · Feb 11, 2009