IP Library Granted Patent US 10,847,144
Granted Patent B1
US 10,847,144 · App. 14/855,165 · Granted Nov 24, 2020

Methods and apparatus for identification and analysis of temporally differing corpora

Inventors: Jens Erik Tellefsen (Mountain View, CA); Ranjeet Singh Bhatia (Sunnyvale, CA)
Assignee: NetBase Solutions, Inc.
G10L15/19G06F16/26G06F16/9535G06Q30/0201G06Q50/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,847,144
App. No.
14/855,165
Granted
Nov 24, 2020
Kind
B1
Abstract

Differences are identified, at the lexical unit and/or phrase level, between time-varying corpora. A corpus for a time period of interest is compared with a reference corpus. N-grams are generated for both the corpus of interest and reference corpus. Numbers of occurrences are counted. An average number of occurrences, for each n-gram of the reference corpus, is determined. A difference value, between number of occurrences in corpus of interest and average number of occurrences, is determined. Each difference value is normalized. N-grams can be selected for display, or for further processing, on the basis of the normalized difference value. Further processing can include selecting a sample period. A plurality of reference corpora are produced, where a begin time, for each sub-corpus of the plurality of reference corpora, differs, from a begin time for the corpus of interest, by an integer multiple of the sample period. Word Cloud visualization is shown.

Claims (56)

1. A method for identifying n-grams about an object, comprising:

identifying, as a result of computing hardware and programmable memory, an object-specific corpus, that is a subset of a first corpus, where approximately all statements of the object-specific corpus are about a same first object;

identifying, as a result of computing hardware and programmable memory, statements of the object-specific corpus, for inclusion in a corpus of interest, upon a basis of a statement relating to a time period of interest;

identifying, as a result of computing hardware and programmable memory, statements of the object-specific corpus, for inclusion in a reference corpus, upon a basis of a statement relating to a reference time period that is different from the time period of interest;

identifying for a selected list of n-grams, as a result of computing hardware and programmable memory, n-grams that appear in both the corpus of interest and the reference corpus;

identifying, as a result of computing hardware and programmable memory, for each n-gram of the selected list of n-grams, for subsequent access in conjunction with an n-gram, a number of occurrences of the n-gram in the corpus of interest;

identifying, as a result of computing hardware and programmable memory, for each n-gram of the selected list of n-grams, for subsequent access in conjunction with an n-gram, a number of occurrences of the n-gram in the reference corpus;

determining, as a result of computing hardware and programmable memory, for each n-gram of the selected list of n-grams, for subsequent access in conjunction with an n-gram, an average number of occurrences of the n-gram, in the reference-corpus;

determining, as a result of computing hardware and programmable memory, for each n-gram of the selected list of n-grams, for subsequent access in conjunction with an n-gram, a difference value, between a number of occurrences of the n-gram in the corpus of interest and an average number of occurrences of the n-gram in the reference-corpus;

normalizing, as a result of computing hardware and programmable memory, for each n-gram of the selected list of n-grams, for subsequent access in conjunction with an n-gram, the difference value to produce a normalized difference value; and

determining, as a result of computing hardware and programmable memory, from the selected list of n-grams, a second selected list of n-grams, on a basis of the normalized difference value; and displaying, as a result of computing hardware and programmable memory, n-grams of the second selected list to the user with a visualization technique.

2. The method of claim 1 , wherein a duration of the reference time period is equal to an integer times a duration of the time period of interest.

3. The method of claim 1 , wherein the step of determining a selected list of n-grams further comprises:

determining an n-gram, for inclusion in the selected list of n-grams, on a basis of a number of occurrences of the n-gram in the object-specific corpus.

4. The method of claim 1 , wherein the step of determining an average number of occurrences of an n-gram further comprises:

determining a first multiple that, when a time duration of the corpus of interest is multiplied by, produces a time duration represented by the reference corpus; and

using the first multiple to determine an average number of occurrences of an n-gram in the reference-corpus.

5. The method of claim 4 , wherein the step of using the first multiple further comprises:

dividing, for each n-gram of the selected list of n-grams, a number of occurrences of the n-gram in the reference corpus by the multiple.

6. The method of claim 1 , wherein the step of normalizing, to produce a normalized difference value, further comprises:

dividing, for each n-gram of the selected list of n-grams, a difference value by an average number of occurrences of the n-gram in the reference-corpus.

7. The method of claim 1 , wherein the step of normalizing, to produce a normalized difference value, further comprises:

dividing, for each n-gram of the selected list of n-grams, a difference value by a number of occurrences of the n-gram in the corpus of interest.

8. The method of claim 1 , wherein the step of normalizing, to produce a normalized difference value, further comprises:

dividing, for each n-gram of the selected list of n-grams, a difference value by a standard deviation of occurrences of the n-gram in the reference-corpus.

9. The method of claim 1 , wherein the step of determining a second selected list of n-grams further comprises:

ordering the selected list of n-grams on a basis of decreasing normalized difference value; and

selecting a first predetermined number of n-grams from the ordered selected list of n-grams.

10. The method of claim 1 , further comprising:

varying a first graphical dimension or characteristic, for a display of each n-gram of the second selected list of n-grams, on a basis of the normalized difference value of the n-gram.

11. The method of claim 10 , wherein the first graphical dimension is a font size, for a display of each n-gram of the second selected list of n-grams.

12. The method of claim 10 , further comprising:

varying a second graphical dimension or characteristic, for a display of each n-gram of the second selected list of n-grams, on a basis of a number of occurrences of the n-gram.

13. The method of claim 1 , further comprising:

producing a first Logical Form semantic representation for a first unit of natural language of the first corpus;

determining whether a first frame extraction rule matches the first Logical Form;

producing, if the first frame extraction rule matches, a first instance having at least object and sentiment roles, with values of the first Logical Form assigned to corresponding roles of the first instance; and

including the first instance in the first object-specific corpus if the user entered query matches a value assigned to the object role of the first instance.

14. The method of claim 12 , wherein the first graphical characteristic is a font transparency and the second graphical dimension is a font size, for a display of each n-gram of the second selected list of n-grams.

15. The method of claim 1 , further comprising:

selecting a sample period;

producing a plurality of reference corpora, wherein a begin time, for each sub-corpus of the plurality of reference corpora, differs, from a begin time for the corpus of interest, by an integer multiple of the sample period;

identifying, for each n-gram of the second selected list of n-grams and for each sub-corpus of the plurality of reference corpora, for subsequent access in conjunction with an n-gram, a number of occurrences of the n-gram in the sub-corpus;

determining, for each n-gram of the second selected list of n-grams, for subsequent access in conjunction with an n-gram, a second average number of occurrences of the n-gram, from a number of occurrences of the n-gram in each sub-corpus;

determining, for each n-gram of the second selected list of n-grams, for subsequent access in conjunction with an n-gram, a second difference value, between a number of occurrences of the n-gram in the corpus of interest and a second average number of occurrences of the n-gram;

normalizing, for each n-gram of the second selected list of n-grams, for subsequent access in conjunction with an n-gram, the second difference value to produce a second normalized difference value; and

determining, from the second selected list of n-grams, a third selected list of n-grams, on a basis of the second normalized difference value.

16. The method of claim 15 , wherein the step of normalizing, to produce a second normalized difference value, further comprises:

determining, for each n-gram of the second selected list of n-grams, a standard deviation from a number of occurrences of the n-gram in each sub-corpus;

dividing, for each n-gram of the second selected list of n-grams, a second difference value by a standard deviation.

17. The method of claim 15 , further comprising:

varying a first graphical dimension or characteristic, for a display of each n-gram of the third selected list of n-grams, on a basis of the second normalized difference value of the n-gram.

18. The method of claim 17 , wherein the first graphical dimension is a font size, for a display of each n-gram of the second selected list of n-grams.

19. The method of claim 17 , further comprising:

varying a second graphical dimension or characteristic, for a display of each n-gram of the second selected list of n-grams, on a basis of a number of occurrences of the n-gram.

20. The method of claim 19 , wherein the first graphical characteristic is a font transparency and the second graphical dimension is a font size, for a display of each n-gram of the second selected list of n-grams.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 24, 2021
From: ORIX GROWTH CAPITAL, LLC
To: NETBASE SOLUTIONS, INC.
Reel/Frame 058208/0292 →
SECURITY INTEREST Recorded Nov 18, 2021
From: NETBASE SOLUTIONS, INC.; QUID, LLC
To: EAST WEST BANK, AS ADMINISTRATIVE AGENT
Reel/Frame 058157/0091 →
SECURITY INTEREST Recorded Aug 31, 2018
From: NETBASE SOLUTIONS, INC.
To: ORIX GROWTH CAPITAL, LLC
Reel/Frame 046770/0639 →
Continuity (1)
Continuation 13836416 · Mar 15, 2013