IP Library Granted Patent US 11,775,596
Granted Patent B1
US 11,775,596 · App. 17/726,482 · Granted Oct 3, 2023

Models for classifying documents

Inventors: Ashutosh Joshi (Fremont, CA); Martin Betz (Palo Alto, CA); Rajiv Arora (Gurgaon, IN); Rakesh Kumar Srivastava (New Delhi, IN); David Cooke (Los Altos, CA)
Assignee: Aurea Software, Inc.
G06F16/951G06F7/08G06F16/353
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,596
App. No.
17/726,482
Granted
Oct 3, 2023
Kind
B1
Abstract

Some embodiments provide a method for defining a content relevance model for determining whether a content segment is relevant to a particular category. The method receives a first set of content segments that contain content relevant to the particular category and a second set of content segments that contain content not relevant to the particular category. The method identifies a set of key word sets more likely to appear in the first set of content segments than the second set of content segments. The method defines a content relevance model that comprises a set of groups of word sets and a score for each group, each of the groups of word sets comprising a key word set from the set of key word sets and at least one word set found in a context of the key word set in at least one of the received content segments.

Claims (39)

1. A method for defining and utilizing a content relevance model for a particular category to at least in part determine whether a content segment is relevant to the particular category, the method comprising:

executing code by a processor for a first computer system to cause the processor of the first computer system to perform operations comprising:

sending a first set of content segments that contain content relevant to the particular category and a second set of content segments that contain content not relevant to the particular category; and

receiving, from a second computer system, documents that are relevant to one or more particular categories, wherein the documents are identified by a second computer system executing code by a processor to perform operations comprising:

receiving the first set of content segments;

identifying a set of key word sets more likely to appear in the first set of content segments than the second set of content segments; and

defining a content relevance model that comprises a set of groups of word sets and a score for each group, each of the groups of word sets comprising a key word set from the set of key word sets and at least one word set found in a context of the key word set in at least one of the received content segments, wherein defining the content relevance model further comprises:

determining the set of key word sets for the particular category based on an analysis of (i) a first set of content segments defined as relevant to the particular category and (ii) a second set of content segments defined as not relevant to the particular category;

determining (i) a set of pairs of word sets that each comprise a key word set and a word set that appears in a defined context of the keyword and (ii) a score for each of the word set pairs, the score for a particular word set pair quantifying a likelihood that a content segment containing the word set pair is relevant to the particular category; and

defining a content relevance model for the particular category, the content relevance model comprising (i) a context definition and (ii) the set of word set pairs and corresponding scores;

utilizing the content relevance model in a system to identify content segments in documents for relevancy to the one or more particular categories;

providing the documents that are relevant to the one or more particular categories to the first computer system.

2. The method of claim 1 , wherein the content segments comprise text documents.

3. The method of claim 1 , wherein the particular category comprises one of a company, product, person, industry, or concept.

4. The method of claim 1 , wherein identifying the set of key word sets comprises:

calculating a score for each word set that appears in at least one content segment of the first or second sets of content segments;

identifying a plurality of word sets with the highest scores as the key word sets.

5. The method of claim 4 , wherein calculating a score for a particular word set comprises comparing a probability of finding the particular word set in the first set of content segments and a probability of finding the particular word set in the second set of content segments.

6. The method of claim 5 , wherein calculating the score for the particular word set further comprises accounting for a number of occurrences of the particular word set in the first set of content segments.

7. The method of claim 1 , wherein defining the content relevance model comprises:

identifying the set of groups of word sets;

calculating a score for each group of word sets; and

storing (i) the set of groups of word sets (ii), the calculated scores, and (iii) a set of model parameters in the content relevance model.

8. The method of claim 7 , wherein the content relevance model is stored as a text file.

9. The method of claim 1 , wherein the groups of word sets are pairs of word sets.

10. A method for defining and utilizing a content relevance model for a particular category, the method comprising:

executing code by a processor for a first computer system to cause the processor of the first computer system to perform operations comprising:

determining a set of key word sets for the particular category based on an analysis of (i) a first set of content segments defined as relevant to the particular category and (ii) a second set of content segments defined as not relevant to the particular category;

determining (i) a set of pairs of word sets that each comprise a key word set and a word set that appears in a defined context of the keyword and (ii) a score for each of the word set pairs, the score for a particular word set pair quantifying a likelihood that a content segment containing the word set pair is relevant to the particular category;

defining a content relevance model for the particular category, the content relevance model comprising (i) a context definition and (ii) the set of word set pairs and corresponding scores; and

utilizing the content relevance model in a system to identify content segments in documents for relevancy to one or more particular categories.

11. The method of claim 10 , wherein the defined context of a key word set is a context defined for the content relevance model.

12. The method of claim 10 , wherein a word set appears in the defined context of a key word set when the word set is within a particular number of words surrounding the key word set in a content segment.

13. The method of claim 10 , wherein a word set appears in the defined context of a key word set when the word set is in the same sentence in a content segment as the key word set.

14. The method of claim 10 , wherein a word set appears in the defined context of a key word set when the word set is in the same paragraph in a content segment as the key word set.

15. The method of claim 10 , wherein determining a score for a particular word set pair comprises comparing a function calculated for the particular word set pair in the first set of content segments to the same function calculated for the particular word set pair in the second set of content segments.

16. The method of claim 15 , wherein the particular word set pmr comprises a particular key word set, wherein calculating the function for the particular word set pair in a particular content segment comprises:

comparing a number of occurrences of the particular word set pair in the particular set of content segments to a number of occurrences of the particular key word set in the particular set of content segments; and

comparing a number of content segments in the particular set of content segments in which the particular word set pair appears to a number of content segments in the particular set of content segments in which the particular key word set appears.

Assignments (1)
SECURITY INTEREST Recorded May 31, 2022
From: AUREA SOFTWARE, INC.; NEXTDOCS CORPORATION,; MESSAGEONE, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 060220/0673 →
Continuity (4)
Continuation 16691963 · Nov 22, 2019
Continuation 15662271 · Jul 27, 2017
Continuation 12772168 · Apr 30, 2010
Provisional Application 61316824 · Mar 23, 2010
Cited By (1)
US 12,561,355