IP Library Granted Patent US 10,726,055
Granted Patent B2
US 10,726,055 · App. 15/482,179 · Granted Jul 28, 2020

Multi-term query subsumption for document classification

Inventor: Nick Pendar (San Ramon, CA)
Assignee: GROUPON, INC.
G06F16/3322G06F16/35G06F16/93G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,726,055
App. No.
15/482,179
Granted
Jul 28, 2020
Kind
B2
Abstract

In general, embodiments of the present invention provide systems, methods and computer readable media for generating an optimal classifying query set for categorizing and/or labeling textual data based on a query subsumption calculus to determine, given two queries, whether one of the queries subsumes another. In one aspect, a method includes generating a group of determining queries based on analyzing text within a document; receiving a group of classifying queries; and, for each determining query within the group of determining queries, determining whether at least one of the classifying queries is subsumed by the determining query; and updating the group of classifying queries in an instance in which the classifying query is subsumed by the determining query.

Claims (29)

1. A computer-implemented method, comprising:

generating a determining queries group based on analyzing document text within a document, wherein each determining query within the determining queries group includes at least one term identified within the document text;

receiving a classifying queries group, the classifying queries group representing a particular category;

updating the classifying queries group in an instance in which a classifying query is subsumed by a determining query, wherein subsuming is defined by the determining query and the classifying query sharing at least two common terms;

identifying a classifying queries subsuming subset, the classifying queries subsuming subset being a subset of the classifying queries group, wherein each classifying query of the classifying queries subsuming subset each subsumes at least one of the determining queries of the determining queries group;

calculating a document categorization score based on a normalized sum of respective performance metrics for respective classifying queries of the classifying queries subsuming subset; and

associating the document with the particular category represented by the classifying queries group in an instance in which the document categorization score is greater than a categorization threshold value.

2. The computer-implemented method of claim 1 , wherein the respective performance metrics are respective binormal separation scores for the respective classifying queries of the classifying queries subsuming subset.

3. The computer-implemented method of claim 2 , wherein the respective binormal separation scores are calculated based on training data used to generate the classifying queries group.

4. The computer-implemented method of claim 1 , wherein the categorization threshold value is computed through cross-validation at a time of training a machine learning model.

5. The computer-implemented method of claim 1 , wherein the document categorization threshold value represents a score threshold that yields a minimum desired performance metric.

6. The computer-implemented method of claim 1 , wherein the document categorization threshold value is optimized for an F-score.

7. The computer-implemented method of claim 2 , wherein the binormal separation score of a classifying query represents how well the classifying query separates positive documents from negative documents in a corpus.

8. The computer-implemented method of claim 7 , wherein the positive documents have been categorized as belonging to a category represented by a feature set and the negative documents have been categorized as belonging to a different category than the category represented by the feature set.

9. A system, comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

generating a determining queries group based on analyzing document text within a document, wherein each determining query within the determining queries group includes at least one term identified within the document text;

receiving a classifying queries group, the classifying queries group representing a particular category;

updating the classifying queries group in an instance in which a classifying query is subsumed by a determining query, wherein subsuming is defined by the determining query and the classifying query sharing at least two common terms;

identifying a classifying queries subsuming subset, the classifying queries subsuming subset being a subset of the classifying queries group, wherein each classifying query of the classifying queries subsuming subset each subsumes at least one of the determining queries of the determining queries group;

calculating a document categorization score based on a normalized sum of respective performance metrics for respective classifying queries of the classifying queries subsuming subset; and

associating the document with the particular category represented by the classifying queries group in an instance in which the document categorization score is greater than a categorization threshold value.

10. They system of claim 9 , wherein the respective performance metrics are respective binormal separation scores for the respective classifying queries of the classifying queries subsuming subset.

11. They system of claim 10 , wherein the respective binormal separation scores are calculated based on training data used to generate the classifying queries group.

12. They system of claim 9 , wherein the categorization threshold value is computed through cross-validation at a time of training a machine learning model.

13. They system of claim 9 , wherein the document categorization threshold value represents a score threshold that yields a minimum desired performance metric.

14. They system of claim 9 , wherein the document categorization threshold value is optimized for an F-score.

15. They system of claim 10 , wherein the binormal separation score of a classifying query represents how well the classifying query separates positive documents from negative documents in a corpus.

16. They system of claim 15 , wherein the positive documents have been categorized as belonging to a category represented by a feature set and the negative documents have been categorized as belonging to a different category than the category represented by the feature set.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2024
From: GROUPON, INC.
To: BYTEDANCE INC.
Reel/Frame 068833/0811 →
RELEASE OF SECURITY INTEREST Recorded Feb 26, 2024
From: JPMORGAN CHASE BANK, N.A.
To: GROUPON, INC.; LIVINGSOCIAL, LLC (F/K/A LIVINGSOCIAL, INC.)
Reel/Frame 066676/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RIGHTS Recorded Feb 26, 2024
From: JPMORGAN CHASE BANK, N.A.
To: GROUPON, INC.; LIVINGSOCIAL, LLC (F/K/A LIVINGSOCIAL, INC.)
Reel/Frame 066676/0251 →
SECURITY INTEREST Recorded Jul 23, 2020
From: GROUPON, INC.; LIVINGSOCIAL, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 053294/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 12, 2017
From: PENDAR, NICK
To: GROUPON, INC.
Reel/Frame 044368/0015 →
Continuity (3)
Continuation 15198461 · Jun 30, 2016
Continuation 14038644 · Sep 26, 2013
Related Publication 20180032532A1 · Feb 1, 2018