IP Library Granted Patent US 10,942,963
Granted Patent B1
US 10,942,963 · App. 15/946,400 · Granted Mar 9, 2021

Method and system for generating topic names for groups of terms

Inventors: Bei Huang (Mountain View, CA); Nhung Ho (Redwood City, CA); Meng Chen (Mountain View, CA)
Assignee: Intuit Inc.
G06F16/355G06F16/313G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,942,963
App. No.
15/946,400
Granted
Mar 9, 2021
Kind
B1
Abstract

The invention relates to a method for generating a topic name for accounts grouped under a topic. The method includes obtaining account names associated with the accounts, generating a plurality of n-grams from the account names, and for each n-gram, obtaining a quality score based on relevance and meaningfulness of the n-gram. The relevance is determined using at least one relevance score that reflects how representative the n-gram is for the plurality of n-grams, and the meaningfulness is determined based on whether the n-gram exists in a knowledge base. The method further includes assigning one or more of the n-grams to the topic as the topic name, based on the quality score generated for each of the n-grams.

Claims (79)

1. A method for generating a topic name for accounts grouped under a topic, the method comprising:

obtaining account names associated with the accounts;

generating a plurality of n-grams from the account names;

for each n-gram, generating a quality score based on relevance and meaningfulness of the n-gram,

wherein the relevance is determined using at least one relevance score that reflects how representative the n-gram is for the plurality of n-grams, and

wherein the meaningfulness is determined based on whether the n-gram exists in a knowledge base; and

assigning one or more of the n-grams to the topic as the topic name, based on the quality score generated for each of the plurality of n-grams, wherein assigning comprises one selected from the group consisting of:

assigning the n-gram with a highest quality score to the topic as the topic name,

assigning a plurality of n-grams associated with quality scores above a threshold to the topic as the topic name, and

assigning a set number of n-grams associated with highest quality scores to the topic as the topic name.

2. The method of claim 1 , wherein determining the relevance of the n-gram comprises obtaining at least one statistical measure resulting in the at least one relevance score for the n-gram, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams.

3. The method of claim 2 , wherein the at least one statistical measure is based on at least one selected from a group consisting of a frequency of the n-gram, a term frequency-inverse document frequency (TF-IDF), and a mutual information.

4. The method of claim 1 , wherein determining the meaningfulness of the n-gram comprises:

labeling each of the n-grams as either “quality” or “non-quality”, based on whether the n-gram exists in the knowledge base;

obtaining at least one statistical measure resulting in the at least one relevance score for each of the n-grams, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams;

training classifiers using the statistical measures as features and using the “quality”/“non-quality” labels as ground truth; and

applying the trained classifiers to the n-gram to obtain the quality score.

5. The method of claim 4 , wherein the classifiers are based on one selected from a group consisting of logistic regressions, random forests and gradient boost machines.

6. The method of claim 4 , wherein the quality score comprises a ratio of predictions by the trained classifiers indicating “quality” and all predictions by the trained classifiers, resulting from applying the trained classifiers to the n-gram.

7. The method of claim 1 , wherein the knowledge base comprises expert knowledge in a discipline related to the accounts.

8. The method of claim 1 , wherein the knowledge base is one selected from a group consisting of a text document, a spreadsheet, a database, and a web page.

9. A system for generating a topic name for accounts grouped under a topic, the system comprising:

an account repository storing the accounts;

an account identification repository storing account names associated with the accounts;

a hardware processor and memory; and

software instructions stored in the memory, which when executed by the hardware processor, cause the hardware processor to:

obtain the account names associated with the accounts from the account identification repository;

generate a plurality of n-grams from the account names;

for each n-gram, generate a quality score based on relevance and meaningfulness of the n-gram,

wherein the relevance is determined using at least one relevance score that reflects how representative the n-gram is for the plurality of n-grams, and

wherein the meaningfulness is determined based on whether the n-gram exists in a knowledge base; and

assign one or more of the n-grams to the topic as the topic name, based on the quality score generated for each of the plurality of n-grams, wherein the assigning comprises performing one selected from the group consisting of:

assigning the n-gram with a highest quality score to the topic as the topic name,

assigning a plurality of n-grams associated with quality scores above a threshold to the topic as the topic name, and

assigning a set number of n-grams associated with highest quality scores to the topic as the topic name.

10. The system of claim 9 , wherein determining the relevance of the n-gram comprises obtaining at least one statistical measure resulting in the at least one relevance score for the n-gram, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams.

11. The system of claim 10 , wherein the at least one statistical measure is based on at least one selected from a group consisting of a frequency of the n-gram, a term frequency-inverse document frequency (TF-IDF), and a mutual information.

12. The system of claim 9 , wherein determining the meaningfulness of the n-gram comprises:

labeling each of the n-grams as either “found” or “not found”, based on whether the n-gram exists in the knowledge base;

obtaining at least one statistical measure resulting in the at least one relevance score for each of the n-grams, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams;

training classifiers using the statistical measures as features and using the “found”/“not found” labels as ground truth; and

applying the trained classifiers to the n-gram to obtain the quality score.

13. The system of claim 9 , wherein the knowledge base comprises expert knowledge in a discipline related to the accounts.

14. The system of claim 9 , wherein the knowledge base is one selected from a group consisting of a text document, a spreadsheet, a database, and a web page.

15. A method for generating a topic name for accounts grouped under a topic, the method comprising:

obtaining account names associated with the accounts;

generating a plurality of n-grams from the account names;

for each n-gram, generating a quality score based on relevance and meaningfulness of the n-gram,

wherein the relevance is determined using at least one relevance score that reflects how representative the n-gram is for the plurality of n-grams,

wherein the meaningfulness is determined based on whether the n-gram exists in a knowledge base, and

wherein determining the meaningfulness of the n-gram comprises:

labeling each of the n-grams as either “quality” or “non-quality”, based on whether the n-gram exists in the knowledge base;

obtaining at least one statistical measure resulting in the at least one relevance score for each of the n-grams, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams;

training classifiers using the statistical measures as features and using the “quality”/“non-quality” labels as ground truth; and

applying the trained classifiers to the n-gram to obtain the quality score; and

assigning one or more of the n-grams to the topic as the topic name, based on the quality score generated for each of the plurality of n-grams.

16. The method of claim 15 , wherein determining the relevance of the n-gram comprises obtaining at least one statistical measure resulting in the at least one relevance score for the n-gram, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams.

17. The method of claim 16 , wherein the at least one statistical measure is based on at least one selected from a group consisting of a frequency of the n-gram, a term frequency-inverse document frequency (TF-IDF), and a mutual information.

18. The method of claim 16 , wherein the classifiers are based on one selected from a group consisting of logistic regressions, random forests and gradient boost machines.

19. The method of claim 16 , wherein the quality score comprises a ratio of predictions by the trained classifiers indicating “quality” and all predictions by the trained classifiers, resulting from applying the trained classifiers to the n-gram.

20. A system for generating a topic name for accounts grouped under a topic, the system comprising:

an account repository storing the accounts;

an account identification repository storing account names associated with the accounts;

a hardware processor and memory; and

software instructions stored in the memory, which when executed by the hardware processor, cause the hardware processor to:

obtain the account names associated with the accounts from the account identification repository;

generate a plurality of n-grams from the account names;

for each n-gram, generate a quality score based on relevance and meaningfulness of the n-gram,

wherein the relevance is determined using at least one relevance score that reflects how representative the n-gram is for the plurality of n-grams, and

wherein the meaningfulness is determined based on whether the n-gram exists in a knowledge base;

wherein determining the meaningfulness of the n-gram comprises:

labeling each of the n-grams as either “found” or “not found”, based on whether the n-gram exists in the knowledge base;

obtaining at least one statistical measure resulting in the at least one relevance score for each of the n-grams, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams;

training classifiers using the statistical measures as features and using the “found”/“not found” labels as ground truth; and

applying the trained classifiers to the n-gram to obtain the quality score; and

assign one or more of the n-grams to the topic as the topic name, based on the quality score generated for each of the plurality of n-grams.

21. The system of claim 20 , wherein determining the relevance of the n-gram comprises obtaining at least one statistical measure resulting in the at least one relevance score for the n-gram, based on how relevant the n-gram is in view of all other n-grams of the plurality of n-grams.

22. The system of claim 21 , wherein the at least one statistical measure is based on at least one selected from a group consisting of a frequency of the n-gram, a term frequency-inverse document frequency (TF-IDF), and a mutual information.

23. The system of claim 20 , wherein the knowledge base comprises expert knowledge in a discipline related to the accounts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2018
From: HUANG, BEI; HO, NHUNG; CHEN, MENG
To: INTUIT INC.
Reel/Frame 045483/0927 →
Cited By (2)
US 12,430,437 US 12,432,225