IP Library Granted Patent US 10,762,992
Granted Patent B2
US 10,762,992 · App. 15/365,191 · Granted Sep 1, 2020

Synthetic ground truth expansion

Inventors: Matthew Kellar MacLeod (Denver, CO); Jacque W. Swartz (Loveland, CO); Marissa Victoria Ponder (Denver, CO)
Assignee: Welltok, Inc.
G16H50/20G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,762,992
App. No.
15/365,191
Granted
Sep 1, 2020
Kind
B2
Abstract

A ground truth expansion system that generates an expanded set of synthetic questions and selects a targeted subset of questions for machine learning training. The machine learning may be used to train an automated inquiry system that responds to questions received from individuals about subject matter of interest. The automated inquiry system is particularly suitable for use in, for example, responding to questions raised by insured individuals about their healthcare benefits.

Claims (64)

1. A computer-implemented method for generating a ground truth for training a text classifier, the method comprising:

retrieving a plurality of question forms and a plurality of intent forms, wherein each intent form includes an intent identifier, at least one intent synonyms, and answer information;

generating, based on the question forms and intent forms, an expanded question set comprised of question text and corresponding answers, wherein each question text-answer pair is associated with a topic;

constructing vectors characterizing the questions of the expanded question set by, for each question text:

dividing the question text into one or more text segments;

calculating metrics of significance for the one or more text segments; and

generating a vector associated with the question text comprised of the metrics of significance for the one or more text segments;

determining a number of question text-answer pairs needed to train a text classifier based on a topic associated with the text classifier;

generating a ground truth training set for training the text classifier by:

identifying question text-answer pairs from the expanded question set associated with the topic of the text classifier;

calculating vector distances between the vector characterizations of the identified question text-answer pairs; and

selecting, based on the vector distances, question text-answer pairs for the ground truth training set based on the determined number needed to train the text classifier; and

training the text classifier using the ground truth training set.

2. The method of claim 1 , further comprising:

receiving text classifier feedback; and

retraining the text classifier based on the received feedback.

3. The method of claim 1 , further comprising:

receiving a question associated with the text classifier;

determining a measure of dissimilarity between the received question and the ground truth training set;

adding, based on the measure of dissimilarity, the received question to the ground truth training set; and

retraining the text classifier based on the ground truth training set.

4. The method of claim 1 , wherein the text segment metric of significance is a term frequency-inverse document frequency (TF-IDF) value, and wherein the TF-IDF value for a question text segment is based on the number of occurrences of the text segment in the question text and the number of question texts in the expanded question set in which the text segment occurs.

5. The method of claim 1 , wherein the text segment is a unigram, bigram, or skip-gram.

6. The method of claim 1 , wherein the vector distances are calculated based on the cosine distance between the two vectors.

7. The method of claim 1 , wherein the selection of the question text-answer pairs for the ground truth training set comprises:

determining a threshold dissimilarity vector distance;

initializing an empty set of training questions; and

for each of the identified question text-answer pairs, adding the identified question text-answer pair to the set of training questions when the vector distance between the identified question text-answer pair and a question text-answer pair of the training set exceeds the threshold dissimilarity vector distance.

8. The method of claim 1 , wherein the selection of the question text-answer pairs for the ground truth training set comprises:

determining a threshold similarity vector distance;

constructing a set of training questions comprised of the identified question text-answer pairs; and

for each question text-answer pair of the set of training questions, removing the question text-answer pair from the training set when the vector distance between the question text-answer pair and a second question text-answer pair of the training set exceeds the similarity threshold distance.

9. A non-transitory computer-readable medium containing instruction configured to cause one or more processors to perform a method of generating a ground truth for training a text classifier, the method comprising:

retrieving a plurality of question forms and a plurality of intent forms, wherein each intent form includes an intent identifier, at least one intent synonyms, and answer information;

generating, based on the question forms and intent forms, an expanded question set comprised of question text and corresponding answers, wherein each question text-answer pair is associated with a topic;

constructing vectors characterizing the questions of the expanded question set by, for each question text:

dividing the question text into one or more text segments;

calculating metrics of significance for the one or more text segments; and

generating a vector associated with the question text comprised of the metrics of significance for the one or more text segments;

determining a number of question text-answer pairs needed to train a text classifier based on a topic associated with the text classifier;

generating a ground truth training set for training the text classifier by:

identifying question text-answer pairs from the expanded question set associated with the topic of the text classifier;

calculating vector distances between the vector characterizations of the identified question text-answer pairs; and

selecting, based on the vector distances, question text-answer pairs for the ground truth training set based on the determined number needed to train the text classifier; and

training the text classifier using the ground truth training set.

10. The non-transitory computer-readable medium of claim 9 , wherein the method further comprises:

receiving text classifier feedback; and

retraining the text classifier based on the received feedback.

11. The non-transitory computer-readable medium of claim 9 , wherein the method further comprises:

receiving a question associated with the text classifier;

determining a measure of dissimilarity between the received question and the ground truth training set;

adding, based on the measure of dissimilarity, the received question to the ground truth training set; and

retraining the text classifier based on the ground truth training set.

12. The non-transitory computer-readable medium of claim 9 , wherein the text segment metric of significance is a term frequency-inverse document frequency (TF-IDF) value, and wherein the TF-IDF value for a question text segment is based on the number of occurrences of the text segment in the question text and the number of question texts in the expanded question set in which the text segment occurs.

13. The non-transitory computer-readable medium of claim 9 , wherein the text segment is a unigram, bigram, or skip-gram.

14. The non-transitory computer-readable medium of claim 9 , wherein the vector distances are calculated based on the cosine distance between the two vectors.

15. The non-transitory computer-readable medium of claim 9 , wherein the selection of the question text-answer pairs for the ground truth training set comprises:

determining a threshold dissimilarity vector distance;

initializing an empty set of training questions; and

for each of the identified question text-answer pairs, adding the identified question text-answer pair to the set of training questions when the vector distance between the identified question text-answer pair and a question text-answer pair of the training set exceeds the threshold dissimilarity vector distance.

16. The non-transitory computer-readable medium of claim 9 , wherein the selection of the question text-answer pairs for the ground truth training set comprises:

determining a threshold similarity vector distance;

constructing a set of training questions comprised of the identified question text-answer pairs; and

for each question text-answer pair of the set of training questions, removing the question text-answer pair from the training set when the vector distance between the question text-answer pair and a second question text-answer pair of the training set exceeds the similarity threshold distance.

Assignments (9)
RELEASE OF FIRST LIEN SECURITY INTEREST AT 58671/0463 Recorded Nov 8, 2023
From: KKR LOAN ADMINISTRATION SERVICES LLC
To: WELLTOK, INC.
Reel/Frame 065524/0227 →
RELEASE OF SECURITY INTEREST AT 58671/0470 Recorded Nov 8, 2023
From: JPMORGAN CHASE BANK, N.A.
To: WELLTOK, INC.
Reel/Frame 065524/0232 →
SECURITY INTEREST Recorded Nov 8, 2023
From: WELLTOK, INC.
To: ALTER DOMUS (US) LLC, AS ADMINISTRATIVE AGENT
Reel/Frame 065497/0553 →
PATENT SECURITY AGREEMENT Recorded Jan 10, 2022
From: WELLTOK, INC.
To: KKR LOAN ADMINISTRATION SERVICES LLC, AS A COLLATERAL AGENT
Reel/Frame 058671/0463 →
PATENT SECURITY AGREEMENT Recorded Jan 10, 2022
From: WELLTOK, INC.
To: JPMORGAN CHASE BANK, N.A. AS A COLLATERAL AGENT
Reel/Frame 058671/0470 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded Nov 19, 2021
From: TRUSTMARK GROUP, INC.
To: WELLTOK, INC.
Reel/Frame 058567/0787 →
RELEASE OF SECURITY INTEREST RECORDED AT R/F 044166/0830 Recorded Nov 11, 2021
From: SILICON VALLEY BANK
To: WELLTOK, INC.
Reel/Frame 058105/0085 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Oct 5, 2017
From: WELLTOK, INC.
To: SILICON VALLEY BANK
Reel/Frame 044166/0830 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2016
From: MACLEOD, MATTHEW KELLAR; SWARTZ, JACQUE W.; PONDER, MARISSA VICTORIA
To: WELLTOK, INC.
Reel/Frame 040470/0685 →
Continuity (1)
Related Publication 20180150607A1 · May 31, 2018