IP Library › Granted Patent US 11,704,351
Granted Patent B1
US 11,704,351 · App. 17/976,461 · Granted Jul 18, 2023

Machine-learning model for performing contextual summarization of text data

Inventors: Reza Soleimani (Raleigh, NC); Samuel Leeman-Munk (Durham, NC); David Blake Styles (Raleigh, NC)
Assignee: SAS Institute, Inc.
G06F16/34G06F40/284G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,704,351
App. No.
17/976,461
Granted
Jul 18, 2023
Kind
B1
Abstract

In one example, a system can receive a set of text samples and generate a set of summaries based on the set of text samples. The system can then generate a training dataset by iteratively executing a training-sample generation process. Each iteration can involve selecting multiple text samples from the set of text samples, combining the multiple text samples together into a training sample, determining a text category and a summary corresponding to a selected one of the multiple text samples, and including the text category and the summary in the training sample. After generating the training dataset, the system can use it to train a model. The trained model can then receive a target textual dataset and a target category as input, identify a portion of the target textual dataset corresponding to the target category, and generate a summarization of the portion of that target textual dataset.

Claims (111)

1. A system comprising:

one or more processors; and

one or more storage devices including instructions that are executable by the one or more processors for causing the one or more processors to:

receive a plurality of text samples corresponding to a plurality of text categories, each text sample of the plurality of text samples corresponding to a respective category of the plurality of text categories;

provide the plurality of text samples as input to a summarization model that is configured to generate a plurality of summaries of the plurality of text samples;

generate a training dataset by iteratively performing a training-sample generation process until a stopping condition is met, the training-sample generation process involving:

selecting a first text sample and a second text sample from the plurality of text samples, wherein the first text sample and the second text sample correspond to different categories of the plurality of text categories;

generating a training sample that includes the first text sample and the second text sample;

determining a text category corresponding to the first text sample;

incorporating an identifier of the text category into the training sample;

selecting a summary corresponding to the first text sample from among the plurality of summaries;

incorporating the summary into the training sample; and

storing the training sample as part of the training dataset;

train a model based on the training dataset to generate a trained model, the trained model being configured to:

receive a target textual dataset and a target category as input, the target category being one of the plurality of text categories;

identify a portion of the target textual dataset corresponding to the target category; and

generate an output that includes a summarization of the portion of the target textual dataset corresponding to the target category.

2. The system of claim 1 , wherein the one or more storage devices further include instructions that are executable by the one or more processors for causing the one or more processors to:

generate another text string that includes the first text sample and the second text sample;

select another summary from the plurality of summaries corresponding to the second text sample;

incorporate the other summary into the other text string;

determine another text category corresponding to the second text sample;

incorporate another identifier of the other text category into the other text string; and

store the other text string as another training sample in the training dataset.

3. The system of claim 1 , wherein the plurality of text categories include a plurality of topics.

4. The system of claim 1 , wherein the one or more storage devices further include instructions that are executable by the one or more processors for causing the one or more processors to:

provide the summarization as input to a trained classifier model, the trained classifier model being configured to output a predicted category associated with the summarization;

compare the predicted category to the target category; and

in response to determining that there is a difference between the predicted category and the target category based on the comparison, transmit a notification indicating that the model produced an erroneous summary.

5. The system of claim 4 , wherein the trained classifier model is an interpretable model configured to provide an explanation of why it selected the predicted category, and wherein the notification include the explanation.

6. The system of claim 4 , wherein the one or more storage devices further include instructions that are executable by the one or more processors for causing the one or more processors to generate the trained classifier model by training a classifier model based on the plurality of text samples.

7. The system of claim 1 , wherein the model includes the summarization model.

8. The system of claim 1 , wherein the plurality of text categories correspond to a plurality of sentiments, and wherein the plurality of sentiments include positive sentiment and negative sentiment.

9. The system of claim 1 , wherein the one or more storage devices further include instructions that are executable by the one or more processors for causing the one or more processors to:

select at least a subset of text samples from the plurality of text samples; and

generate the training dataset by performing the training-sample generation process on every possible pair of text samples in the subset of text samples.

10. The system of claim 1 , wherein the one or more storage devices further include instructions that are executable by the one or more processors for causing the one or more processors to:

determine the text category corresponding to the first text sample based on an assignment of the text category to the first text sample in the plurality of text samples.

11. A method comprising:

receiving, by one or more processors, a plurality of text samples corresponding to a plurality of text categories, each text sample of the plurality of text samples corresponding to a respective category of the plurality of text categories;

providing, by the one or more processors, the plurality of text samples as input to a summarization model that is configured to generate a plurality of summaries of the plurality of text samples;

generating, by the one or more processors, a training dataset by iteratively performing a training-sample generation process until a stopping condition is met, the training-sample generation process involving:

selecting a first text sample and a second text sample from the plurality of text samples, wherein the first text sample and the second text sample correspond to different categories of the plurality of text categories;

generating a training sample that includes the first text sample and the second text sample;

determining a text category corresponding to the first text sample;

incorporating an identifier of the text category into the training sample;

selecting a summary corresponding to the first text sample from among the plurality of summaries;

incorporating the summary into the training sample; and

storing the training sample as part of the training dataset;

training, by the one or more processors, a model based on the training dataset to generate a trained model, the trained model being configured to:

receive a target textual dataset and a target category as input, the target category being one of the plurality of text categories;

identify a portion of the target textual dataset corresponding to the target category; and

generate an output that includes a summarization of the portion of the target textual dataset corresponding to the target category.

12. The method of claim 11 , further comprising:

generating another text string that includes the first text sample and the second text sample;

selecting another summary from the plurality of summaries corresponding to the second text sample;

incorporating the other summary into the other text string;

determining another text category corresponding to the second text sample;

incorporating another identifier of the other text category into the other text string; and

storing the other text string as another training sample in the training dataset.

13. The method of claim 11 , wherein the plurality of text categories include a plurality of topics.

14. The method of claim 11 , further comprising:

providing the summarization as input to a trained classifier model, the trained classifier model being configured to output a predicted category associated with the summarization;

comparing the predicted category to the target category; and

in response to determining that there is a difference between the predicted category and the target category based on the comparison, transmitting a notification indicating that the model produced an erroneous summary.

15. The method of claim 14 , wherein the trained classifier model is an interpretable model configured to provide an explanation of why it selected the predicted category, and wherein the notification include the explanation.

16. The method of claim 14 , further comprising generating the trained classifier model by training a classifier model based on the plurality of text samples.

17. The method of claim 11 , wherein the model includes the summarization model.

18. The method of claim 11 , wherein the plurality of text categories correspond to a plurality of sentiments, and wherein the plurality of sentiments include positive sentiment and negative sentiment.

19. The method of claim 11 , further comprising:

selecting at least a subset of text samples from the plurality of text samples; and

generating the training dataset by performing the training-sample generation process on every possible pair of text samples in the subset of text samples.

20. The method of claim 11 , further comprising:

determining the text category corresponding to the first text sample based on an assignment of the text category to the first text sample in the plurality of text samples.

21. A non-transitory computer-readable medium comprising program code that is executable by one or more processors for causing the one or more processors to:

receive a plurality of text samples corresponding to a plurality of text categories, each text sample of the plurality of text samples corresponding to a respective category of the plurality of text categories;

provide the plurality of text samples as input to a summarization model that is configured to generate a plurality of summaries of the plurality of text samples;

generate a training dataset by iteratively performing a training-sample generation process until a stopping condition is met, the training-sample generation process involving:

selecting a first text sample and a second text sample from the plurality of text samples, wherein the first text sample and the second text sample correspond to different categories of the plurality of text categories;

generating a training sample that includes the first text sample and the second text sample;

determining a text category corresponding to the first text sample;

incorporating an identifier of the text category into the training sample;

selecting a summary corresponding to the first text sample from among the plurality of summaries;

incorporating the summary into the training sample; and

storing the training sample as part of the training dataset;

train a model based on the training dataset to generate a trained model, the trained model being configured to:

receive a target textual dataset and a target category as input, the target category being one of the plurality of text categories;

identify a portion of the target textual dataset corresponding to the target category; and

generate an output that includes a summarization of the portion of the target textual dataset corresponding to the target category.

22. The non-transitory computer-readable medium of claim 21 , further comprising program code that is executable by the one or more processors for causing the one or more processors to:

generate another text string that includes the first text sample and the second text sample;

select another summary from the plurality of summaries corresponding to the second text sample;

incorporate the other summary into the other text string;

determine another text category corresponding to the second text sample;

incorporate another identifier of the other text category into the other text string; and

store the other text string as another training sample in the training dataset.

23. The non-transitory computer-readable medium of claim 21 , wherein the plurality of text categories include a plurality of topics.

24. The non-transitory computer-readable medium of claim 21 , further comprising program code that is executable by the one or more processors for causing the one or more processors to:

provide the summarization as input to a trained classifier model, the trained classifier model being configured to output a predicted category associated with the summarization;

compare the predicted category to the target category; and

in response to determining that there is a difference between the predicted category and the target category based on the comparison, transmit a notification indicating that the model produced an erroneous summary.

25. The non-transitory computer-readable medium of claim 24 , wherein the trained classifier model is an interpretable model configured to provide an explanation of why it selected the predicted category, and wherein the notification include the explanation.

26. The non-transitory computer-readable medium of claim 24 , further comprising program code that is executable by the one or more processors for causing the one or more processors to:

generate the trained classifier model by training a classifier model based on the plurality of text samples.

27. The non-transitory computer-readable medium of claim 21 , wherein the model includes the summarization model.

28. The non-transitory computer-readable medium of claim 21 , wherein the plurality of text categories correspond to a plurality of sentiments, and wherein the plurality of sentiments include positive sentiment and negative sentiment.

29. The non-transitory computer-readable medium of claim 21 , further comprising program code that is executable by the one or more processors for causing the one or more processors to:

select at least a subset of text samples from the plurality of text samples; and

generate the training dataset by performing the training-sample generation process on every possible pair of text samples in the subset of text samples.

30. The non-transitory computer-readable medium of claim 21 , further comprising program code that is executable by the one or more processors for causing the one or more processors to:

determine the text category corresponding to the first text sample based on an assignment of the text category to the first text sample in the plurality of text samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: SOLEIMANI, REZA; LEEMAN-MUNK, SAMUEL PAUL; STYLES, DAVID BLAKE
To: SAS INSTITUTE INC.
Reel/Frame 061584/0778 →
Continuity (1)
Provisional Application 63353772 · Jun 20, 2022
Cited By (2)
US 12,596,889 US 12,695,704