IP Library › Granted Patent US 12,045,272
Granted Patent B2
US 12,045,272 · App. 17/370,899 · Granted Jul 23, 2024

Auto-creation of custom models for text summarization

Inventors: Saurabh Mahapatra (Sunnyvale, CA); Niyati Chhaya (Hyderabad, IN); Snehal Raj (Odisha, IN); Sharmila Reddy Nangi (Hyderabad, IN); Sapthotharan Nair (Cupertino, CA); Sagnik Mukherjee (Kolkata, IN); Jay Mundra (Chittorgarh Rajasthan, IN); Fan Du (Milpitas, CA); Atharv Tyagi (Uttar Pradesh, IN); Aparna Garimella (Karnataka, IN)
Assignee: ADOBE INC.
G06F16/345G06F16/3329G06F40/30G06N3/04G06N3/044G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,045,272
App. No.
17/370,899
Filed
Jul 8, 2021
Granted
Jul 23, 2024
Kind
B2
Art Unit
2657
USPC
704/9
Abstract

A text summarization system auto-generates text summarization models using a combination of neural architecture search and knowledge distillation. Given an input dataset for generating/training a text summarization model, neural architecture search is used to sample a search space to select a network architecture for the text summarization model. Knowledge distillation includes fine-tuning a language model for a given text summarization task using the input dataset, and using the fine-tuned language model as a teacher model to inform the selection of the network architecture and the training of the text summarization model. Once a text summarization model has been generated, the text summarization model can be used to generate summaries for given text.

Claims (43)

1. One or more computer storage media storing computer-useable instructions that, when used by a computing device, cause the computing device to perform operations, the operations comprising:

receiving an input dataset;

determining a type of text summarization task as an extractive summarization task or an abstractive summarization task;

fine-tuning a language model for the determined type of text summarization task using the input dataset; and

generating a text summarization model for the determined type of text summarization task by:

using neural architecture search to learn a network architecture for an encoder of the text summarization model for only the determined type of text summarization task,

selecting a pre-defined decoder of the text summarization model based on the determined type of text summarization task, and

using knowledge distillation to train the text summarization model on the input dataset using the fine-tuned language model as a teacher model.

2. The one or more computer storage media of claim 1 , wherein the input dataset comprises a plurality of examples, each example comprising an example text and an example summary of the example text.

3. The one or more computer storage media of claim 2 , wherein the type of text summarization task is automatically determined from the examples in the input dataset.

4. The one or more computer storage media of claim 1 , wherein the type of text summarization task is determined based on user input specifying the type of text summarization task.

5. The one or more computer storage media of claim 1 , wherein the language model comprises a bidirectional encoder representations from transformers (BERT) model.

6. The one or more computer storage media of claim 1 , wherein the neural architecture search employs reinforcement learning to train a controller to learn the network architecture for the text summarization model using a reward determined at each time step based on a validation loss.

7. The one or more computer storage media of claim 6 , wherein the controller selects the network architecture of the text summarization model from a search space defining types of cells for the network architecture and how the cells can be connected in the network architecture.

8. The one or more computer storage media of claim 7 , wherein the types of cells defined by the search space include: convolutional neural network, recurrent neural network, pooling layers, and multi-head self-attention.

9. The one or more computer storage media of claim 1 , wherein the text summarization model is trained using an overall loss based on a weighted contribution from a knowledge distillation loss determined using soft labels from the fine-tuned language model and a cross-entropy loss determined using ground truth labels from the input dataset.

10. The one or more computer storage media of claim 9 , wherein the overall loss is determined at sentence level for extractive summarization and at vocab level for abstractive summarization.

11. The one or more computer storage media of claim 6 , wherein the controller is a recurrent neural network based controller.

12. The one or more computer storage media of claim 1 , wherein, for extractive summarization, the pre-defined decoder of the text summarization model based on the determined type of text summarization task is a scorer function with sigmoid activation that takes in text representations from the encoder and scores each sentence.

13. The one or more computer storage media of claim 1 , wherein, for abstractive summarization, the pre-defined decoder of the text summarization model based on the determined type of text summarization task is a recurrent neural network that takes in text representations from the encoder and outputs a generated summary in an auto-regressive manner by decoding a word at each time step.

14. The one or more computer storage media of claim 1 , wherein the operations further comprise:

receiving input text; and

generating a summary of the input text using the text summarization model.

15. A computer-implemented method comprising:

receiving input, the input including an input dataset;

fine-tuning a language model for a text summarization task using the input dataset to provide a fine-tuned language model; and

generating a text summarization model for the text summarization task using neural architecture search and knowledge distillation, the text summarization model comprising an encoder and a decoder, the text summarization model being generated by iteratively:

using a controller to select a network architecture for the encoder from a search space;

training the text summarization model to minimize a total loss as a function of a knowledge distillation loss determined using soft labels from the fine-tuned language model and a cross-entropy loss determined using ground truth labels from the input dataset; and

updating the controller using a reward determined based on performance of the text summarization model.

16. The computer-implemented method of claim 15 , wherein the input further comprises one or more selected from the following: an indication of the text summarization task, a summary size, a model size, a number of layers, and a number of epochs.

17. The computer-implemented method of claim 15 , wherein, for extractive summarization, the decoder is a scorer function with sigmoid activation that takes in text representations from the encoder and scores each sentence, and wherein, for abstractive summarization, the decoder is a recurrent neural network that takes in text representations from the encoder and outputs a generated summary in an auto-regressive manner by decoding a word at each time step.

18. The computer-implemented method of claim 15 , wherein the search space defines types of cells for the network architecture and how the cells can be connected in the network architecture, and wherein the types of cells defined by the search space include: convolutional neural network, recurrent neural network, pooling layers, and multi-head self-attention.

19. The computer-implemented method of claim 15 , wherein the performance of the text summarization model is based on a validation loss determined for the text summarization model.

20. A computer system comprising:

a processor; and

a computer storage medium storing computer-useable instructions that, when used by the processor, causes the computer system to perform operations comprising:

receiving user input comprising input text;

feeding the input text to a text summarization model, the text summarization model generated by:

using neural architecture search to learn a network architecture of an encoder of the text summarization model for only a determined type of text summarization task,

selecting a pre-defined decoder of the text summarization model based on the determined type of text summarization task,

and using knowledge distillation from a language model fine-tuned for the determined type of text summarization task based on an input dataset; and

providing, in response to the user input, a summary of the input text generated by the text summarization model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: MAHAPATRA, SAURABH; CHHAYA, NIYATI; RAJ, SNEHAL; NANGI, SHARMILA REDDY; NAIR, SAPTHOTHARAN; MUKHERJEE, SAGNIK; MUNDRA, JAY; DU, FAN; TYAGI, ATHARV; GARIMELLA, APARNA
To: ADOBE INC.
Reel/Frame 056796/0506 →
Continuity (1)
Related Publication 20230020886A1 · Jan 19, 2023
Cited By (3)
US 12,387,053 US 12,485,795 US 12,511,547