IP Library Granted Patent US 10,521,465
Granted Patent B2
US 10,521,465 · App. 16/452,339 · Granted Dec 31, 2019

Deep reinforced model for abstractive summarization

Inventor: Romain Paulus (Menlo Park, CA)
Assignee: salesforce.com, inc.
G06F16/345G06F17/277G06F17/289G06N3/006G06N3/0445G06N3/0454G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,521,465
App. No.
16/452,339
Granted
Dec 31, 2019
Kind
B2
Abstract

A system for text summarization includes an encoder for encoding input tokens of a document and a decoder for emitting summary tokens which summarize the document based on the encoded input tokens. At each iteration the decoder generates attention scores between a current hidden state of the decoder and previous hidden states of the decoder, generates a current decoder context from the attention scores and the previous hidden states of the decoder, and selects a next summary token based on the current decoder context and a current encoder context of the encoder. The attention scores penalize candidate summary tokens having high attention scores in previous iterations. In some embodiments, the attention scores include an attention score for each of the previous hidden states of the decoder. In some embodiments, the selection of the next summary token prevents emission of repeated summary phrases in a summary of the document.

Claims (42)

1. A text summarization system comprising:

an encoder for encoding input tokens of a document to be summarized; and

a decoder for emitting summary tokens which summarize the document based on the encoded input tokens, wherein at each iteration the decoder:

generates attention scores between a current hidden state of the decoder and previous hidden states of the decoder;

generates a current decoder context from the attention scores and the previous hidden states of the decoder; and

selects a next summary token based on the current decoder context and a current encoder context of the encoder;

wherein the attention scores penalize candidate summary tokens having high attention scores in previous iterations.

2. The text summarization system of claim 1 , wherein the attention scores include an attention score for each of the previous hidden states of the decoder.

3. The text summarization system of claim 1 , wherein at each iteration, the decoder further normalizes the attention scores.

4. The text summarization system of claim 3 , wherein the decoder normalizes the attention scores using a softmax layer.

5. The text summarization system of claim 1 , wherein the current decoder context is a convex combination of the attention scores and the previous hidden states of the decoder.

6. The text summarization system of claim 1 , wherein the selection of the next summary token prevents emission of repeated summary phrases in a summary of the document.

7. The text summarization system of claim 1 , wherein the next summary token is further based on the current hidden state of the decoder.

8. A method for summarizing text, the method comprising:

receiving a document to be summarized;

encoding, using an encoder, input tokens of the document;

generating, using a decoder, attention scores between a current hidden state of the decoder and previous hidden states of the decoder;

generating, using the decoder, a current decoder context from the attention scores and the previous hidden states of the decoder; and

selecting, using the decoder, a next summary token based on the current decoder context and a current encoder context of the encoder;

wherein:

the next summary token from each iteration summarizes the document; and

the attention scores penalize candidate summary tokens having high attention scores in previous iterations.

9. The method of claim 8 , wherein the attention scores include an attention score for each of the previous hidden states of the decoder.

10. The method of claim 8 , further comprising, at each iteration, normalizing the attention scores.

11. The method of claim 10 , wherein normalizing the attention scores comprises using a softmax layer.

12. The method of claim 8 , wherein generating the current decoder context comprises generating a convex combination of the attention scores and the previous hidden states of the decoder.

13. The method of claim 8 , wherein the selecting of the next summary token prevents emission of repeated summary phrases in a summary of the document.

14. The method of claim 8 , wherein selecting the next summary token is further based on the current hidden state of the decoder.

15. A tangible non-transitory computer readable storage medium impressed with computer program instructions that, when executed on a processor, implement a method comprising:

receiving a document to be summarized;

encoding, using an encoder, input tokens of the document;

generating, using a decoder, attention scores between a current hidden state of the decoder and previous hidden states of the decoder;

generating, using the decoder, a current decoder context from the attention scores and the previous hidden states of the decoder; and

selecting, using the decoder, a next summary token based on the current decoder context and a current encoder context of the encoder;

wherein:

the next summary token from each iteration summarizes the document; and

the attention scores penalize candidate summary tokens having high attention scores in previous iterations.

16. The tangible non-transitory computer readable storage medium of claim 15 , wherein the attention scores include an attention score for each of the previous hidden states of the decoder.

17. The tangible non-transitory computer readable storage medium of claim 15 , further comprising, at each iteration, normalizing the attention scores.

18. The tangible non-transitory computer readable storage medium of claim 15 , wherein generating the current decoder context comprises generating a convex combination of the attention scores and the previous hidden states of the decoder.

19. The tangible non-transitory computer readable storage medium of claim 15 , wherein the selecting of the next summary token prevents emission of repeated summary phrases in a summary of the document.

20. The tangible non-transitory computer readable storage medium of claim 15 , wherein selecting the next summary token is further based on the current hidden state of the decoder.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2019
From: PAULUS, ROMAIN
To: SALESFORCE.COM, INC.
Reel/Frame 049584/0589 →
Continuity (3)
Continuation 15815686 · Nov 16, 2017
Provisional Application 62485876 · Apr 14, 2017
Related Publication 20190311002A1 · Oct 10, 2019
Cited By (3)
US 12,265,909 US 12,299,982 US 12,530,560