IP Library Granted Patent US 10,783,314
Granted Patent B2
US 10,783,314 · App. 16/024,212 · Granted Sep 22, 2020

Emphasizing key points in a speech file and structuring an associated transcription

Inventors: Franck Dernoncourt (San Jose, CA); Walter Wei-Tuh Chang (San Jose, CA); Seokhwan Kim (San Jose, CA); Sean Fitzgerald (Campbell, CA); Ragunandan Rao Malangully (San Jose, CA); Laurie Marie Byrum (Pleasanton, CA); Frederic Thevenet (San Francisco, CA); Carl Iwan Dockhorn (San Jose, CA)
Assignee: Adobe Inc.
G06F40/106G06F40/14G06F40/166G10L15/26G06F40/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,783,314
App. No.
16/024,212
Granted
Sep 22, 2020
Kind
B2
Abstract

Techniques are disclosed for generating a structured transcription from a speech file. In an example embodiment, a structured transcription system receives a speech file comprising speech from one or more people and generates a navigable structured transcription object. The navigable structured transcription object may comprise one or more data structures representing multimedia content with which a user may navigate and interact via a user interface. Text and/or speech relating to the speech file can be selectively presented to the user (e.g., the text can be presented via a display, and the speech can be aurally presented via a speaker).

Claims (57)

1. A method for generating a structured transcription of a speech file, the method comprising:

converting said speech file to text;

processing said speech file to determine at least one sentence;

processing said at least one sentence to generate a document tree structure comprising a plurality of sections;

generating a highlighted representation of said text by:

computing term-frequency vectors based on said text, and,

performing a highlighting operation on each sentence by performing a binary classification based upon a maximum term-frequency vector associated with said sentence and acoustic features associated with said sentence, and in response to said binary classification outputting a pre-determined value, performing a formatting operation to highlight said sentence; and,

performing an interactive navigation of said structured transcription based upon said speech file, said document tree structure and said highlighted representation.

2. The method according to claim 1 , wherein processing said at least one sentence to generate said document tree structure comprises outputting a binary value indicating whether said at least one sentence concludes one of said sections.

3. The method according to claim 1 , wherein processing said at least one sentence to generate said document tree structure comprises performing a TextTiling process.

4. The method according to claim 1 , further comprising generating, based upon said at least one sentence, an extractive summary, or an abstractive summary, or both an extractive summary and an abstractive summary.

5. The method according to claim 4 , wherein said extractive summary is utilized to perform text highlighting.

6. The method according to claim 3 , wherein processing said at least one sentence to generate said document tree structure further comprises:

segmenting said at least one sentence into at least one segment;

utilizing said at least one segment to generate at least one of said sections; and,

constructing said document tree from said at least one generated section.

7. The method according to claim 1 , wherein processing said speech file to determine at least one sentence comprises:

performing an automatic speech recognition (“ASR”) on said speech file to generate a first file;

processing said first file to determine said at least one sentence to generate a second file; and,

processing said second file to decompose run-on sentences into smaller logical sentences to generate a third file.

8. The method according to claim 7 , further comprising processing said third file to associate said at least one sentence with a respective speaker to generate a fourth file.

9. A system for processing a structured transcription of a speech file, the system comprising:

a sentence identifier, wherein said sentence identifier generates at least one sentence from said speech file;

a summarizer, wherein said summarizer generates at least one summary based upon said at least one sentence;

a document tree analyzer, wherein said document tree analyzer generates a document tree structure;

a document highlighting module for generating a highlighted textual representation of said speech file, wherein said document highlighting module further comprises:

a term-frequency vector computation module for generating term-frequency vectors, and

a binary classifier for performing a binary classification of each sentence based upon a maximum term-frequency vector associated with said sentence and acoustic features associated with said sentence; and,

a navigator, wherein said navigator performs an interactive navigation of said structured transcription based upon said speech file, said document tree structure and said highlighted textual representation.

10. The system according to claim 9 , wherein each of said at least one summary comprises an abstractive summary and an extractive summary.

11. The system according to claim 9 , wherein said document tree analyzer generates said document tree structure using a TextTiling process.

12. The system according to claim 9 , wherein said sentence identifier comprises:

a speech recognition engine, wherein said speech recognition engine generates a text representation of said speech file;

a sentence boundary detector, wherein said sentence boundary detector generates said at least one sentence based upon said text representation;

a run-on sentence detector, wherein said run-on sentence-detector splits a run-on sentence into at least two shorter sentences; and,

a speaker sentence identifier, wherein said speaker sentence identifier associates each of said at last one sentence with a respective speaker.

13. The system according to claim 9 , wherein said document tree analyzer:

segments said at least one sentence into at least one segment;

utilizes said at least one segment to generate at least one section; and,

constructs said document tree structure from said at least one section.

14. The system according to claim 9 , wherein said document tree analyzer outputs a binary value indicating whether said at least one sentence concludes a section.

15. A computer program product including one or more non-transitory machine-readable mediums encoded with instructions that when executed by one or more processors cause a process to be carried out for processing a structured transcription of a speech file, the process comprising:

converting said speech file to text;

processing said speech file to determine at least one sentence;

processing said at least one sentence to generate a document tree structure comprising a plurality of sections;

generating a highlighted representation of said text by:

computing term-frequency vectors based on said text, and,

performing a highlighting operation on each sentence by performing a binary classification based upon a maximum term-frequency vector associated with said sentence and acoustic features associated with said sentence, and in response to said binary classification outputting a pre-determined value, performing a formatting operation to highlight said sentence; and,

performing an interactive navigation of said structured transcription based upon said speech file, said document tree structure and said highlighted representation.

16. The computer program product according to claim 15 , wherein processing said at least one sentence to generate said document tree structure comprises outputting a binary value indicating whether said at least one sentence concludes one of said sections.

17. The computer program product according to claim 15 , wherein processing said at least one sentence to generate said document tree structure comprises performing a TextTiling process.

18. The computer program product according to claim 15 , wherein said process further comprises generating, based upon said at least one sentence, an extractive summary, or an abstractive summary, or both an extractive summary and an abstractive summary.

19. The computer program product according to claim 18 , wherein said extractive summary is utilized to perform text highlighting.

20. The computer program product according to claim 15 , wherein processing said speech file to determine said at least one sentence further comprises:

performing an automatic speech recognition (“ASR”) on said speech file to generate a first file;

processing said first file to determine said at least one sentence to generate a second file; and,

processing said second file to decompose run-on sentences into smaller logical sentences to generate a third file.

Assignments (2)
CHANGE OF NAME Recorded Nov 30, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047688/0530 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2018
From: DERNONCOURT, FRANCK; CHANG, WALTER WEI-TUH; KIM, SEOKHWAN; FITZGERALD, SEAN; MALANGULLY, RAGUNANDAN RAO; BYRUM, LAURIE MARIE; THEVENET, FREDERIC; DOCKHORN, CARL IWAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 046520/0484 →
Continuity (1)
Related Publication 20200004803A1 · Jan 2, 2020
Cited By (2)
US 12,210,818 US 12,327,078