IP Library › Granted Patent US 11,893,990
Granted Patent B2
US 11,893,990 · App. 17/486,661 · Granted Feb 6, 2024

Audio file annotation

Inventor: Hans-Martin Ramsl (Mannheim, DE)
Assignee: SAP SE
G10L15/22G06F40/295G10L15/26G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,893,990
App. No.
17/486,661
Granted
Feb 6, 2024
Kind
B2
Abstract

Text-to-speech translation is used to generate a transcript for an audio file. Text segments are associated with time segments in the transcript. A trained machine learning model determines, based on the text in the transcript, one or more topics for the audio file. The transcript is modified to include the determined one or more topics. A user interface may be presented that allows a user to search for portions of an audio file that relate to a particular topic. In response to the selected or entered topic, the user interface presents segments having a matching topic. The user may use voice or other user interface commands to modify the annotation of the audio file. User commands may also be used to extract data from the transcript and copy the data to a clipboard or to another application.

Claims (48)

1. A method comprising:

accessing, by one or more processors, an annotation file for an audio file, the annotation file comprising a text transcription of the audio file;

determining, by the one or more processors, based on the text transcription, a topic for each segment of a plurality of segments of a first predetermined length;

determining, by the one or more processors, a confidence level of the topic for each segment of the plurality of segments;

determining, by the one or more processors, a topic for each larger segment of a plurality of larger segments, each larger segment comprising a predetermined number of consecutive component segments of the plurality of segments, the determining of the topic for a larger segment based on the topics of the component segments and the confidence levels of the topics of the component segments; and

modifying the annotation file to include the determined topics for the plurality of larger segments.

2. The method of claim 1 , wherein the first predetermined length is one minute.

3. The method of claim 1 , further comprising:

based on a voice command, modifying the annotation file to indicate that a portion of the text transcription is highlighted.

4. The method of claim 3 , wherein the voice command comprises an indication of a duration of the audio file to highlight.

5. The method of claim 3 , wherein the modifying of the annotation file includes storing a timestamp that indicates when the portion of the text transcription was highlighted.

6. The method of claim 1 , further comprising:

based on a voice command, modifying the annotation file to include a comment.

7. The method of claim 1 , further comprising:

identifying, based on a segment, a named entity and a type of the named entity; and

modifying the annotation file to include an indication that the segment includes the named entity, the indication comprising the type of the named entity.

8. The method of claim 7 , further comprising:

based on search criteria and the indication, identifying the segment, the search criteria comprising a type of named entity and a period of time within the audio file.

9. The method of claim 1 , further comprising:

based on a voice command, copying a portion of the text transcription to a clipboard.

10. The method of claim 1 , further comprising:

based on a voice command, modifying the annotation file to include at least a portion of text corresponding to the voice command.

11. The method of claim 1 , wherein the determining of the topic for each larger segment of the plurality of larger segments comprises selecting from among the topics of the component segments based on the confidence levels of the topics of the component segments.

12. The method of claim 1 , wherein the multiple is three.

13. The method of claim 1 , wherein the multiple is seven.

14. A system comprising:

a memory that stores instructions; and

one or more processors configured by the instructions to perform operations comprising:

accessing an annotation file for an audio file, the annotation file comprising a text transcription of the audio file;

determining based on the text transcription, a topic for each segment of a plurality of segments of a first predetermined length;

determining a confidence level of the topic for each segment of the plurality of segments;

determining a topic for each larger segment of a plurality of larger segments of a second predetermined length, each larger segment comprising a predetermined number of consecutive component segments of the plurality of segments, the determining of the topic for a larger segment based on the topics of the component segments and the confidence levels of the topics of the component segments; and

modifying the annotation file to include the determined topics for the plurality of larger segments.

15. The system of claim 14 , wherein the first predetermined length is one minute.

16. The system of claim 14 , wherein the operations further comprise:

based on a voice command, modifying the annotation file to include a comment.

17. The system of claim 14 , wherein the operations further comprise:

identifying, based on a segment, a named entity and a type of the named entity; and

modifying the annotation file to include an indication that the segment includes the named entity, the indication comprising the type of the named entity.

18. A non-transitory computer-readable medium that stores instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

accessing an annotation file for an audio file, the annotation file comprising a text transcription of the audio file;

determining based on the text transcription, a topic for each segment of a plurality of segments of a first predetermined length;

determining a confidence level of the topic for each segment of the plurality of segments;

determining a topic for each larger segment of a plurality of larger segments of a second predetermined length, each larger segment comprising a predetermined number of consecutive component segments of the plurality of segments, the determining of the topic for a larger segment based on the topics of the component segments and the confidence levels of the topics of the component segments; and

modifying the annotation file to include the determined topics for the plurality of larger segments.

19. The non-transitory computer-readable medium of claim 18 , wherein the operations further comprise:

based on a voice command, modifying the annotation file to indicate that a portion of the text transcription is highlighted.

20. The non-transitory computer-readable medium of claim 19 , wherein the voice command comprises an indication of a duration of the audio file to highlight.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2021
From: RAMSL, HANS-MARTIN
To: SAP SE
Reel/Frame 057690/0449 →
Continuity (1)
Related Publication 20230094828A1 · Mar 30, 2023
Cited By (1)
US 12,431,112