IP Library Granted Patent US 11,594,227
Granted Patent B2
US 11,594,227 · App. 17/354,455 · Granted Feb 28, 2023

Computer-implemented method of transcribing an audio stream and transcription mechanism

Inventors: Lars Hermanns (Simmerath, DE); Thomas Nass (Hürth, DE); Stefan Moers (Baesweiler, DE); Frank Reif (Aachen, DE)
Assignee: Unify Patente GmbH & Co. KG
G10L15/26G06K9/6201G06V30/153H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,594,227
App. No.
17/354,455
Granted
Feb 28, 2023
Kind
B2
Abstract

A computer-implemented method of transcribing an audio stream can include transcribing the audio stream using a first transcribing instance having a first predetermined transcription size that is smaller than the total length of the audio stream. The first transcribing instance can provide a plurality of consecutive first transcribed text data snippets of the audio stream and the size of the first transcribed text data snippets can respectively corresponding to the first predetermined transcription size. The audio stream can also be transcribed using at least a second transcribing instance having a second predetermined transcription size that is smaller than the length of the audio stream. The second transcribing instance can provide a plurality of consecutive second transcribed text data snippets each corresponding to the second predetermined transcription size.

Claims (61)

1. A computer-implemented method of transcribing an audio stream comprises:

transcribing the audio stream using a first transcribing instance of a transcription service, the first transcribing instance having a first predetermined transcription size that is smaller than a total length of the audio stream, the first transcribing instance providing a plurality of consecutive first transcribed text data snippets of the audio stream, each of the first transcribed text data snippets having a size corresponding to the first predetermined transcription size;

transcribing the audio stream using at least a second transcribing instance having a second predetermined transcription size that is smaller than the length of the audio stream, the second transcribing instance providing a plurality of consecutive second transcribed text data snippets of the audio stream, each of the second transcribed text data snippets having a size corresponding to the second predetermined transcription size;

wherein the first transcribing instance starts transcription of the audio stream at a first point of time and the second transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the first transcribing instance;

wherein the predetermined delay is selected such that each one of the plurality of second text data snippets overlaps with at least an ending portion of a respective first transcribed text data snippet of the plurality of the first text data snippets ends and also overlaps with a starting portion of the first transcribed text data snippet of the plurality of the first text data snippets that is consecutive to the respective first transcribed text data snippet;

identifying matching text passages in overlapping portions of the first and second transcribed text data snippets via at least one of:

identifying at least one word pattern in the first transcribed text data snippet, the at least one word pattern comprising at least two long words with a predetermined number of short words in between the two long words and, in response to the at least one word pattern being identified in the first transcribed text data snippet, searching the identified at least one word pattern in the second transcribed text data snippets; and

identifying at least one syllable pattern according to a Porter-Stemmer algorithm in the first transcribed text data snippets and, in response to the at least one syllable pattern being identified in the first transcribed text data snippets, searching the identified at least one syllable pattern in the second transcribed text data snippets.

2. The computer implemented-method of claim 1 , wherein the method further comprises:

transcribing the audio stream using a third transcribing instance, the third transcribing instance having a third predetermined transcription size that is smaller than the total length of the audio stream, the third transcribing instance providing a plurality of consecutive third transcribed text data snippets of the audio stream, a size of the third transcribed text data snippets corresponding to the third predetermined transcription size;

wherein the third transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the second transcribing instance;

wherein the predetermined delay is selected such that each one of the plurality of third text data snippets respectively overlaps at least a portion at which a first transcribed text data snippet of the plurality of the first text data snippets ends and a consecutive first transcribed text data snippet of the plurality of the first text data snippets starts.

3. The computer-implemented method according to claim 1 , wherein the transcription size of the first transcribing instance is equal to the transcription size of the second transcribing instance.

4. The computer-implemented method according to claim 1 , wherein the identifying of the matching text passages also comprises:

identifying at least one word having a predetermined minimum length, and

in response to the at least one word having the predetermined minimum length being identified in the first transcribed text data snippet, searching for the identified at least one word in the second transcribed text data snippets.

5. The computer-implemented method according to claim 4 , comprising:

transcribing the audio stream using a third transcribing instance, the third transcribing instance having a third predetermined transcription size that is smaller than the total length of the audio stream, the third transcribing instance providing a plurality of consecutive third transcribed text data snippets of the audio stream, a size of the third transcribed text data snippets corresponding to the third predetermined transcription size;

wherein the third transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the second transcribing instance;

wherein the predetermined delay is selected such that each one of the plurality of third text data snippets respectively overlaps at least a portion at which a first transcribed text data snippet of the plurality of the first text data snippets ends and a consecutive first transcribed text data snippet of the plurality of the first text data snippets starts; and

in response to the at least one word having the predetermined minimum length being identified in the first transcribed text data snippet, searching the identified at least one word in the third transcribed text data snippets.

6. The computer-implemented method according to claim 4 , wherein the identified matching words and/or text passages are correlated.

7. The computer-implemented method of claim 1 , wherein

wherein the predetermined delay is selected such that one of the second text data snippets overlaps with at least an ending portion of a the first snippet of the plurality of the first text data snippets and also overlaps with a starting portion of a the second snippet of the plurality of the first text data snippets.

8. The computer implemented-method of claim 7 , wherein the method further comprises:

transcribing the audio stream using a third transcribing instance, the third transcribing instance having a third predetermined transcription size that is smaller than the total length of the audio stream, the third transcribing instance providing a plurality of consecutive third transcribed text data snippets of the audio stream, a size of the third transcribed text data snippets corresponding to the third predetermined transcription size;

wherein the third transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the second transcribing instance;

wherein the predetermined delay is selected such that one of the plurality of third text data snippets overlaps with at least an ending portion of the first snippet of the plurality of the first text data snippets and also overlaps with a starting portion of the second snippet of the plurality of the first text data snippets.

9. A computer-implemented method of transcribing an audio stream comprising:

transcribing the audio stream using a first transcribing instance of a transcription service, the first transcribing instance having a first predetermined transcription size that is smaller than a total length of the audio stream, the first transcribing instance providing a plurality of consecutive first transcribed text data snippets of the audio stream, each of the first transcribed text data snippets having a size corresponding to the first predetermined transcription size;

transcribing the audio stream using at least a second transcribing instance having a second predetermined transcription size that is smaller than the length of the audio stream, the second transcribing instance providing a plurality of consecutive second transcribed text data snippets of the audio stream, each of the second transcribed text data snippets having a size corresponding to the second predetermined transcription size;

wherein the first transcribing instance starts transcription of the audio stream at a first point of time and the second transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the first transcribing instance:

wherein the predetermined delay is selected such that each one of the plurality of second text data snippets overlaps with at least an ending portion of a respective first transcribed text data snippet of the plurality of the first text data snippets ends and also overlaps with a starting portion of the first transcribed text data snippet of the plurality of the first text data snippets that is consecutive to the respective first transcribed text data snippet;

concatenating the first transcribed text data snippets and the second transcribed text data snippets, the concatenating of the first and second transcribed text data snippets comprising identifying matching text passages in overlapping portions of the first and second transcribed text data snippets, wherein the identifying of the matching text passages comprises:

identifying at least one word pattern in the first transcribed text data snippet, the at least one word pattern comprising at least two long words with a predetermined number of short words in between the two long words, and

in response to the at least one word pattern being identified in the first transcribed text data snippet, searching the identified at least one word pattern in the second transcribed text data snippets.

10. The computer-implemented method of claim 9 , wherein the transcription service is a real-time transcription service or an Automatic Speech Recognition (ASR) service.

11. The computer-implemented method of claim 9 , comprising:

displaying the transcribed audio stream via a display device.

12. A transcription mechanism for a communication system for carrying out a video and/or audio conference with at least two participants, wherein the transcription mechanism is adapted to carry out the method of claim 9 .

13. A computer-implemented method of transcribing an audio stream comprising:

transcribing the audio stream using a first transcribing instance of a transcription service, the first transcribing instance having a first predetermined transcription size that is smaller than a total length of the audio stream, the first transcribing instance providing a plurality of consecutive first transcribed text data snippets of the audio stream, each of the first transcribed text data snippets having a size corresponding to the first predetermined transcription size;

transcribing the audio stream using at least a second transcribing instance having a second predetermined transcription size that is smaller than the length of the audio stream, the second transcribing instance providing a plurality of consecutive second transcribed text data snippets of the audio stream, each of the second transcribed text data snippets having a size corresponding to the second predetermined transcription size;

wherein the first transcribing instance starts transcription of the audio stream at a first point of time and the second transcribing instance starts transcription of the audio stream at a second point of time with a predetermined delay with respect to the first transcribing instance:

wherein the predetermined delay is selected such that each one of the plurality of second text data snippets overlaps with at least an ending portion of a respective first transcribed text data snippet of the plurality of the first text data snippets ends and also overlaps with a starting portion of the first transcribed text data snippet of the plurality of the first text data snippets that is consecutive to the respective first transcribed text data snippet;

concatenating the first transcribed text data snippets and the second transcribed text data snippets, the concatenating of the first and second transcribed text data snippets comprising identifying matching text passages in overlapping portions of the first and second transcribed text data snippets, wherein the identifying of the matching text passages comprises:

identifying at least one syllable pattern according to a Porter-Stemmer algorithm in the first transcribed text data snippets; and

in response to the at least one syllable pattern being identified in the first transcribed text data snippets, searching the identified at least one syllable pattern in the second transcribed text data snippets.

14. The computer-implemented method according to claim 13 , wherein the transcription service is a non real-time transcription service.

15. A transcription mechanism for a communication system for carrying out a video and/or audio conference with at least two participants, wherein the transcription mechanism is adapted to carry out the method of claim 13 .

16. The computer-implemented method of claim 13 , comprising: displaying the transcribed audio stream via a display device.

17. A transcription mechanism for a communication system for carrying out a video and/or audio conference with at least two participants, the transcription mechanism comprising:

a computer device having a processor connected to a non-transitory computer readable medium, the computer device positionable in a communication network and communicatively connectable to at least two participant communication devices of the at least two participants to a video and/or audio conference, the computer device configured to:

transcribe an audio stream of the video and/or audio conference using a first transcribing instance, the first transcribing instance having a first predetermined transcription size that is smaller than a total length of the audio stream, the first transcribing instance providing a plurality of consecutive first transcribed text data snippets of the audio stream, each of the first transcribed text data snippets having a size corresponding to the first predetermined transcription size;

transcribe the audio stream using at least a second transcribing instance having a second predetermined transcription size that is smaller than the length of the audio stream, the second transcribing instance providing a plurality of consecutive second transcribed text data snippets of the audio stream, each of the second transcribed text data snippets having a size corresponding to the second predetermined transcription size;

wherein the first transcribing instance is configured to start transcription of the audio stream at a first point of time and the second transcribing instance is configured to start transcription of the audio stream at a second point of time with a predetermined delay with respect to the first transcribing instance, predetermined delay being configured such that each one of the plurality of second text data snippets overlaps with at least an ending portion of a respective first transcribed text data snippet of the plurality of the first text data snippets ends and also overlaps with a starting portion of the first transcribed text data snippet of the plurality of the first text data snippets that is consecutive to the respective first transcribed text data snippet;

the computer device also configured to:

identify matching text passages in overlapping portions of the first and second transcribed text data snippets via at least one of:

identify at least one word pattern in the first transcribed text data snippet, the at least one word pattern comprising at least two long words with a predetermined number of short words in between the two long words and, in response to the at least one word pattern being identified in the first transcribed text data snippet, search the identified at least one word pattern in the second transcribed text data snippets; and

identify at least one syllable pattern according to a Porter-Stemmer algorithm in the first transcribed text data snippets and, in response to the at least one syllable pattern being identified in the first transcribed text data snippets, search the identified at least one syllable pattern in the second transcribed text data snippets.

18. The transcription mechanism of claim 17 , wherein the at least two participant communication devices comprises laptop computers, telephones, tablets, and/or smart phones.

Assignments (9)
RELEASE OF SECURITY INTEREST Recorded Jun 24, 2025
From: WILMINGTON SAVINGS FUND SOCIETY, FSB
To: MITEL (DELAWARE), INC.; MITEL COMMUNICATIONS, INC.; MITEL NETWORKS, INC.; MITEL NETWORKS CORPORATION
Reel/Frame 071712/0821 →
NOTICE OF SUCCCESSION OF AGENCY - PL Recorded Jan 14, 2025
From: UBS AG, STAMFORD BRANCH, AS LEGAL SUCCESSOR TO CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: WILMINGTON SAVINGS FUND SOCIETY, FSB
Reel/Frame 069895/0755 →
NOTICE OF SUCCCESSION OF AGENCY - 3L Recorded Jan 14, 2025
From: UBS AG, STAMFORD BRANCH, AS LEGAL SUCCESSOR TO CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: WILMINGTON SAVINGS FUND SOCIETY, FSB
Reel/Frame 070006/0268 →
NOTICE OF SUCCCESSION OF AGENCY - 2L Recorded Jan 14, 2025
From: UBS AG, STAMFORD BRANCH, AS LEGAL SUCCESSOR TO CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
To: WILMINGTON SAVINGS FUND SOCIETY, FSB
Reel/Frame 069896/0001 →
CHANGE OF NAME Recorded Oct 24, 2024
From: UNIFY PATENTE GMBH & CO. KG
To: UNIFY BETEILIGUNGSVERWALTUNG GMBH & CO. KG
Reel/Frame 069242/0312 →
SECURITY INTEREST Recorded Jan 5, 2024
From: UNIFY PATENTE GMBH & CO. KG
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 066197/0333 →
SECURITY INTEREST Recorded Jan 5, 2024
From: UNIFY PATENTE GMBH & CO. KG
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 066197/0299 →
SECURITY INTEREST Recorded Jan 5, 2024
From: UNIFY PATENTE GMBH & CO. KG
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH, AS COLLATERAL AGENT
Reel/Frame 066197/0073 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2021
From: HERMANNS, LARS; NASS, THOMAS; MOERS, STEFAN; REIF, FRANK
To: UNIFY PATENTE GMBH & CO. KG
Reel/Frame 056630/0782 →
Priority Claims (1)
EP 20182036 · Jun 24, 2020 · regional
Continuity (1)
Related Publication 20210407515A1 · Dec 30, 2021
Cited By (1)
US 12,412,581