IP Library Granted Patent US 11,574,639
Granted Patent B2
US 11,574,639 · App. 17/127,938 · Granted Feb 7, 2023

Hypothesis stitcher for speech recognition of long-form audio

Inventors: Naoyuki Kanda (Bellevue, WA); Xuankai Chang (Baltimore, MD); Yashesh Gaur (Redmond, WA); Xiaofei Wang (Bellevue, WA); Zhong Meng (Mercer Island, WA); Takuya Yoshioka (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L17/02G10L15/22G10L15/26G10L19/022G10L21/0272
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,574,639
App. No.
17/127,938
Granted
Feb 7, 2023
Kind
B2
Abstract

A hypothesis stitcher for speech recognition of long-form audio provides superior performance, such as higher accuracy and reduced computational cost. An example disclosed operation includes: segmenting the audio stream into a plurality of audio segments; identifying a plurality of speakers within each of the plurality of audio segments; performing automatic speech recognition (ASR) on each of the plurality of audio segments to generate a plurality of short-segment hypotheses; merging at least a portion of the short-segment hypotheses into a first merged hypothesis set; inserting stitching symbols into the first merged hypothesis set, the stitching symbols including a window change (WC) symbol; and consolidating, with a network-based hypothesis stitcher, the first merged hypothesis set into a first consolidated hypothesis. Multiple variations are disclosed, including alignment-based stitchers and serialized stitchers, which may operate as speaker-specific stitchers or multi-speaker stitchers, and may further support multiple options for differing hypothesis configurations.

Claims (48)

1. A method of speech recognition, the method comprising:

segmenting an audio stream into a plurality of audio segments;

identifying a plurality of speakers within the audio stream;

performing automatic speech recognition (ASR) on each of the plurality of audio segments to generate a plurality of short-segment hypotheses;

merging a first portion of the short-segment hypotheses into a first merged hypothesis set specific to a first speaker of the plurality of speakers;

merging a second portion of the short-segment hypotheses into a second merged hypothesis set specific to a second speaker of the plurality of speakers, the second speaker different than the first speaker;

inserting stitching symbols into the first merged hypothesis set and the second merged hypothesis set, the stitching symbols including a window change (WC) symbol; and

consolidating, with a network-based hypothesis stitcher, the first merged hypothesis set into a first consolidated hypothesis and the second merged hypothesis set into a second consolidated hypothesis.

2. The method of claim 1 , further comprising:

outputting the first consolidated hypothesis as a first transcription and the second consolidated hypothesis as a second transcription.

3. The method of claim 1 , wherein the first merged hypothesis set is specific to a first speaker of the plurality of speakers, wherein the first consolidated hypothesis is specific to the first speaker.

4. The method of claim 1 , wherein the first merged hypothesis set comprises a multi-speaker merged hypothesis set, and wherein the stitching symbols further include a speaker identification.

5. The method of claim 1 , wherein the hypothesis stitcher comprises an alignment-based stitcher, wherein the first merged hypothesis set comprises an odd hypothesis sequence and an even hypothesis sequence, and wherein the method further comprises:

aligning the odd hypothesis sequence with the even hypothesis sequence.

6. The method of claim 1 , wherein the hypothesis stitcher comprises a serialized stitcher that does not use an alignment of odd and even hypothesis sequences.

7. The method of claim 6 , wherein the hypothesis stitcher uses 25% overlap or less.

8. A system for speech recognition, the system comprising:

a processor; and

a computer-readable medium storing instructions that are operative upon execution by the processor to:

segment an audio stream into a plurality of audio segments;

identify a plurality of speakers within the audio stream;

perform automatic speech recognition (ASR) on each of the plurality of audio segments to generate a plurality of short-segment hypotheses;

merge a first portion of the short-segment hypotheses into a first merged hypothesis set specific to a first speaker of the plurality of speakers;

merging a second portion of the short-segment hypotheses into a second merged hypothesis set specific to a second speaker of the plurality of speakers, the second speaker different than the first speaker;

insert stitching symbols into the first merged hypothesis set and the second merged hypothesis set, the stitching symbols including a window change (WC) symbol; and

consolidate, with a network-based hypothesis stitcher, the first merged hypothesis set into a first consolidated hypothesis and the second merged hypothesis set into a second consolidated hypothesis.

9. The system of claim 8 , wherein the first merged hypothesis set comprises hypotheses ranks.

10. The system of claim 8 , wherein the first merged hypothesis set is specific to a first speaker of the plurality of speakers, wherein the first consolidated hypothesis is specific to the first speaker.

11. The system of claim 8 , wherein the first merged hypothesis set comprises a multi-speaker merged hypothesis set, and wherein the stitching symbols further include a speaker identification.

12. The system of claim 8 , wherein the hypothesis stitcher comprises an alignment-based stitcher, wherein the first merged hypothesis set comprises an odd hypothesis sequence and an even hypothesis sequence, and wherein the instructions are further operative to:

align the odd hypothesis sequence with the even hypothesis sequence.

13. The system of claim 8 , wherein the hypothesis stitcher comprises a serialized stitcher that does not use an alignment of odd and even hypothesis sequences.

14. The system of claim 13 , wherein the hypothesis stitcher uses 25% overlap or less.

15. One or more computer storage devices having computer-executable instructions stored thereon, which, on execution by a computer, cause the computer to perform operations comprising:

segmenting an audio stream into a plurality of audio segments;

identifying a plurality of speakers within the audio stream;

performing automatic speech recognition (ASR) on each of the plurality of audio segments to generate a plurality of short-segment hypotheses;

merging a first portion of the short-segment hypotheses into a first merged hypothesis set specific to a first speaker of the plurality of speakers;

merging a second portion of the short-segment hypotheses into a second merged hypothesis set specific to a second speaker of the plurality of speakers, the second speaker different than the first speaker;

inserting stitching symbols into the first merged hypothesis set and the second merged hypothesis set, the stitching symbols including a window change (WC) symbol; and

consolidating, with a network-based hypothesis stitcher, the first merged hypothesis set into a first consolidated hypothesis and the second merged hypothesis set into a second consolidated hypothesis.

16. The one or more computer storage devices of claim 15 , wherein the operations further comprise:

outputting the first consolidated hypothesis as a first transcription and the second consolidated hypothesis as a second transcription.

17. The one or more computer storage devices of claim 15 , wherein the first merged hypothesis set is specific to a first speaker of the plurality of speakers, wherein the first consolidated hypothesis is specific to the first speaker.

18. The one or more computer storage devices of claim 15 , wherein the first merged hypothesis set comprises a multi-speaker merged hypothesis set, and wherein the stitching symbols further include a speaker identification.

19. The one or more computer storage devices of claim 15 , wherein the hypothesis stitcher comprises an alignment-based stitcher, wherein the first merged hypothesis set comprises an odd hypothesis sequence and an even hypothesis sequence, and wherein the operations further comprise:

aligning the odd hypothesis sequence with the even hypothesis sequence.

20. The one or more computer storage devices of claim 15 , wherein the hypothesis stitcher comprises a serialized stitcher that does not use an alignment of odd and even hypothesis sequences.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2020
From: KANDA, NAOYUKI; CHANG, XUANKAI; GAUR, YASHESH; WANG, XIAOFEI; MENG, ZHONG; YOSHIOKA, TAKUYA
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 054702/0029 →
Continuity (1)
Related Publication 20220199091A1 · Jun 23, 2022