IP Library Granted Patent US 12,266,362
Granted Patent B2
US 12,266,362 · App. 17/087,330 · Granted Apr 1, 2025

Systems and methods for a two pass diarization, automatic speech recognition, and transcript generation

Inventors: Jean-Philippe Robichaud (Mercier, CA); Alexei Skurikhin (Redwood City, CA); Migüel Jetté (Squamish, CA); Petrov Evgeny Stanislavovich (Saint Petersburg, RU)
Assignee: Rev.com, Inc.
G10L15/26G10L17/00G10L19/038G10L15/02G10L15/22G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,266,362
App. No.
17/087,330
Granted
Apr 1, 2025
Kind
B2
Abstract

In one embodiment, a method for transcript generation includes receiving an audio file and dividing it into a plurality of chunks. The method further includes sending each instance of the plurality of chunks to a speech service module. The method further includes converting speech to text for each instance of the plurality of chunks and returning the text for each instance of the plurality of chunks. The method further includes merging the text for each instance of the plurality of chunks to yield an audio file transcript and sending the audio file and chunks to a diarization module. The method further includes performing first pass diarization on the chunks to yield a plurality of diarized chunks and performing second pass diarization on the plurality of diarized chunks and the audio file to yield a diarized audio file. The method further includes merging the files to yield a final transcript.

Claims (36)

1. A method of performing diarization on a sound recording, the method comprising:

receiving a sound recording;

breaking the sound recording into a plurality of chunks;

performing a first diarization on the plurality of chunks, wherein the performing the first diarization on the plurality of chunks occurs simultaneously, and wherein the performing includes breaking each of the plurality of chunks into a plurality of segments, for each of the plurality of segments generating statistical speaker information descriptive of the sound characteristics in that segment, and clustering, within each chunk of the plurality of chunks, segments having similar statistical speaker information to generate within each chunk of the plurality of chunks groups of segments grouped according to the similar statistical speaker information;

performing a second diarization by clustering between the plurality of chunks, the groups of segments according to grouped similar statistical speaker information, the grouped similar statistical speaker information being characteristics of speech of each group for the groups of segments, wherein the second diarization performs a modified I-Vector scoring, based on the groups of segments according to grouped similar statistical speaker information, I-vectors of the groups of segments according to grouped similar statistical speaker information are averaged and then compared to other averaged I-vectors, where a closeness of two or more averaged I-vectors is compared accordingly clustered based on similarity;

creating a new i-vector for the groups of segments according to grouped similar statistical speaker information.

2. The method of claim 1 , further comprising:

transcoding the sound recording according to a known codec.

3. The method of claim 1 , further comprising:

creating a sound recording transcript from the sound recording;

sending the sound recording transcript to a post process module;

applying punctuation and casing to the sound recording transcript.

4. The method of claim 1 , wherein the speaker identification information is an I-vector.

5. The method of claim 1 , wherein the second diarization includes giving each of a plurality of speakers for each of the plurality of diarized chunks a unique identifier.

6. The method of claim 5 , wherein the second diarization includes, for associated segments of the plurality of segments for each unique identifier, averaging the speaker identification information of the associated segments to yield averaged speaker identification information.

7. The method of claim 6 , wherein the second diarization includes, assigning identified segments of the plurality of segments from all of the plurality of chunks a final speaker based on correlation between the averaged speaker identification information for the associated segments of the plurality of segments for each unique identifier.

8. The method of claim 3 , further comprising: creating a final transcript from the sound recording transcript; and outputting a final transcript in a fixed and tangible format.

9. A system for performing diarization on a sound recording, the system comprising:

a diarization module configured to

receive a sound recording;

break the sound recording into a plurality of chunks;

perform a first diarization on the plurality of chunks, wherein the first diarization on the plurality of chunks occurs simultaneously, and wherein the performing includes breaking each of the plurality of chunks into a plurality of segments, for each of the plurality of segments generating statistical speaker information descriptive of the sound characteristics in that segment, and clustering, within each chunk of the plurality of chunks, segments having similar statistical speaker information to generate within each chunk of the plurality of chunks groups of segments grouped according to the similar statistical speaker information;

perform a second diarization by clustering between the plurality of chunks, the groups of segments according to grouped similar statistical speaker information, the grouped similar statistical speaker information being characteristics of speech of each group for the groups of segments, wherein the second diarization performs a modified I-Vector scoring, based on the groups of segments according to grouped similar statistical speaker information, I-vectors of the groups of segments according to grouped similar statistical speaker information are averaged and then compared to other averaged I-vectors, where a closeness of two or more averaged I-vectors is compared accordingly clustered based on similarity;

creating a new i-vector for the groups of segments according to grouped similar statistical speaker information.

10. The system of claim 9 , wherein the diarization module is further configured to transcode the sound recording according to a known codec.

11. The method of claim 9 , wherein the diarization module is further configured to

create a sound recording transcript from the sound recording;

send the sound recording transcript to a post process module;

apply punctuation and casing to the sound recording transcript.

12. The method of claim 9 , wherein the speaker identification information is an I-vector.

13. The method of claim 9 , wherein the second diarization includes giving each of a plurality of speakers for each of the plurality of diarized chunks a unique identifier.

14. The method of claim 13 , wherein the second diarization includes, for associated segments of the plurality of segments for each unique identifier, averaging the speaker identification information of the associated segments to yield averaged speaker identification information.

15. The method of claim 14 , wherein the second diarization includes, assigning identified segments of the plurality of segments from all of the plurality of chunks a final speaker based on correlation between the averaged speaker identification information for the associated segments of the plurality of segments for each unique identifier.

16. The method of claim 11 , wherein the diarization module is further configured to create a final transcript from the sound recording transcript; and outputting a final transcript in a fixed and tangible format.

17. The method of claim 1 , wherein the groups of segments grouped according to the similar statistical speaker information are clustered as speakers.

18. The method of claim 9 , wherein the groups of segments grouped according to the similar statistical speaker information are clustered as speakers.

Assignments (4)
SECURITY INTEREST Recorded Aug 8, 2024
From: REV.COM, INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 068224/0202 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2022
From: ROBICHAUD, JEAN-PHILIPPE; SKURIKHIN, ALEXEI; JETTÉ, MIGÜEL; STANISLAVOVICH, PETROV EVGENY
To: REV.COM, INC.
Reel/Frame 061338/0069 →
SECURITY INTEREST Recorded Nov 30, 2021
From: REV.COM, INC.
To: SILICON VALLEY BANK, AS AGENT
Reel/Frame 058246/0707 →
SECURITY INTEREST Recorded Nov 30, 2021
From: REV.COM, INC.
To: SILICON VALLEY BANK
Reel/Frame 058246/0828 →
Continuity (2)
Continuation 16177061 · Oct 31, 2018
Related Publication 20210050015A1 · Feb 18, 2021
References Cited (26)
US 10476872B2 · McLaren · 2019 [cited by examiner]
US 10964329B2 · Ghaemmaghami · 2021 [cited by examiner]
US 11056118B2 · Martínez · 2021 [cited by examiner]
US 11138334B1 · Garrod · 2021 [cited by examiner]
US 11410175B2 · Howald · 2022 [cited by examiner]
US 20110302489A1 · Zimmerman · 2011 [cited by examiner]
US 20120059656A1 · Garland et al. · 2012 [cited by applicant]
US 20130166285A1 · Chang et al. · 2013 [cited by applicant]
US 20140074467A1 · Ziv et al. · 2014 [cited by applicant]
US 20160217792A1 · Gorodetski · 2016 [cited by examiner]
US 20160283185A1 · McLaren et al. · 2016 [cited by applicant]
US 20170084295A1 · Tsiartas et al. · 2017 [cited by applicant]
US 20170372706A1 · Shepstone et al. · 2017 [cited by applicant]
US 20180211670A1 · Gorodetski et al. · 2018 [cited by applicant]
CN 103700370A · 2014 [cited by applicant]
CN 104485105A · 2015 [cited by applicant]
CN 107210038A · 2017 [cited by applicant]
WO WO2018009969A1 · 2018 [cited by applicant]
International Search Report and Written Opinion dated Mar. 3, 2020 issued in related PCT App. No. PCT/US19/58870 (19 pages). [cited by applicant]
Extended European Search Report dated Jun. 29, 2022 issued in co-pending European patent app. No. 19877574.4 (15 pages). [cited by applicant]
Bohac et al: “Post-processing of the recognized speech for web presentation of large audio archive”, Telecommunications and Signal Processing (TSP) 2012 35 [cited by applicant]
Silovsky et al. “Speaker diarization f broadcast streams using two-stage clustering based on i-vectors and cosine distance scoring”, 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP … [cited by applicant]
Patent Examination Report dated Jun. 30, 2022 issued in co-pending New Zealand patent app. No. 774716 (3 pages). [cited by applicant]
Office Action dated Feb. 14, 2023 issued in co-pending New Zealand patent application No. 774716 (11 pages). [cited by applicant]
Marek Bohec et al, ‘Post processing of the recognized speech for we presentation of large audio archive’, Telecommunications and Signal Processing (TSP), 2012 35th International Conference on, IEEE, Jul. 3, 2012, pp. 44… [cited by applicant]
Office Action dated Aug. 10, 2023 issued in related Chinese patent application No. 201980070755.X (11 pages). [cited by applicant]