IP Library Granted Patent US 10,522,135
Granted Patent B2
US 10,522,135 · App. 15/859,611 · Granted Dec 31, 2019

System and method for segmenting audio files for transcription

Inventors: Tom Livne (Ramat Gan, IL); Kobi Ben Tzvi (Ramat Hasharon, IL); Eric Shellef (Givaatayim, IL)
Assignee: Verbit Software Ltd.
G10L15/04G10L15/16G10L15/1815G10L25/48G10L15/005G10L15/02G10L17/005G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,522,135
App. No.
15/859,611
Granted
Dec 31, 2019
Kind
B2
Abstract

A system and method for segmenting an audio file. The method includes analyzing an audio file, wherein the analyzing includes identifying speech recognition features within the audio file; generating metadata based on the audio file, wherein the metadata includes transcription characteristics of the audio file; and determining a segmenting interval for the audio file based on the speech recognition features and the metadata.

Claims (24)

1. A method for segmenting an audio file at optimal points that allow efficient transcribing, comprising:

analyzing an audio file, wherein the analyzing includes (i) identifying number of speakers in the audio file based on generating a signature for each voice determined to be unique within the audio file, and (ii) identifying accent of each speaker based on at least one of a Gaussian mixture model (GMM), a GMM Support Vector Machine (GMM-SVM) and GMM Universal Background Model (GMM-UBM);

generating metadata based on the audio file, wherein the metadata includes language spoken by each speaker; and

segmenting the audio file at optimal points without contextual or linguistic interruptions between segments, based on the number of speakers, the accent of each speaker, and the metadata.

2. The method of claim 1 , further comprising identifying that the audio file contains confidential information, limiting the pool of eligible candidates for providing the transcription services to include only those who have been identified as having passed a confidentiality clearance sufficient for the relevant audio file, and selecting at least one candidate from the pool of eligible candidates.

3. The method of claim 1 , wherein a first speaker within the audio file is speaking English, a second speaker within the audio file is speaking French, and the audio file is segmented into a first segment and a second segment, such that the first segment includes the speech by the first speaker that is assigned to a first transcription provider capable of English transcription, and the second segment includes the speech by the second speaker that is assigned to a second transcription provider capable of French transcription.

4. The method of claim 1 , wherein the analyzing of the audio file further utilizes at least one of the following techniques: tonal context, linguistic context, voice activity, Linear Predictive Coding (LPC), Perceptional Linear Predictive Coefficients (PLP), Mel-Frequency Cepstral Coefficients (MFCC), Linear Prediction Cepstral Coefficients (LPCC), Wavelet Based Features and Non-Negative Matrix Factorization features.

5. The method of claim 1 , wherein at least one of the optimal points comprises a transition between a first speaker and a second speaker.

6. The method of claim 5 , wherein at least one of the optimal points comprises a shift in recording location.

7. The method of claim 1 , wherein at least one of the optimal points comprises a change in spoken language.

8. The method of claim 1 , wherein the analyzing the at least one audio file further comprises: employing a deep learning technique, and further comprising forwarding a first segment of the segments to a first transcription provider and forwarding a second segment of the segments to a second transcription provider.

9. The method of claim 8 , wherein the deep learning technique includes as least one of: a neural network algorithm, decision tree learning, clustering, homomorphic filtering, wideband reducing filtering, and sound wave anti-aliasing algorithms.

10. A system for segmenting an audio file at optimal points that allow efficient transcribing, comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

analyze an audio file, wherein the analyzing includes (i) identifying number of speakers in the audio file based on generating a signature for each voice determined to be unique within the audio file, and (ii) identifying accent of each speaker based on at least one of a Gaussian mixture model (GMM), a GMM Support Vector Machine (GMM-SVM) and GMM Universal Background Model (GMM-UBM);

generate metadata based on the audio file, wherein the metadata includes language spoken by each speaker; and

segmenting the audio file at optimal points without contextual or linguistic interruptions between segments, based on the number of speakers, the accent of each speaker, and the metadata.

11. The system of claim 10 , wherein the system is further configured to identify that the audio file contains confidential information, limit the pool of eligible candidates for providing the transcription services to include only those who have been identified as having passed a confidentiality clearance sufficient for the relevant audio file, and select at least one candidate from the pool of eligible candidates.

12. The system of claim 10 , wherein at least one of the optimal points comprises a transition between a first speaker and a second speaker.

13. The system of claim 12 , wherein at least one of the optimal points comprises a shift in recording location.

14. The system of claim 10 , wherein at least one of the optimal points comprises a change in spoken language.

15. The system of claim 10 , wherein the analyzing the at least one audio file further comprises: employ a deep learning technique, and further comprising forwarding a first segment of the segments to a first transcription provider and forwarding a second segment of the segments to a second transcription provider.

16. The system of claim 12 , wherein the deep learning technique includes as least one of: a neural network algorithm, decision tree learning, clustering, homomorphic filtering, wideband reducing filtering, and sound wave anti-aliasing algorithms.

Assignments (4)
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Feb 3, 2021
From: VERBIT SOFTWARE LTD
To: SILICON VALLEY BANK
Reel/Frame 055209/0474 →
SECURITY INTEREST Recorded Jun 4, 2020
From: VERBIT SOFTWARE LTD.
To: SILICON VALLEY BANK
Reel/Frame 052836/0856 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 28, 2019
From: LIVNE, TOM; BEN TZVI, KOBI; SHELLEF, ERIC
To: VERBIT SOFTWARE LTD.
Reel/Frame 049293/0389 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded May 24, 2019
From: VERBIT SOFTWARE LTD
To: SILICON VALLEY BANK
Reel/Frame 049285/0338 →
Continuity (2)
Provisional Application 62510293 · May 24, 2017
Related Publication 20180342235A1 · Nov 29, 2018