IP Library Granted Patent US 11,900,947
Granted Patent B2
US 11,900,947 · App. 17/184,323 · Granted Feb 13, 2024

Method and system for automatically diarising a sound recording

Inventors: Houman Ghaemmaghami (Tarragindi, AU); Shahram Kalantari (Kangaroo Point, AU); David Dean (Auchenflower, AU); Subramanian Sridharan (Brisbane, AU)
Assignee: FTR LABS PTY LTD
G10L17/06G10L17/02G10L17/04G10L21/0272G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,900,947
App. No.
17/184,323
Granted
Feb 13, 2024
Kind
B2
Abstract

Embodiments of the present invention provide methods and systems for performing automatic diarisation of sound recordings including speech from one more speakers. The automatic diarisation has a development or training phase and a utilisation or evaluation phase. In the development or training phase background models and hyperparameters are generated from already annotated sound recordings. These models and hyperparameters are applied during the evaluation or utilisation phase to diarise new or not previously diarised or annotated recordings.

Claims (16)

1. A method of developing models in an automatic diarisation system comprising a system controller, a digital signal processor and a memory, the method comprising the steps of:

a) retrieving from the memory one or more annotated audio recordings by the controller, each annotated audio recording comprising an audio file and associated annotation data identifying timing in the file for changes between speakers, changes from active speech to non-speech, and differentiation of individual speakers;

b) performing, using the digital signal processor, feature extraction on the one or more annotated audio recordings, wherein the feature extraction is performed based on frames of samples, processed to extract speech feature components compatible for reliable modelling using Gaussian mixture modelling (GMM);

c) using the controller, the controller using the annotation data for accumulating a set of all non-speech features for each recording;

d) using the controller, the controller forming from the annotation data a set comprising speech features for all speakers for each recording;

e) developing a universal background model (UBM) for each recording using the set comprising speech features for all speakers for each recording, wherein the universal background model (UBM) is developed as an extensive Gaussian mixture model using an Expectation-Maximisation (EM) algorithm;

f) storing the UBM for input to automated diarisation processing;

g) performing speech Gaussian mixture modelling (GMM) for each recording using sampled data from the set comprising speech features for all speakers for the recording and storing the speech GMM for each recording;

h) performing GMM for each recording using sampled data from the set comprising all non-speech features for the recording and storing the non-speech GMM for each recording, wherein the speech GMMs, non-speech GMMs are stored for input to voice activity detection in an automated diarisation process;

i) using the controller, the controller forming from the annotation data for each recording a plurality of sets of speech features for each speaker, each set representing speech features for a set short time duration, and computing statistics for each set;

j) performing joint factor analysis (JFA) hyperparameter training for each speaker using the plurality of sets of speech features for the speaker, the computed statistics for the speaker and the UBM to determine characterising speaker JFA hyperparameters for the speaker, and storing the speaker JFA hyperparameters for each speaker;

k) computing statistics for non-speech features for each set of all non-speech features for each recording;

l) performing JFA hyperparameter training for non-speech using the set of all non-speech features for each recording, the computed statistics for the non-speech and the UBM to determine characterising non-speech JFA hyperparameters for non-speech, and storing the non-speech JFA hyperparameters for each recording;

whereby the speaker JFA Hyperparameters and non-speech JFA hyperparameters are input to speaker identification in the automated diarisation process.

2. The method as claimed in claim 1 wherein in step i) the short time duration l is chosen as a length sufficient to exhibit linguistic variation over the length of the short time duration and to divide a longer period of speech into short time duration segments that can each be treated as a separate session.

3. The method as claimed in claim 2 wherein the short time duration l is chosen from within the range of 5 to 20 seconds.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 14, 2026
From: FTR LABS PTY LTD ACN 146 022 175
To: FTR LLC
Reel/Frame 075267/0638 →
NUNC PRO TUNC ASSIGNMENT Recorded Nov 29, 2022
From: GHAEMMAGHAMI, HOUMAN; KALANTARI, SHAHRAM; DEAN, DAVID; SRIDHARAN, SUBRAMANIAN
To: FTR PTY LTD
Reel/Frame 061909/0932 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 29, 2022
From: FTR PTY LTD
To: FTR LABS PTY LTD
Reel/Frame 061909/0947 →
Priority Claims (1)
AU 2016902710 · Jul 11, 2016 · national
Continuity (2)
Division 16316708
Related Publication 20210183395A1 · Jun 17, 2021
Cited By (3)
US 12,482,487 US 12,488,778 US 12,573,370