Processing of audio data in multi-speaker multi-channel environments
Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker recognition, verification, and/or diarization. The techniques include processing audio data channels (ADCs) using a voice detection model to determine voice activity likelihoods (VALs) that individual ADCs include speech, obtaining, using VALs, a second set of ADC(s), and processing, using an audio processing a neural network (NN) model, the second set of ADCs to obtain association of the speech to the one or more speakers. The techniques also include generating a plurality of embeddings associated with the ADCs, processing the plurality of embeddings to obtain aggregated embedding(s) that represent audio data of multiple ADCs, and processing the aggregated embedding(s), using the audio processing NN model, to obtain association of the speech to the one or more speakers.
1 . A method comprising:
obtaining a first set of embeddings, each embedding of the first set of embeddings associated with (i) a respective audio data channel (ADC) of a set of ADCs jointly capturing speech produced by one or more speakers and (ii) a respective time frame of a plurality of time frames;
identifying, for individual time frames of the plurality of time frames, a subset of embeddings from the first set based at least on similarity among the first set of embeddings;
obtaining a second set of embeddings, each embedding of the second set of embeddings obtained by combining embeddings of the subset associated with a corresponding time frame; and
processing, using an audio processing neural network (NN) model, the second set of embeddings to obtain an association of the speech to the one or more speakers.
2 . The method of claim 1 , wherein the obtaining the first set of embeddings comprises:
determining, using a voice detection model, voice activity likelihoods (VALs) of individual ADCs of the set of ADCs; and
eliminating ADCs of the set of ADCs having a VAL below a VAL threshold.
3 . The method of claim 2 , wherein an input into the voice detection model comprises the first set of embeddings.
4 . The method of claim 1 , wherein the second set of embeddings is obtained for a sliding time window comprising the plurality of time frames, and wherein an input into the audio processing NN model comprises
the second set of embeddings obtained for the sliding time window.
5 . The method of claim 1 , wherein the set of ADCs is obtained from an initial set of ADCs, wherein individual ADCs of the set of ADCs represent one or more channels of the initial set of ADCs, and wherein at least one ADC of the set of ADCs represents a cluster of two or more ADCs of the initial set of ADCs, the two or more ADCs being selected based on similarity of audio data of the two or more ADCs.
6 . The method of claim 5 , wherein the set of ADCs is obtained using at least one of:
obtaining a similarity matrix that characterizes similarity of the audio data of individual ADCs of the initial set of ADCs; or
processing, using a clustering NN model, the audio data of the initial set of ADCs to estimate similarity of the audio data of individual ADCs of the initial set of ADCs.
7 . The method of claim 6 , wherein an element (j, k) of the similarity matrix characterizes similarity of the audio data of j-th ADC of the initial set of ADCs and the audio data of k-th ADC of the initial set of ADCs;
and wherein obtaining the set of ADCs further comprises:
identifying, using the similarity matrix, one or more clusters of ADCs of the initial set of ADCs; and
using the one or more clusters of ADCs to obtain the set of ADCs.
8 . A method comprising:
generating a plurality of embeddings associated with a set of audio data channels (ADCs), the set of ADCs jointly capturing a speech produced by one or more speakers;
processing, using one or more neural network (NN) models, the plurality of embeddings to obtain one or more aggregated embeddings, individual aggregated embeddings obtained by aggregating a subset of the plurality of embeddings, the subset selected, for individual temporal units of the speech, based on similarity of the plurality of embeddings associated with said individual temporal units of the speech; and
processing, using an audio processing NN model, the one or more aggregated embeddings to obtain an association of the speech to the one or more speakers.
9 . The method of claim 8 , wherein the one or more NN models comprise a voice detection model, and wherein generating the plurality of embeddings comprises:
processing, using the voice detection model, a plurality of initial embeddings to determine voice activity likelihoods (VALs) that individual embeddings of the plurality of initial embeddings represent speech; and
eliminating, using the VALs, one or more embeddings from the plurality of initial embeddings to obtain the plurality of embeddings.
10 . The method of claim 9 , wherein the eliminating the one or more embeddings from the plurality of initial embeddings to obtain the plurality of embeddings comprises:
eliminating one or more low-VAL embeddings from the plurality of initial embeddings.
11 . The method of claim 8 , wherein the subset of the plurality of embeddings has a predetermined number of embeddings.
12 . The method of claim 9 , wherein individual temporal units of the speech correspond to individual spectrogram frames of the speech.
13 . The method of claim 8 , wherein the generating the plurality of embeddings comprises:
obtaining, using an initial set of ADCs, the set of ADCs, wherein individual ADCs of the set of ADCs represent one or more channels of the initial set of ADCs, and wherein at least one ADC of the set of ADCs represents a cluster of two or more ADCs of the initial set of ADCs, the two or more ADCs being selected based on similarity of audio data of the two or more ADCs; and
wherein individual embeddings of the plurality of embeddings are generated for respective ADCs of the set of ADCs.
14 . The method of claim 13 , wherein the obtaining the set of ADCs comprises at least one of:
obtaining a similarity matrix that characterizes similarity of the audio data of individual ADCs of the initial set of ADCs; or
processing, using a clustering NN model of the one or more NN models, the audio data of the initial set of ADCs to estimate similarity of the audio data of individual ADCs of the initial set of ADCs.
15 . The method of claim 14 , wherein an element (j, k) of the similarity matrix characterizes similarity of the audio data of j-th ADC of the initial set of ADCs and the audio data of k-th ADC of the initial set of ADCs;
and wherein obtaining the set of ADCs further comprises:
identifying, using the similarity matrix, one or more clusters of ADCs of the initial set of ADCs; and
using the one or more clusters of ADCs to obtain the set of ADCs.
16 . A system comprising:
one or more processing units to:
obtain a first set of embeddings, each embedding of the first set of embeddings associated with (i) a respective audio data channel (ADC) of a set of ADCs jointly capturing speech produced by one or more speakers and (ii) a respective time frame of a plurality of time frames;
process the first set of embeddings to identify, for individual time frames of the plurality of time frames, a subset of embeddings from the first set based at least on similarity among the first set of embeddings;
obtain a second set of embeddings, each embedding of the second set of embeddings obtained by combining embeddings of the subset associated with a corresponding time frame; and
process, using an audio processing neural network (NN) model, the second set of embeddings to obtain an association of the speech to the one or more speakers.
17 . The system of claim 16 , wherein to obtain the first set of embeddings, the one or more processing units are to:
determine, using a voice detection model, voice activity likelihoods (VALs) of individual ADCs of the set of ADCs; and
eliminate ADCs of the set of ADCs having a VAL below a VAL threshold.
18 . The system of claim 16 , wherein the second set of embeddings is obtained for a sliding time window comprising the plurality of time framess, and wherein an input into the audio processing NN model comprises:
the second set of embeddings obtained for the sliding time window.
19 . The system of claim 16 , wherein the system is comprised in at least one of:
an in-vehicle infotainment system for an autonomous or semi-autonomous machine;
a system for performing one or more simulation operations;
a system for performing one or more digital twin operations;
a system for performing light transport simulation;
a system for performing collaborative content creation for 3D assets;
a system for performing one or more deep learning operations;
a system implemented using an edge device;
a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;
a system implemented using a robot;
a system for performing one or more conversational AI operations;
a system implementing one or more large language models (LLMs);
a system implementing one or more language models;
a system for performing one or more generative AI operations;
a system for generating synthetic data;
a system incorporating one or more virtual machines (VMs);
a system implemented at least partially in a data center; or
a system implemented at least partially using cloud computing resources.