IP Library Granted Patent US 12688856
Granted Patent B2
US 12688856 · App. 18/410,789 · Granted Jul 21, 2026

Processing of audio data in multi-speaker multi-channel environments

Inventors: Taejin Park (San Jose, CA); Ante Jukic (Culver City, CA); He Huang (Greenville, SC); Venkata Naga Krishna Chaitanya Puvvada (San Jose, CA); Kunal Dhawan (San Jose, CA); Nithin Rao Koluguri (Milpitas, CA); Nikolay Karpov (Moscow, AM); Aleksandr Laptev (Erevan, AM); Jagadeesh Balam (Campbell, CA)
Assignee: NVIDIA Corporation
G10L17/18G10L17/02G10L17/20G10L25/30G10L25/78G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688856
App. No.
18/410,789
Granted
Jul 21, 2026
Kind
B2
Abstract

Disclosed are apparatuses, systems, and techniques that may use machine learning for implementing speaker recognition, verification, and/or diarization. The techniques include processing audio data channels (ADCs) using a voice detection model to determine voice activity likelihoods (VALs) that individual ADCs include speech, obtaining, using VALs, a second set of ADC(s), and processing, using an audio processing a neural network (NN) model, the second set of ADCs to obtain association of the speech to the one or more speakers. The techniques also include generating a plurality of embeddings associated with the ADCs, processing the plurality of embeddings to obtain aggregated embedding(s) that represent audio data of multiple ADCs, and processing the aggregated embedding(s), using the audio processing NN model, to obtain association of the speech to the one or more speakers.

Claims (69)

1 . A method comprising:

obtaining a first set of embeddings, each embedding of the first set of embeddings associated with (i) a respective audio data channel (ADC) of a set of ADCs jointly capturing speech produced by one or more speakers and (ii) a respective time frame of a plurality of time frames;

identifying, for individual time frames of the plurality of time frames, a subset of embeddings from the first set based at least on similarity among the first set of embeddings;

obtaining a second set of embeddings, each embedding of the second set of embeddings obtained by combining embeddings of the subset associated with a corresponding time frame; and

processing, using an audio processing neural network (NN) model, the second set of embeddings to obtain an association of the speech to the one or more speakers.

2 . The method of claim 1 , wherein the obtaining the first set of embeddings comprises:

determining, using a voice detection model, voice activity likelihoods (VALs) of individual ADCs of the set of ADCs; and

eliminating ADCs of the set of ADCs having a VAL below a VAL threshold.

3 . The method of claim 2 , wherein an input into the voice detection model comprises the first set of embeddings.

4 . The method of claim 1 , wherein the second set of embeddings is obtained for a sliding time window comprising the plurality of time frames, and wherein an input into the audio processing NN model comprises

the second set of embeddings obtained for the sliding time window.

5 . The method of claim 1 , wherein the set of ADCs is obtained from an initial set of ADCs, wherein individual ADCs of the set of ADCs represent one or more channels of the initial set of ADCs, and wherein at least one ADC of the set of ADCs represents a cluster of two or more ADCs of the initial set of ADCs, the two or more ADCs being selected based on similarity of audio data of the two or more ADCs.

6 . The method of claim 5 , wherein the set of ADCs is obtained using at least one of:

obtaining a similarity matrix that characterizes similarity of the audio data of individual ADCs of the initial set of ADCs; or

processing, using a clustering NN model, the audio data of the initial set of ADCs to estimate similarity of the audio data of individual ADCs of the initial set of ADCs.

7 . The method of claim 6 , wherein an element (j, k) of the similarity matrix characterizes similarity of the audio data of j-th ADC of the initial set of ADCs and the audio data of k-th ADC of the initial set of ADCs;

and wherein obtaining the set of ADCs further comprises:

identifying, using the similarity matrix, one or more clusters of ADCs of the initial set of ADCs; and

using the one or more clusters of ADCs to obtain the set of ADCs.

8 . A method comprising:

generating a plurality of embeddings associated with a set of audio data channels (ADCs), the set of ADCs jointly capturing a speech produced by one or more speakers;

processing, using one or more neural network (NN) models, the plurality of embeddings to obtain one or more aggregated embeddings, individual aggregated embeddings obtained by aggregating a subset of the plurality of embeddings, the subset selected, for individual temporal units of the speech, based on similarity of the plurality of embeddings associated with said individual temporal units of the speech; and

processing, using an audio processing NN model, the one or more aggregated embeddings to obtain an association of the speech to the one or more speakers.

9 . The method of claim 8 , wherein the one or more NN models comprise a voice detection model, and wherein generating the plurality of embeddings comprises:

processing, using the voice detection model, a plurality of initial embeddings to determine voice activity likelihoods (VALs) that individual embeddings of the plurality of initial embeddings represent speech; and

eliminating, using the VALs, one or more embeddings from the plurality of initial embeddings to obtain the plurality of embeddings.

10 . The method of claim 9 , wherein the eliminating the one or more embeddings from the plurality of initial embeddings to obtain the plurality of embeddings comprises:

eliminating one or more low-VAL embeddings from the plurality of initial embeddings.

11 . The method of claim 8 , wherein the subset of the plurality of embeddings has a predetermined number of embeddings.

12 . The method of claim 9 , wherein individual temporal units of the speech correspond to individual spectrogram frames of the speech.

13 . The method of claim 8 , wherein the generating the plurality of embeddings comprises:

obtaining, using an initial set of ADCs, the set of ADCs, wherein individual ADCs of the set of ADCs represent one or more channels of the initial set of ADCs, and wherein at least one ADC of the set of ADCs represents a cluster of two or more ADCs of the initial set of ADCs, the two or more ADCs being selected based on similarity of audio data of the two or more ADCs; and

wherein individual embeddings of the plurality of embeddings are generated for respective ADCs of the set of ADCs.

14 . The method of claim 13 , wherein the obtaining the set of ADCs comprises at least one of:

obtaining a similarity matrix that characterizes similarity of the audio data of individual ADCs of the initial set of ADCs; or

processing, using a clustering NN model of the one or more NN models, the audio data of the initial set of ADCs to estimate similarity of the audio data of individual ADCs of the initial set of ADCs.

15 . The method of claim 14 , wherein an element (j, k) of the similarity matrix characterizes similarity of the audio data of j-th ADC of the initial set of ADCs and the audio data of k-th ADC of the initial set of ADCs;

and wherein obtaining the set of ADCs further comprises:

identifying, using the similarity matrix, one or more clusters of ADCs of the initial set of ADCs; and

using the one or more clusters of ADCs to obtain the set of ADCs.

16 . A system comprising:

one or more processing units to:

obtain a first set of embeddings, each embedding of the first set of embeddings associated with (i) a respective audio data channel (ADC) of a set of ADCs jointly capturing speech produced by one or more speakers and (ii) a respective time frame of a plurality of time frames;

process the first set of embeddings to identify, for individual time frames of the plurality of time frames, a subset of embeddings from the first set based at least on similarity among the first set of embeddings;

obtain a second set of embeddings, each embedding of the second set of embeddings obtained by combining embeddings of the subset associated with a corresponding time frame; and

process, using an audio processing neural network (NN) model, the second set of embeddings to obtain an association of the speech to the one or more speakers.

17 . The system of claim 16 , wherein to obtain the first set of embeddings, the one or more processing units are to:

determine, using a voice detection model, voice activity likelihoods (VALs) of individual ADCs of the set of ADCs; and

eliminate ADCs of the set of ADCs having a VAL below a VAL threshold.

18 . The system of claim 16 , wherein the second set of embeddings is obtained for a sliding time window comprising the plurality of time framess, and wherein an input into the audio processing NN model comprises:

the second set of embeddings obtained for the sliding time window.

19 . The system of claim 16 , wherein the system is comprised in at least one of:

an in-vehicle infotainment system for an autonomous or semi-autonomous machine;

a system for performing one or more simulation operations;

a system for performing one or more digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing one or more deep learning operations;

a system implemented using an edge device;

a system for generating or presenting at least one of virtual reality content, mixed reality content, or augmented reality content;

a system implemented using a robot;

a system for performing one or more conversational AI operations;

a system implementing one or more large language models (LLMs);

a system implementing one or more language models;

a system for performing one or more generative AI operations;

a system for generating synthetic data;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.