IP Library Granted Patent US 12,700,413
Granted Patent B2
US 12,700,413 · App. 17/678,405 · Granted Aug 4, 2026

Frequency mapping in the voiceprint domain

Inventors: Claudio Vair (Borgone Susa, IT); Haydar Talib (Montreal, CA); Kevin Robert Farrell (Brick, NJ); Daniele Ernesto Colibro (Alessandria, IT)
Assignee: Microsoft Technology Licensing, LLC
G10L17/18G06N3/047G10L17/00G10L17/04G10L17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,700,413
App. No.
17/678,405
Filed
Feb 23, 2022
Granted
Aug 4, 2026
Kind
B2
Art Unit
2654
USPC
704/232
Abstract

There is provided a method that includes (a) obtaining a first voice vector that was derived from a signal of a voice that was sampled at a first sampling frequency, (b) obtaining a second voice vector that was derived from a signal of a voice that was sampled at a second sampling frequency, (c) mapping the second voice vector into a mapped voice vector in accordance with a machine learning model, and (d) comparing the first voice vector to the mapped voice vector to yield a score that indicates a probability that the first voice vector and the second voice vector originated from a same person.

Claims (68)

1 . A method comprising:

training, during a training process, a neural net mapping model of a wideband (WB) speaker recognition system to perform narrowband-to-wideband (NB-to-WB) voice vector transformation exclusively using a WB training dataset that excludes known narrowband (NB) voice vectors associated with known speaker identities, the WB training dataset being composed of WB audio samples produced by sampling utterances of people at a WB sampling frequency,

wherein training the neural net mapping model comprises:

processing one of the WB audio samples from the WB training dataset via a WB processing path to obtain a target WB voice vector:

processing the WB audio sample via a NB processing path, that is independent from the WB processing path, to obtain an input NB voice vector; and

regressively training the neural net mapping model to transform the input NB voice vector into an approximated WB voice vector that reduces distortion between the approximated WB voice vector and the target WB voice vector,

wherein the neural net mapping model is regressively trained using a loss function measuring the distortion between the approximated WB voice vector and the target WB voice vector; and

after the training process has concluded, receiving an input WB voice vector of a speaker and performing, by the WB speaker recognition system based on the known NB voice vectors, WB speaker identification on the input WB voice vector using the trained neural net mapping model, the input WB voice vector produced by a user device directly sampling human voice audio of the speaker at the WB sampling frequency,

wherein performing WB speaker identification on the input WB voice vector includes:

mapping the input WB voice vector to one of the known NB voice vectors;

using the trained neural net mapping model to transform the known NB voice vector into a representative WB voice vector:

comparing the input WB voice vector with the representative WB voice vector to determine a score that indicates a probability that the input WB voice vector and the known NB voice vector originated from a same person; and

identifying the speaker based on a corresponding one of the known speaker identities associated with the known NB voice vector when the score exceeds a second threshold,

wherein the trained neural net mapping model transforms the known NB voice vector into the representative WB voice vector during a time period between reception of the input WB voice vector and identification of the speaker by the WB speaker recognition system.

2 . The method of claim 1 , wherein processing the WB audio sample via the NB processing path to generate the input NB voice vector comprises down sampling the WB audio sample to obtain a NB audio sample and converting the NB audio sample into the input NB voice vector.

3 . The method of claim 2 , wherein the WB audio sample has a sampling frequency of 16 kilohertz (kHz) and the NB audio sample has a sampling frequency of 8 kHz.

4 . The method of claim 1 , wherein processing the WB audio sample via the WB processing path comprises:

extracting acoustic features from the WB audio sample;

converting the extracted acoustic features into feature coefficients; and

generating the target WB voice vector based on the feature coefficients using a deep neural network (DNN).

5 . The method of claim 4 , wherein the feature coefficients include Mel frequency cepstral coefficients (MFCCs).

6 . The method of claim 4 , wherein the feature coefficients include linear prediction cepstral coefficients (LPCCs).

7 . The method of claim 4 , wherein the feature coefficients include perceptual linear predictive (PLP) cepstral coefficients.

8 . The method of claim 4 , wherein the target WB voice vector includes floating-point values conveying biometric information associated with the WB audio sample.

9 . The method of claim 8 , wherein the target WB voice vector is a DNN-Embeddings wideband voice vector.

10 . The method of claim 4 , wherein the DNN is a time Delay DNN or a factorized time Delay DNN.

11 . The method of claim 1 , wherein training the neural net mapping model to transform the input NB voice vector into the approximated WB voice vector further comprises:

performing regression neural network training of the neural net mapping model to minimize distortion between the approximated WB voice vector and the target WB voice vector, wherein the distortion is measured using a Maximum Absolute Error or Mean Square Error loss function to compare the approximated WB voice vector with the target WB voice vector.

12 . The method of claim 1 , wherein the neural net mapping model is regressively trained using a loss function measuring the distortion between the approximated WB voice vector and the target WB voice vector.

13 . The method of claim 1 , wherein the known NB voice vectors are derived from NB audio samples produced from directly sampling human voice audio at a NB sampling frequency, and wherein the WB training dataset excludes the NB audio samples.

14 . A system comprising:

a processor; and

a memory storing programming instructions for execution by the processor, wherein the programming instructions, upon execution by the processor, cause the processor to perform the following operations:

training, during a training process, a neural net mapping model of a wideband (WB) speaker recognition system to perform narrowband-to-wideband (NB-to-WB) voice vector transformation exclusively using a WB training dataset that excludes known narrowband (NB) voice vectors associated with known speaker identities, the WB training dataset being composed of WB audio samples produced by sampling utterances of people at a WB sampling frequency,

wherein training the neural net mapping model comprises;

processing one of the WB audio samples from the WB training dataset via a WB processing path to obtain a target WB voice vector:

processing the WB audio sample via a NB processing path, that is independent from the WB processing path, to obtain an input NB voice vector; and

regressively training the neural net mapping model to transform the input NB voice vector into an approximated WB voice vector that reduces distortion between the approximated WB voice vector and the target WB voice vector,

wherein the neural net mapping model is regressively trained using a loss function measuring the distortion between the approximated WB voice vector and the target WB voice vector; and

after the training process has concluded, receiving an input WB voice vector of a speaker and performing, by the WB speaker recognition system based on the known NB voice vectors, WB speaker identification on the input WB voice vector using the trained neural net mapping model, the input WB voice vector produced by a user device directly sampling human voice audio of the speaker at the WB sampling frequency,

wherein performing WB speaker identification on the input WB voice vector includes:

mapping the input WB voice vector to one of the known NB voice vectors;

using the trained neural net mapping model to transform the known NB voice vector into a representative WB voice vector:

comparing the input WB voice vector with the representative WB voice vector to determine a score that indicates a probability that the input WB voice vector and the known NB voice vector originated from a same person; and

identifying the speaker based on a corresponding one of the known speaker identities associated with the known NB voice vector when the score exceeds a second threshold,

wherein the trained neural net mapping model transforms the known NB voice vector into the representative WB voice vector during a time period between reception of the input WB voice vector and identification of the speaker by the WB speaker recognition system.

15 . The system of claim 14 , wherein processing the WB audio sample via the NB processing path to generate the input NB voice vector includes down sampling the WB audio sample to obtain a NB audio sample and converting the NB audio sample into the input NB voice vector.

16 . The system of claim 14 , wherein processing the WB audio sample via the WB processing path includes:

extracting acoustic features from the WB audio sample;

converting the extracted acoustic features into feature coefficients; and

generating the target WB voice vector based on the feature coefficients using a deep neural network (DNN).

17 . The system of claim 16 , wherein the feature coefficients include Mel frequency cepstral coefficients (MFCCs).

18 . The system of claim 16 , wherein the feature coefficients include linear prediction cepstral coefficients (LPCCs).

19 . The system of claim 16 , wherein the feature coefficients include perceptual linear predictive (PLP) cepstral coefficients.

20 . A storage device that is non-transitory, comprising instructions that are readable by a processor to cause said processor to perform operations of:

training, during a training process, a neural net mapping model of a wideband (WB) speaker recognition system to perform narrowband-to-wideband (NB-to-WB) voice vector transformation exclusively using a WB training dataset that excludes known narrowband (NB) voice vectors associated with known speaker identities, the WB training dataset being composed of WB audio samples produced by sampling utterances of people at a WB sampling frequency,

wherein training the neural net mapping model comprises;

processing one of the WB audio samples from the WB training dataset via a WB processing path to obtain a target WB voice vector;

processing the WB audio sample via a NB processing path, that is independent from the WB processing path, to obtain an input NB voice vector, and

regressively training the neural net mapping model to transform the input NB voice vector into an approximated WB voice vector that reduces distortion between the approximated WB voice vector and the target WB voice vector,

wherein the neural net mapping model is regressively trained using a loss function measuring the distortion between the approximated WB voice vector and the target WB voice vector; and

after the training process has concluded, receiving an input WB voice vector of a speaker and performing, by the WB speaker recognition system based on the known NB voice vectors, WB speaker identification on the input WB voice vector using the trained neural net mapping model, the input WB voice vector produced by a user device directly sampling human voice audio of the speaker at the WB sampling frequency,

wherein performing WB speaker identification on the input WB voice vector includes:

mapping the input WB voice vector to one of the known NB voice vectors;

using the trained neural net mapping model to transform the known NB voice vector into a representative WB voice vector;

comparing the input WB voice vector with the representative WB voice vector to determine a score that indicates a probability that the input WB voice vector and the known NB voice vector originated from a same person; and

identifying the speaker based on a corresponding one of the known speaker identities associated with the known NB voice vector when the score exceeds a second threshold,

wherein the trained neural net mapping model transforms the known NB voice vector into the representative WB voice vector during a time period between reception of the input WB voice vector and identification of the speaker by the WB speaker recognition system.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065210/0694 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2022
From: VAIR, CLAUDIO; TALIB, HAYDAR; FARRELL, KEVIN ROBERT; CALIBRO, DANIELE ERNESTO
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 059076/0694 →
Continuity (1)
Related Publication 20230267936A1 · Aug 24, 2023
References Cited (21)
US 6988066B2 · Malah · 2006 [cited by examiner]
US 11985179B1 · Tacer · 2024 [cited by examiner]
US 20120201399A1 · Mitsufuji · 2012 [cited by examiner]
US 20140233725A1 · Kim · 2014 [cited by examiner]
US 20190043488A1 · Bocklet · 2019 [cited by examiner]
US 20210005182A1 · Han · 2021 [cited by examiner]
US 20210182357A1 · Partee · 2021 [cited by examiner]
US 20210183358A1 · Mao · 2021 [cited by examiner]
US 20210241776A1 · Sivaraman · 2021 [cited by examiner]
US 20210264923A1 · Lesso · 2021 [cited by examiner]
US 20230005486A1 · Chen · 2023 [cited by examiner]
US 20230260531A1 · Srivastava · 2023 [cited by examiner]
US 20230267936A1 · Vair · 2023 [cited by examiner]
US 20230371889A1 · Weston · 2023 [cited by examiner]
CN 113516987A · 2021 [cited by examiner]
CN 113611295B · 2024 [cited by examiner]
E. Variani, E. McDermott and G. Heigold, “A Gaussian Mixture Model layer jointly optimized with discriminative features within a Deep Neural Network architecture,” 2015 IEEE International Conference on Acoustics, Speech… [cited by examiner]
Shahina et al., “Mapping Neural Networks for Bandwidth Extension of Narrowband Speech”, Interspeech 2006, pp. 1435-1438. [cited by applicant]
Macho, Duan, “Narrowband to Wideband Feature Expansion for Robust Multilingual ASR”, Interspeech 2007, pp. 1118-1121. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US23/060672”, Mailed Date: Mar. 31, 2023, 12 Pages. [cited by applicant]
Chen et al., “Speaker Embedding Conversion for Backward and Cross-Channel Compatibility,” ICASSP 2022—2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), May 23-27, 2022, 5 pages. [cited by applicant]