IP Library Granted Patent US 12,277,949
Granted Patent B2
US 12,277,949 · App. 18/342,025 · Granted Apr 15, 2025

Method and audio processing device for voice anonymization

Inventor: Knut Inge Hvidsten (Oslo, NO)
Assignee: PEXIP AS
G10L21/013H04L12/1831G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,949
App. No.
18/342,025
Granted
Apr 15, 2025
Kind
B2
Abstract

A method and audio processing device for voice anonymization in an audio- or videoconferencing session. The method comprises receiving a plurality of input audio samples comprising speech, calculating a frequency spectrum of each the plurality of input audio samples, calculating a smoothed spectral magnitude envelope of a first of the plurality of frequency spectrums to determine a plurality of formant features of the speech, each of the plurality of formant features being located at different frequencies in the frequency spectrum, determining one random scaling factor for the audio- or videoconferencing session, determining, based on the one random scaling factor, a voice anonymization function shifting the formant location of at least one of the plurality of formants, and applying the voice anonymization function on the frequency spectrum of each the subsequent plurality of input audio samples in the audio- or videoconferencing session.

Claims (35)

1. A method for voice anonymization in an audio- or videoconferencing session, the method comprising:

receiving a plurality of input audio samples comprising speech;

calculating a frequency spectrum of each the plurality of input audio samples;

calculating a smoothed spectral magnitude envelope of a first of the plurality of frequency spectrums to determine a plurality of formant features of the speech, each of the plurality of formant features being located at different frequencies in the frequency spectrum;

determining one random scaling factor for the audio- or videoconferencing session;

determining, based on the one random scaling factor, a voice anonymization function shifting the formant location of at least one of the plurality of formants;

applying the voice anonymization function on the frequency spectrum of each the subsequent plurality of input audio samples in the audio- or videoconferencing session;

determining the one random scaling factor by using a random function to pick a number from two or more ranges of scaling factors;

wherein the voice anonymization function is a linear segment warping function performing linear scaling in the range 0-4 kHz;

the voice anonymization function is tapering off to zero warp at one half of a sampling frequency; and

determining a plurality of frequency gains for the voice anonymization function by calculating a ratio between the smoothed spectral magnitude envelope and a spectral magnitude envelope of the voice anonymization function applied on the smoothed spectral magnitude envelope.

2. A method for voice anonymization in an audio- or videoconferencing session, the method comprising:

receiving a plurality of input audio samples comprising speech;

calculating a frequency spectrum of each the plurality of input audio samples;

calculating a smoothed spectral magnitude envelope of a first of the plurality of frequency spectrums to determine a plurality of formant features of the speech, each of the plurality of formant features being located at different frequencies in the frequency spectrum;

determining one random scaling factor for the audio- or videoconferencing session;

determining, based on the one random scaling factor, a voice anonymization function shifting the formant location of at least one of the plurality of formants;

applying the voice anonymization function on the frequency spectrum of each the subsequent plurality of input audio samples in the audio- or videoconferencing session; and

determining a plurality of frequency gains for the voice anonymization function by calculating a ratio between the smoothed spectral magnitude envelope and a spectral magnitude envelope of the voice anonymization function applied on the smoothed spectral magnitude envelope.

3. The method of claim 2 , wherein the calculating of the frequency spectrum of each of the plurality of input samples is performed by a filterbank.

4. The method of claim 3 , wherein the filterbank is a Short-Time Fourier Transform filterbank.

5. An audio processing device for an audio- or videoconferencing session, the audio processing device:

receiving a plurality of input audio samples comprising speech:

calculating a frequency spectrum of each the plurality of input audio samples;

calculating a smoothed spectral magnitude envelope of a first of the plurality of frequency spectrums to determine a plurality of formant features of the speech, each of the plurality of formant features being located at different frequencies in the frequency spectrum;

determining one random scaling factor for the audio- or videoconferencing session;

determining, based on the one random scaling factor, a voice anonymization function shifting the formant location of at least one of the plurality of formants;

applying the voice anonymization function on the frequency spectrum of each the subsequent plurality of input audio samples in the audio- or videoconferencing session; and

determining a plurality of frequency gains for the voice anonymization function by calculating a ratio between the smoothed spectral magnitude envelope and a spectral magnitude envelope of the voice anonymization function applied on the smoothed spectral magnitude envelope.

6. The audio processing device of claim 5 , adapted to determining the one random scaling factor by using a random function to pick a number from two or more ranges of scaling factors.

7. The audio processing device of claim 5 , wherein the voice anonymization function is a linear segment warping function performing linear scaling in the range 0-4 kHz.

8. The audio processing device of claim 7 , wherein the voice anonymization function is tapering off to zero warp at one half of a sampling frequency.

9. The audio processing device of claim 5 , comprising a filterbank calculating the frequency spectrum of each of the plurality of input samples.

10. The audio processing device of claim 9 , wherein the filterbank is a Short-Time Fourier Transform filterbank.

11. The audio processing device of claim 10 , wherein the audio processing device is integrated in at least one of a multipoint conferencing node, MCN, and a videoconferencing terminal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: HVIDSTEN, KNUT INGE
To: PEXIP AS
Reel/Frame 064077/0201 →
Priority Claims (1)
NO 20220759 · Jul 1, 2022 · national
Continuity (1)
Related Publication 20240005936A1 · Jan 4, 2024
References Cited (13)
US 8892448B2 · Vos · 2014 [cited by examiner]
US 10432687B1 · Hanes · 2019 [cited by examiner]
US 20050249272A1 · Kirkeby · 2005 [cited by examiner]
US 20210089626A1 · Ross · 2021 [cited by examiner]
US 20210400142A1 · Jorasch et al. · 2021 [cited by applicant]
US 20230351059A1 · Springer · 2023 [cited by examiner]
WO 2021152566A1 · 2021 [cited by applicant]
Patino, J., Tomashenko, N., Todisco, M., Nautsch, A., & Evans, N. (2020). Speaker anonymisation using the McAdams coefficient. arXiv preprint arXiv:2011.01130. [cited by examiner]
Jose Patino et al., Speaker anonymisation using the McAdams coefficient, arxiv.org, Cornell University Library, 201, Sep. 1, 2021, 5-pages. [cited by applicant]
Kai Hiroto et al., Lightweight Voice Anonymization Based on Data-Driven Optimization of Cascaded Voice Modification Modules, Tokyo Metropolitan University, 2021, 7-pages. [cited by applicant]
Lal Srivastava Brij Mohan et al., Evaluating Voice Conversion-Based Privacy Protection Against Informed Attackers, 2020 IEEE, 5 pages. [cited by applicant]
European Patent Office, International-Type Search Report for corresponding Norwegian Application No. 20220759, dated Feb. 2, 2023, 14-pages. [cited by applicant]
Norwegian Search Report for corresponding Norwegian Application No. 20220759, dated Jan. 26, 2023, 2-pages. [cited by applicant]