IP Library Granted Patent US 12,475,904
Granted Patent B2
US 12,475,904 · App. 18/032,529 · Granted Nov 18, 2025

Audio signal conversion model learning apparatus, audio signal conversion apparatus, audio signal conversion model learning method and program

Inventors: Takuhiro Kaneko (Musashino, JP); Hirokazu Kameoka (Musashino, JP); Ko Tanaka (Musashino, JP); Nobukatsu Hojo (Musashino, JP)
Assignee: NTT, Inc.
G10L21/007G06N20/00G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,904
App. No.
18/032,529
Granted
Nov 18, 2025
Kind
B2
Abstract

The present invention is a voice signal conversion model learning device provided with: a learning data acquisition unit for acquiring learning input data, which is an input voice signal; and a learning stage conversion unit for executing a conversion learning model, which is a machine learning model, including learning stage conversion processing for converting the learning input data into learning stage conversion-destination data, which is a conversion-destination voice signal. The learning stage conversion processing includes local feature acquisition processing for acquiring, on the basis of input data to be processed, a feature for each learning input-side subset. The conversion learning model further includes adjustment-parameter-value acquisition processing for acquiring, on the basis of the learning input data, an adjustment parameter value. The learning stage conversion processing converts the learning input data into the learning stage conversion-destination data by using a result of a predetermined computation based on the adjustment parameter value.

Claims (23)

1 . A voice signal conversion model learning device comprising:

a processor; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

acquiring learning input data which is an input voice signal; and

executing a conversion learning model which is a model of machine learning including learning stage conversion processing of converting the learning input data into learning stage conversion destination data which is a voice signal of a conversion destination, wherein

the learning stage conversion processing includes local feature quantity acquisition processing of acquiring a feature quantity for each learning input-side subset which is a subset of processing target input data having the processing target input data as a population, based on the processing target input data which is data to be processed,

the conversion learning model further includes adjustment parameter value acquisition processing of acquiring an adjustment parameter value, which is a value of a parameter for adjusting a statistical value of a distribution of the feature quantity, based on the learning input data, and

the learning stage conversion processing converts the learning input data into the learning stage conversion destination data using a result of a predetermined calculation based on the adjustment parameter value, where the predetermined calculation converts the feature quantity according to the adjustment parameter value using affine conversion.

2 . The voice signal conversion model learning device according to claim 1 , wherein

the processing of converting the feature quantity is executed for each element of the feature quantity.

3 . A voice signal conversion device comprising:

a processor; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

acquiring a voice signal to be converted; and

converting the conversion target using a learned conversion learning model obtained by a voice signal conversion model learning device including a processor; and a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of: acquiring learning input data which is an input voice signal; and executing a conversion learning model which is a model of machine learning including learning stage conversion processing of converting the learning input data into learning stage conversion destination data which is a voice signal of a conversion destination, wherein the learning stage conversion processing includes local feature quantity acquisition processing of acquiring a feature quantity for each learning input-side subset which is a subset of processing target input data having the processing target input data as a population, based on the processing target input data which is data to be processed, the conversion learning model further includes adjustment parameter value acquisition processing of acquiring an adjustment parameter value, which is a value of a parameter for adjusting a statistical value of a distribution of the feature quantity, based on the learning input data, and the learning stage conversion processing converts the learning input data into the learning stage conversion destination data using a result of a predetermined calculation based on the adjustment parameter value, where the predetermined calculation converts the feature quantity according to the adjustment parameter value using affine conversion.

4 . A voice signal conversion model learning method comprising:

acquiring learning input data which is an input voice signal; and

executing a conversion learning model which is a model of machine learning including learning stage conversion processing of converting the learning input data into learning stage conversion destination data which is a voice signal of a conversion destination, wherein

the learning stage conversion processing includes local feature quantity acquisition processing of acquiring a feature quantity for each learning input-side subset which is a subset of processing target input data having the processing target input data as a population, based on the processing target input data which is data to be processed,

the conversion learning model further includes adjustment parameter value acquisition processing of acquiring an adjustment parameter value, which is a value of a parameter for adjusting a statistical value of a distribution of the feature quantity, based on the learning input data, and

the learning stage conversion processing converts the learning input data into the learning stage conversion destination data using a result of a predetermined calculation based on the adjustment parameter value, where the predetermined calculation converts the feature quantity according to the adjustment parameter value using affine conversion.

5 . A non-transitory computer readable medium which stores a program for causing a computer to function as the voice signal conversion model learning device according to claim 1 .

6 . The voice signal conversion model learning device of claim 1 wherein the feature quantity includes the input voice signal or acoustic feature quantities derived from the input voice signal.

Assignments (2)
CHANGE OF NAME Recorded Sep 30, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 072989/0863 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2023
From: KANEKO, TAKUHIRO; KAMEOKA, HIROKAZU; TANAKA, KO; HOJO, NOBUKATSU
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 063367/0121 →
Continuity (1)
Related Publication 20230386489A1 · Nov 30, 2023
References Cited (23)
US 11003955B1 · Tan · 2021 [cited by applicant]
US 20120059654A1 · Nishimura et al. · 2012 [cited by applicant]
US 20160078859A1 · Luan · 2016 [cited by examiner]
US 20190172480A1 · Kaskari et al. · 2019 [cited by applicant]
US 20190385628A1 · Nakashika · 2019 [cited by examiner]
US 20200395028A1 · Kameoka et al. · 2020 [cited by applicant]
US 20200410976A1 · Zhou · 2020 [cited by examiner]
US 20210201890A1 · Wang et al. · 2021 [cited by applicant]
US 20220068259A1 · Pan · 2022 [cited by examiner]
US 20220156552A1 · Kaneko et al. · 2022 [cited by applicant]
US 20220246136A1 · Yang · 2022 [cited by examiner]
US 20230360631A1 · Takamichi et al. · 2023 [cited by applicant]
JP 2019035902A · 2019 [cited by applicant]
JP 2019101391A · 2019 [cited by applicant]
JP 2019144402A · 2019 [cited by applicant]
JP 2020140244A · 2020 [cited by applicant]
“Shindong Lee et al., ““Many-To-Many Voice Conversion Using Conditional Cycle-Consistent Adversarial Networks””, ICASSP 2020—2020 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 9… [cited by applicant]
“Hirokazu Kameoka et al., ““ConvS2S-VC: Fully Convolutional Sequenceto-Sequence Voice Conversion””, IEEE/ACM Transactions on Audio, Speech, and Language Processing, Jun. 23, 2020, vol. 28 pp. 1849-1863 DOI: 10.1109/TASL… [cited by applicant]
Jon Magnus Momrak Haug, “Voice Conversion using Deep Learning”, 2017, Norwegian University of Science and Technology (Year: 2017). [cited by applicant]
Takuhiro Kaneko et al., CycleGAN-VC3: Examining and Improving CycleGAN-VCs for Mel-spectrogram Conversion, arXiv preprint, arXiv:2010.11672, Online, Oct. 22, 2020, retrieved on Nov. 5, 2020, Internet: URL: https://arxiv… [cited by applicant]
Hirokazu Kameoka et al., StarGAN-VC: Non-Parallel Many-to-Many Voice Conversion With Star Generative Adversarial Networks, arXiv:1806.02169v2, 2018. [cited by applicant]
Takuhiro Kaneko et al., CycleGAN-VC2: Improved CycleGAN-Based Non-Parallel Voice Conversion, ICASSP2019. [cited by applicant]
Takuhiro Kaneko et al., StarGAN-VC2: Rethinking Conditional Methods for StarGAN-Based Voice Conversion, arXiv:1907.12279v2, 2019. [cited by applicant]