IP Library Patent Application 18033758
Patent Application
App. No. 18/033,758

AUDIO SIGNAL CONVERSION MODEL LEARNING APPARATUS, AUDIO SIGNAL CONVERSION APPARATUS, AUDIO SIGNAL CONVERSION MODEL LEARNING METHOD AND PROGRAM

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/033,758
Abstract

A voice signal conversion model learning device including: a data-for-learning acquisition unit that acquires input data for learning that is a voice signal input; a conversion learning model execution unit that executes a conversion learning model that converts the input data for learning into learning stage conversion destination data; and an update unit that updates the conversion learning model by learning, in which: a probability density function is defined as a target feature amount distribution function, the probability density function being a function on a vector space representing a series of voice feature amounts and representing a distribution of a series of voice feature amounts of a target voice signal that is a voice signal having a predetermined attribute; a point is defined as an initial value point, the point being in the vector space and representing a series of feature amounts of the input data for learning; a function is defined as a score function, the function having a point x in the vector space as an independent variable and indicating a gradient of a path from the point x to a stationary point that is on the target feature amount distribution function and is nearest to the initial value point; the conversion learning model execution unit performs conversion of the input data for learning on the basis of the score function; and the update unit updates the score function in updating the conversion learning model.

Claims (38)

1 . A voice signal conversion model learning device comprising:

a processor; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

acquiring input data for learning, the input data being a voice signal input;

executing a conversion learning model that is a model of machine learning that converts the input data for learning into learning stage conversion destination data that is a voice signal of a conversion destination; and

updating the conversion learning model by learning,

wherein

a probability density function is defined as a target feature amount distribution function, the probability density function being a function on a vector space representing a series of voice feature amounts that are feature amounts obtained from a voice signal and representing a distribution of a series of voice feature amounts of a target voice signal that is a voice signal having a predetermined attribute,

a point is defined as an initial value point, the point being in the vector space and representing a series of feature amounts of the input data for learning,

a function is defined as a score function, the function having a point x in the vector space as an independent variable and indicating a gradient of a path from the point x to a nearest stationary point that is a stationary point on the target feature amount distribution function and is a stationary point nearest to the initial value point,

the input data for learning is converted into the learning stage conversion destination data on a basis of the score function in the executing, and

the score function in updating the conversion learning model in the updating.

2 . The voice signal conversion model learning device according to claim 1 , wherein

a neural network is defined as a score approximator, the neural network representing a function that includes a parameter θ and in which a result of predetermined optimization processing of updating the parameter θ is substantially identical to a score function,

a neural network representing the conversion learning model includes a plurality of the score approximators, and

the score function is updated in the updating on a basis of a sum of differences for the respective score approximators, wherein each of the differences is a difference between a value of the score function and a difference between data of the point x to which noise is added and data of the point x of the space before the noise is added.

3 . The voice signal conversion model learning device according to claim 2 , wherein

a method for updating the score function on a basis of the sum is weighted Denoising Score Matching (DSM).

4 . The voice signal conversion model learning device according to claim 1 , wherein

a neural network is defined as a score approximator, the neural network representing a function that includes a parameter θ and in which a result of predetermined optimization processing of updating the parameter θ is substantially identical to a score function,

a neural network representing the conversion learning model includes a single piece of the score approximator, and

the score function is updated in the updating on a basis of a sum of a plurality of differences included in the score approximator, wherein each of the differences is a difference between a value of the score function and a difference between data of the point x to which noise is added and data of the point x of the space before the noise is added.

5 . A voice signal conversion device comprising:

a processor; and

a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of:

acquiring a voice signal of a conversion target; and

performing conversion of the conversion target by using a learned conversion learning model obtained by a voice signal conversion model learning device comprising: a processor; and a storage medium having computer program instructions stored thereon, wherein the computer program instruction, when executed by the processor, perform processing of: acquiring input data for learning, the input data being a voice signal input; executing a conversion learning model that is a model of machine learning that converts the input data for learning into learning stage conversion destination data that is a voice signal of a conversion destination; and updating the conversion learning model by learning, wherein a probability density function is defined as a target feature amount distribution function, the probability density function being a function on a vector space representing a series of voice feature amounts that are feature amounts obtained from a voice signal and representing a distribution of a series of voice feature amounts of a target voice signal that is a voice signal having a predetermined attribute, a point is defined as an initial value point, the point being in the vector space and representing a series of feature amounts of the input data for learning, a function is defined as a score function, the function having a point x in the vector space as an independent variable and indicating a gradient of a path from the point x to a nearest stationary point that is a stationary point on the target feature amount distribution function and is a stationary point nearest to the initial value point, the input data for learning is converted into the learning stage conversion destination data on a basis of the score function in the executing, and the score function in updating the conversion learning model in the updating.

6 . A voice signal conversion model learning method comprising:

acquiring input data for learning, the input data being a voice signal input;

executing a conversion learning model that is a model of machine learning that converts the input data for learning into learning stage conversion destination data that is a voice signal of a conversion destination; and

updating the conversion learning model by learning,

wherein

a probability density function is defined as a target feature amount distribution function, the probability density function being a function on a vector space representing a series of voice feature amounts that are feature amounts obtained from a voice signal and representing a distribution of a series of voice feature amounts of a target voice signal that is a voice signal having a predetermined attribute,

a point is defined as an initial value point, the point being in the vector space and representing a series of feature amounts of the input data for learning,

a function is defined as a score function, the function having a point x in the vector space as an independent variable and indicating a gradient of a path from the point x to a nearest stationary point that is a stationary point on the target feature amount distribution function and is a stationary point nearest to the initial value point,

in the executing, the input data for learning is converted into the learning stage conversion destination data on a basis of the score function, and

in the updating, the score function is updated in updating the conversion learning model.

7 . A non-transitory computer readable medium which stores a program for causing a computer to function as the voice signal conversion model learning device according to claim 1 .

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073007/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 25, 2023
From: KAMEOKA, HIROKAZU
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 063445/0303 →