IP Library › Granted Patent US 12,217,755
Granted Patent B2
US 12,217,755 · App. 17/640,221 · Granted Feb 4, 2025

Voice conversion apparatus, voice conversion learning apparatus, image generation apparatus, image generation learning apparatus, voice conversion method, voice conversion learning method, image generation method, image generation learning method, and computer program

Inventors: Hirokazu Kameoka (Musashino, JP); Ko Tanaka (Musashino, JP); Yasunori Oishi (Musashino, JP); Takuhiro Kaneko (Musashino, JP); Aaron Valero Puche (Musashino, JP)
Assignee: Nippon Telegraph and Telephone Corporation
G10L15/25G06V40/176G10L15/02G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,755
App. No.
17/640,221
Granted
Feb 4, 2025
Kind
B2
Abstract

A voice conversion device is provided with a linguistic information extraction unit that extracts linguistic information corresponding to utterance content from a conversion source voice signal, an appearance feature extraction unit that extracts appearance features expressing features related to the look of a person's face from a captured image of the person, and a converted voice generation unit that generates a converted voice on a basis of the linguistic information and the appearance features.

Claims (37)

1. A voice device comprising one or more processors configured to perform operations comprising:

receiving a voice signal;

extracting linguistic information corresponding to utterance content from the voice signal;

receiving a captured image of a person;

extracting appearance features expressing features related to the look of the person's face from the captured image;

determining a target timbre conforming to the appearance features of the person in the captured image; and

generating a converted voice based on the linguistic information and the appearance features, wherein the converted voice is in the target timbre that conforms to the appearance features of the person in the captured image.

2. The voice device according to claim 1 , wherein the operations comprise:

accepting linguistic information corresponding to the utterance content extracted from the voice signal and appearance features expressing features related to the look of the person's face extracted from the captured image of the person as input, and train parameters of a linguistic information extraction unit configured to extract the linguistic information, an appearance feature extraction unit configured to extract the appearance features, and a converted voice generation unit configured to generate the converted voice on a basis of the linguistic information and the appearance features.

3. The voice device according to claim 2 , wherein the operations comprise:

performing learning such that the converted voice obtained when the linguistic information and the appearance features are input is as close as possible to the voice signal from which the linguistic information is extracted.

4. A non-transitory computer readable medium storing one or more instructions causing a computer to function as the voice device according to claim 1 .

5. An image device comprising:

a timbre feature extraction unit, including one or more processors, configured to extract timbre features expressing features related to vocal timbre from a voice signal; and

an image generation unit, including one or more processors, configured to generate a face image on a basis of the timbre features and appearance features expressing features related to the look of a person's face obtained from a captured image of the person, wherein the generated face image conforms to the timbre features of the voice signal.

6. The image device according to claim 5 , comprising:

a learning unit, including one or more processors, configured to accept timbre features expressing features related to vocal timbre extracted from the voice signal, appearance features expressing features related to the look of a person's face obtained from the captured image of the person, and a captured image of a person as input, and train parameters of the timbre feature extraction unit configured to extract the timbre features.

7. The image device according to claim 6 , wherein

the learning unit is configured to perform learning on a basis of the appearance features and the captured image such that the face image generated when the appearance features from any given captured image are input is as close as possible to the captured image from which the appearance features are extracted.

8. The image device according to claim 6 , wherein

the learning unit is configured to perform learning on a basis of the timbre features and the captured image such that the timbre features are as close as possible to appearance features obtained from the captured image used to generate a face image.

9. A method comprising:

receiving a voice signal;

extracting linguistic information corresponding to utterance content from the voice signal;

receiving a captured image of a person;

extracting appearance features expressing features related to the look of the person's face from the captured image of the person;

determining a target timbre conforming to the appearance features of the person in the captured image; and

generating a converted voice based on the linguistic information and the appearance features, wherein the converted voice is in the target timbre that conforms to the appearance features of the person in the captured image.

10. The method according to claim 9 , comprising:

performing learning that accepts the linguistic information corresponding to utterance content extracted from the voice signal and appearance features expressing features related to the look of a person's face extracted from the captured image of the person as input, and training parameters of a linguistic information extraction unit that extracts the linguistic information, an appearance feature extraction unit that extracts the appearance features, and a converted voice generation unit that generates a converted voice on a basis of the linguistic information and the appearance features.

11. The method according to claim 9 , comprising:

extracting timbre features expressing features related to vocal timbre from a voice signal; and

generating a face image on a basis of the timbre features and appearance features expressing features related to the look of the person's face obtained from the captured image of the person.

12. The method according to claim 9 , comprising:

performing learning that accepts timbre features expressing features related to vocal timbre extracted from a voice signal, appearance features expressing features related to the look of a person's face obtained from the captured image of the person, and the captured image of a person as input, and training parameters of an image generation unit that generates a face image on a basis of the timbre features and the appearance features and a timbre feature extraction unit that extracts the timbre features.

13. The method according to claim 10 , comprising:

performing learning such that the converted voice obtained when the linguistic information and the appearance features are input is as close as possible to the voice signal from which the linguistic information is extracted.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2025
From: KAMEOKA, HIROKAZU; TANAKA, KO; OISHI, YASUNORI; KANEKO, TAKUHIRO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 069733/0710 →
Priority Claims (1)
JP 2019-163418 · Sep 6, 2019 · national
Continuity (1)
Related Publication 20220335944A1 · Oct 20, 2022
References Cited (7)
US 20160180155A1 · Zhang · 2016 [cited by examiner]
US 20170270948A1 · Li · 2017 [cited by examiner]
US 20190213400A1 · Kim · 2019 [cited by examiner]
Fang et al., “Audiovisual Speaker Conversion: Jointly and Simultaneously Transforming Facial Expression and Acoustic Characteristics,” 2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASS… [cited by applicant]
Kameoka et al., “ACVAE-VC: Non-Parallel Voice Conversion With Auxiliary Classifier Variational Autoencoder,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, 2019, 27(9):1432-1443. [cited by applicant]
Kameoka et al., “Crossmodal Voice Conversion,” Submitted to Interspeech 2019, dated Apr. 9, 2019, retrieved from URL <https://arxiv.org/abs/1904.04540v1>, 6 pages. [cited by applicant]
Kameoka, “Face-to-voice conversion and voice-to-face conversion,” NTT Communication Science Laboratories Open House 2019, May 30, 2019, 5 pages (with English Translation). [cited by applicant]