IP Library › Granted Patent US 12,548,573
Granted Patent B2
US 12,548,573 · App. 18/274,296 · Granted Feb 10, 2026

Speaker embedding device, speaker embedding method, and speaker embedding program

Inventors: Yusuke Ijima (Tokyo, JP); Kenichi Fujita (Tokyo, JP); Atsushi Ando (Tokyo, JP)
Assignee: NTT, Inc.
G10L17/04G10L17/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,573
App. No.
18/274,296
Granted
Feb 10, 2026
Kind
B2
Abstract

A speaker embedding apparatus includes processing circuitry configured to accept input of voice data, generate utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data, and use a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and train a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input.

Claims (37)

1 . A speaker embedding apparatus comprising:

processing circuitry configured to:

accept input of voice data;

generate utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and

use a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and train a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,

the processing circuitry is further configured to:

accept input of voice data to be converted into a speaker vector,

generate the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and

input a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.

2 . The speaker embedding apparatus according to claim 1 , wherein the processing circuitry is further configured to:

use utterances and a duration length for each of the utterances indicated in the utterance unit segmentation information as training data, and when utterances of a speaker and a duration length for each of the utterances are input, train a speaker identification model for outputting an identification result of the speaker.

3 . The speaker embedding apparatus according to claim 2 , wherein the processing circuitry is further configured to:

accept input of voice data to be converted into a speaker vector,

generate the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and

input utterances and a duration length for each of the utterances indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.

4 . The speaker embedding apparatus according to claim 1 , wherein the processing circuitry is further configured to:

convert a duration length for each utterance indicated in the generated utterance unit segmentation information into a one-dimensional numerical expression, and

train the speaker identification model using the converted one-dimensional numerical expression of the duration length for each utterance as training data.

5 . The speaker embedding apparatus according to claim 4 , wherein the processing circuitry is further configured to:

convert utterances and the duration length for each of the utterances indicated in the generated utterance unit segmentation information into a one-dimensional numerical expression, and

train the speaker identification model using the converted one-dimensional numerical expression of the utterances and the duration length for each of the utterances as training data.

6 . A speaker embedding method executed by a speaker embedding apparatus, comprising:

accepting input of voice data;

generating utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and

using a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and training a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,

the speaker embedding method is further configured to:

accepting input of voice data to be converted into a speaker vector,

generating the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and

inputting a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.

7 . A non-transitory computer-readable recording medium storing therein a speaker embedding program that causes a computer to execute a process comprising:

accepting input of voice data;

generating utterance unit segmentation information indicating a duration length for each utterance of a speaker in the input voice data; and

using a duration length for each utterance indicated in the generated utterance unit segmentation information as training data and training a speaker identification model for outputting an identification result of a speaker when a duration length for each utterance of the speaker is input,

the process is further configured to:

accepting input of voice data to be converted into a speaker vector,

generating the utterance unit segmentation information of the input voice data to be converted into a speaker vector, and

inputting a duration length for each utterance indicated in the generated utterance unit segmentation information to the speaker identification model after training and output, as a speaker vector of the voice data, an output in an intermediate layer of the speaker identification model.

Assignments (2)
CHANGE OF NAME Recorded Oct 3, 2025
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 073007/0308 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2023
From: IJIMA, YUSUKE; FUJITA, KENICHI; ANDO, ATSUSHI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 064395/0303 →
Continuity (1)
Related Publication 20240312465A1 · Sep 19, 2024
References Cited (8)
US 11417343B2 · Cohen · 2022 [cited by examiner]
US 20030144841A1 · Shao · 2003 [cited by examiner]
US 20230055597A1 · Niiro · 2023 [cited by examiner]
Takaki et al. (2015) “Deep Auto-encoder based Low-dimensional Feature Extraction using FFT Spectral Envelopes in Statistical Parametric Speech Synthesis” IEICE Technical Report, vol. 115, No. 346, pp. 99-104, ISSN 0913-… [cited by applicant]
Sone et al. (2017) “Pre-training Method for DNN-based Speech Recognition and Synthesis Based on Bidirectional Conversion between Text and Speech” IPSJ SIG Technical Report, vol. 2017-MUS-115, No. 40, pp. 1-6, ISSN 2188-… [cited by applicant]
Snyder et al. (2018) “X-Vectors: Robust DNN Embeddings for Speaker Recognition” 2018 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP). [cited by applicant]
Otani et al. (2020) “Multi-language multi-speaker modeling for speech synthesis using generative adversarial networks” Lecture proceedings of the Acoustical Society of Japan, ISSN 1880-7658, pp. 695-696. [cited by applicant]
Japanese Patent Application No. 2022-579192, Office Action mailed Jun. 25, 2024, 3 pages. [cited by applicant]