IP Library Granted Patent US 10,748,544
Granted Patent B2
US 10,748,544 · App. 15/934,372 · Granted Aug 18, 2020

Voice processing device, voice processing method, and program

Inventors: Kazuhiro Nakadai (Wako, JP); Tomoyuki Sahata (Tokyo, JP)
Assignee: HONDA MOTOR CO., LTD.
G10L17/20G01S3/8006G10L17/00G10L21/028G10L21/0232G06K9/00228G10L15/26G10L21/0272G10L25/78
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,748,544
App. No.
15/934,372
Granted
Aug 18, 2020
Kind
B2
Abstract

A voice processing device includes: a sound source localization unit configured to determine a direction of each sound source on the basis of voice signals of a plurality of channels; a sound source separation unit configured to separate signals for respective sound sources indicating components of respective sound sources from the voice signals of the plurality of channels; a speech section detection unit configured to detect a speech section in which the number of speakers is 1 from the signals for respective sound sources; and a speaker identification unit configured to identify a speaker on the basis of the signals for respective sound sources in the speech section.

Claims (24)

1. A voice processing device, comprising:

a sound source localization unit configured to determine a direction of each sound source on the basis of voice signals of a plurality of channels;

a sound source separation unit configured to separate signals for respective sound sources indicating components of respective sound sources from the voice signals of the plurality of channels;

a speech section detection unit configured to detect speech sections from the signals for respective sound sources and to determine a speech section in which a number of speakers is 1 among the speech sections as a single speech section; and

a speaker identification unit configured to identify a speaker on the basis of the signals for respective sound sources in the single speech section.

2. The voice processing device according to claim 1 , wherein the speech section detection unit detects the single speech section from sections in which a number of sound sources, of which directions are determined by the sound source localization unit, is 1.

3. The voice processing device according to claim 1 , wherein the speaker identification unit estimates speakers of the speech sections, in which directions of sound sources determined by the sound source localization unit are within a predetermined range from a direction of a sound source identified in the single speech section, to be identical to the speaker of the single speech section.

4. The voice processing device according to claim 1 , comprising an image processing unit configured to determine a direction of a speaker on the basis of a captured image,

wherein the speaker identification unit selects sound sources, for which the direction of the speaker determined by the image processing unit is within a predetermined range, from a direction of each sound source determined by the sound source localization unit and detects the single speech section from sections in which a number of selected sound sources is 1.

5. The voice processing device according to claim 1 , comprising a voice recognition unit configured to perform a voice recognition process on the signals for respective sound sources,

wherein the voice recognition unit provides speech information indicating contents of speech to each speaker determined by the speaker identification unit.

6. The voice processing device according to claim 1 ,

wherein the speaker identification unit estimates speakers of the speech sections, in which directions of sound sources determined by the sound source localization unit are within a predetermined range from a direction of a sound source identified in the single speech section and which are out of the single speech section, to be identical to the speaker of the single speech section.

7. The voice processing device according to claim 1 , wherein the speaker identification unit calculates a likelihood for each registered speaker with respect to sound feature quantities of the signals for respective sound sources in the single speech section with reference to speaker identification data of registered speakers stored in advance in a speaker identification data storage unit and identifies a speaker on the basis of the calculated likelihood.

8. A voice processing method in a voice processing device, comprising:

a sound source localization step of determining a direction of each sound source on the basis of voice signals of a plurality of channels;

a sound source separation step of separating signals for respective sound sources indicating components of respective sound sources from the voice signals of the plurality of channels;

a speech section detection step of detecting speech sections from the signals for respective sound sources and determining a speech section in which a number of speakers is 1 among the speech sections as a single speech section; and

a speaker identification step of identifying a speaker on the basis of the signals for respective sound sources in the single speech section.

9. A non-transitory computer-readable storage medium storing a program for causing a computer of a voice processing to execute:

a sound source localization process of determining a direction of each sound source on the basis of voice signals of a plurality of channels;

a sound source separation process of separating signals for respective sound sources indicating components of respective sound sources from the voice signals of the plurality of channels;

a speech section detection process of detecting speech sections from the signals for respective sound sources and determining a speech section in which a number of speakers is 1 among the speech sections as a single speech section; and

a speaker identification process of identifying a speaker on the basis of the signals for respective sound sources in the single speech section.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 23, 2018
From: NAKADAI, KAZUHIRO; SAHATA, TOMOYUKI
To: HONDA MOTOR CO., LTD.
Reel/Frame 045336/0377 →
Priority Claims (1)
JP 2017-065932 · Mar 29, 2017 · national
Continuity (1)
Related Publication 20180286411A1 · Oct 4, 2018
Cited By (2)
US 12,395,809 US 12,445,793