IP Library › Granted Patent US 11,942,074
Granted Patent B2
US 11,942,074 · App. 17/429,737 · Granted Mar 26, 2024

Learning data acquisition apparatus, model learning apparatus, methods and programs for the same

Inventors: Takaaki Fukutomi (Tokyo, JP); Takashi Nakamura (Tokyo, JP); Kiyoaki Matsui (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G10L15/063G06N20/00G10L21/0208G10L25/78G10L25/81G10L25/84G10L25/87G10L2025/783G10L2025/786
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,942,074
App. No.
17/429,737
Granted
Mar 26, 2024
Kind
B2
Abstract

A learning data acquisition device or the like, capable of acquiring learning data by superimposing noise data on clean voice data at an appropriate SN ratio, is provided. The learning data acquisition device includes a voice recognition influence degree calculation unit and a learning data acquisition unit. The voice recognition influence degree calculation unit calculates an influence degree on voice recognition accuracy caused by a change of a signal-to-noise ratio, based on a result of voice recognition on the k th noise superimposed voice data and a result of voice recognition on the k−1 th noise superimposed voice data, where K is an integer of 2 or larger, k=2, 3, . . . , K, and a signal-to-noise ratio of the the k th noise superimposed voice data is smaller than a signal-to-noise ratio of the k−1 th noise superimposed voice data, and obtains a largest signal-to-noise ratio SNR apply among signal-to-noise ratios of the k−1 th noise superimposed voice data when the influence degree meets a given threshold condition. The learning data acquisition unit acquires noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply , as learning data.

Claims (44)

1. A learning data acquisition device comprising a processor configured to execute operations comprising:

determining an influence degree on voice recognition accuracy caused by a change of a signal-to-noise ratio, based on a result of voice recognition on k th noise superimposed voice data and a result of voice recognition on k−1 th noise superimposed voice data, wherein K is an integer of 2 or larger, k=2, 3, . . . , K, and a signal-to-noise ratio of the k th noise superimposed voice data is smaller than a signal-to-noise ratio of the k−1 th noise superimposed voice data;

obtaining a largest signal-to-noise ratio SNR apply among signal-to-noise ratios of the k−1 th noise superimposed voice data when the influence degree meets a given threshold condition; and

acquiring noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply , as learning data.

2. The learning data acquisition device according to claim 1 , wherein the acquiring further comprises superimposing predetermined noise data on clean voice data to have a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply for generating and acquiring the learning data.

3. The learning data acquisition device according to claim 1 , the processor further configured to execute operations comprising:

superimposing predetermined noise data on clean voice data by changing a signal-to-noise ratio of the predetermined noise data in K steps; and

generating K pieces of the noise superimposed voice data,

wherein the acquiring further comprises selecting and acquiring, from the K pieces of the noise superimposed voice data, and the noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply as learning data.

4. The learning data acquisition device according to claim 1 , wherein the learning data distinguishes utterance.

5. A model learning device comprising a processor configured to execute operations comprising:

determining an influence degree on voice recognition accuracy caused by a change of a signal-to-noise ratio, based on a result of voice recognition on k th noise superimposed voice data and a result of voice recognition on k−1 th noise superimposed voice data, wherein K is an integer of 2 or larger, k=2, 3, . . . , K, and a signal-to-noise ratio of the k th noise superimposed voice data is smaller than a signal-to-noise ratio of the k−1 th noise superimposed voice data;

obtaining a largest signal-to-noise ratio SNR apply among signal-to-noise ratios of the k−1 th noise superimposed voice data when the influence degree meets a given threshold condition;

acquiring noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply , as learning data; and

learning a model for detecting voice or non-voice with use of the learning data.

6. The model learning device according to claim 5 ,

wherein, when a learned model does not satisfy a preset convergence condition, repeating processing comprising the determining the influence degree, the acquiring the learning data, and the learning the model, and

wherein the determining the influence degree further comprises using the model learned by the model learner when performing voice recognition.

7. The model learning device according to claim 6 , wherein the acquiring further comprises superimposing predetermined noise data on clean voice data to have a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply for generating and acquiring the learning data.

8. The model learning device according to claim 6 , the processor further configured to execute operations comprising:

superimposing predetermined noise data on clean voice data by changing a signal-to-noise ratio of the predetermined noise data in K steps; and

generating K pieces of the noise superimposed voice data,

wherein the acquiring further comprises selecting and acquiring, from the K pieces of the noise superimposed voice data, noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply as learning data.

9. The model learning device according to claim 5 , wherein the acquiring further comprises superimposing predetermined noise data on clean voice data to have a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply for generating and acquiring the learning data.

10. The model learning device according to claim 5 , the processor further configured to execute operations comprising:

superimposing predetermined noise data on clean voice data by changing a signal-to-noise ratio of the predetermined noise data in K steps; and

generating K pieces of the noise superimposed voice data,

wherein the acquiring further comprising selecting and acquiring, from the K pieces of the noise superimposed voice data, noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply as learning data.

11. A method comprising:

determining an influence degree on voice recognition accuracy caused by a change of a signal-to-noise ratio, based on a result of voice recognition on k th noise superimposed voice data and a result of voice recognition on k−1 th noise superimposed voice data, wherein K is an integer of 2 or larger, k=2, 3, . . . , K, and a signal-to-noise ratio of the k th noise superimposed voice data is smaller than a signal-to-noise ratio of the k−1 th noise superimposed voice data:

obtaining a largest signal-to-noise ratio SNR apply among signal-to-noise ratios of the k−1″ noise superimposed voice data when the influence degree meets a given threshold condition; and

acquiring noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply , as learning data.

12. The method according to claim 11 , the method further comprising:

learning a model for detecting voice or non-voice with use of the learning data.

13. The method according to claim 12 , wherein the acquiring further comprises superimposing predetermined noise data on clean voice data to have a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply for generating and acquiring the learning data.

14. The method according to claim 12 , further comprising:

superimposing predetermined noise data on clean voice data by changing a signal-to-noise ratio of the predetermined noise data in K steps; and

generating K pieces of the noise superimposed voice data,

wherein the acquiring further comprises selecting and acquiring, from the K pieces of the noise superimposed voice data, noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply as learning data.

15. The method according to claim 11 , wherein the acquiring further comprises superimposing predetermined noise data on clean voice data to have a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply for generating and acquiring the learning data.

16. The method according to claim 11 , further comprising:

superimposing predetermined noise data on clean voice data by changing a signal-to-noise ratio of the predetermined noise data in K steps; and

generating K pieces of the noise superimposed voice data,

wherein the acquiring further comprises selecting and acquiring, from the K pieces of the noise superimposed voice data, noise superimposed voice data having a signal-to-noise ratio that is equal to or larger than the signal-to-noise ratio SNR apply as learning data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 10, 2021
From: FUKUTOMI, TAKAAKI; NAKAMURA, TAKASHI; MATSUI, KIYOAKI
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 057133/0349 →
Priority Claims (1)
JP 2019-022516 · Feb 12, 2019 · national
Continuity (1)
Related Publication 20220101828A1 · Mar 31, 2022