IP Library Granted Patent US 9,530,431
Granted Patent B2
US 9,530,431 · App. 14/193,099 · Granted Dec 27, 2016

Device method, and computer program product for calculating score representing correctness of voice

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,530,431
App. No.
14/193,099
Granted
Dec 27, 2016
Kind
B2
Abstract

According to an embodiment, a voice processor includes a presenting unit to present text to an operator; a voice acquisition unit to acquire a voice of the operator reading aloud the text; an identifying unit to identify output intervals of phonemes included in the voice; a determination unit to determine whether each of time lengths of the output intervals is normal; a frequency acquisition unit to acquire frequency values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, the context including the phoneme and another phoneme adjacent to at least one side of the phoneme; and a score calculator to calculate a score representing correctness of the voice on the basis of the determination results of the time lengths of the output intervals and the frequency values of the contexts acquired respectively corresponding to the phonemes.

Claims (45)

1. A voice processing device, comprising:

a processor that operates as:

a presenting unit that presents text to an operator;

a voice acquisition unit that acquires a voice of the operator reading aloud the text;

an identifying unit that identifies output intervals of phonemes included in the voice of the operator;

a determination unit that determines whether each of time lengths of the output intervals is normal;

a frequency acquisition unit that acquires frequency values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, each context including the phoneme and another phoneme adjacent to at least one side of the phoneme;

a weight calculator that calculates a weight corresponding to each of the phonemes in accordance with a frequency value of the context; and

a score calculator that calculates, as a score representing correctness of the voice of the operator, a value in accordance with a ratio of a first value to a second value, the first value representing a sum of weights corresponding to the phonemes, the second value representing a sum of weights corresponding to the phonemes having the time lengths of the output intervals that are determined as normal.

2. The voice processing device according to claim 1 , wherein the weight calculator calculates the weight such that a weight corresponding to a first phoneme is larger than a weight corresponding to a second phoneme, the first phoneme having a value that is equal to or larger than a value for which a corresponding frequency value is preset, the second phoneme having a value that is smaller than the value for which the corresponding frequency value is preset.

3. The voice processing device according to claim 1 , wherein the processor further operates as:

a notifier that notifies the operator of content according to the score.

4. The voice processing device according to claim 1 , wherein the processor further operates as:

a frequency storage unit that stores therein occurrence frequencies of a plurality of contexts included in voices acquired in the past, as the frequency values;

an updating unit that updates the frequency values, which are stored in the frequency storage unit, of the contexts corresponding to the phonemes included in the voice of the operator reading aloud the text in accordance with the score; and

a text selector that selects, as the text, one piece of text from among a plurality of pieces of candidate text, wherein

the text selector selects the text on the basis of the frequency values of contexts corresponding to a plurality of phonemes included in the pieces of candidate text when the pieces of candidate text are read aloud.

5. The voice processing device according to claim 4 , wherein the text selector selects the candidate text in preference to the other candidate text, the preferred candidate text including the phonemes for which the contexts have the frequency values larger than a threshold at the head of and the end of the text and the phonemes for which the contexts have the frequency values smaller than the threshold at a part of the text other than the head and the end of the text.

6. A voice processing method, comprising: presenting, by a processor, text to an operator; acquiring, by the processor, a voice of the operator reading aloud the text;

identifying, by the processor, output intervals of phonemes included in the voice of the operator;

determining, by the processor, whether each of time lengths of the output intervals is normal;

acquiring, by the processor, frequency values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, each context including the corresponding phoneme and another phoneme adjacent to at least one side of the phoneme;

calculating, by the processor, a weight corresponding to each of the phonemes in accordance with a frequency value of the context; and

calculating, by the processor, as a score representing correctness of the voice of the operator, a value in accordance with a ratio of a first value to a second value, the first value representing a sum of weights corresponding to the phonemes, the second value representing a sum of weights corresponding to the phonemes having the time lengths of the output intervals that are determined as normal.

7. A computer program product comprising a non-transitory

computer-readable medium containing a voice processing program that causes a computer to function as:

a presenting unit that presents text to an operator;

a voice acquisition unit that acquires a voice of the operator reading aloud the text;

an identifying unit that identifies output intervals of phonemes included in the voice of the operator;

a determination unit that determines whether each of time lengths of the output intervals is normal;

a frequency acquisition unit that acquires frequency values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, each context including the phoneme and another phoneme adjacent to at least one side of the phoneme;

a weight calculator that calculates a weight corresponding to each of the phonemes in accordance with a frequency value of the context; and

a score calculator that calculates, as a score representing correctness of the voice of the operator, a value in accordance with a ratio of a first value to a second value, the first value representing a sum of weights corresponding to the phonemes, the second value representing a sum of weights corresponding to the phonemes having the time lengths of the output intervals that are determined as normal.

8. A voice processing device, comprising:

a processor that operates as:

a presenting unit that presents text to an operator;

a voice acquisition unit that acquires a voice of the operator reading aloud the text;

an identifying unit that identifies output intervals of phonemes included in the voice of the operator;

a determination unit that determines whether each of time lengths of the output intervals is normal;

a frequency acquisition unit that acquires frequency Values respectively representing occurrence frequencies of contexts, respectively corresponding to the phonemes, each context including the phoneme and another phoneme adjacent to at least one side of the phoneme;

a score calculator that calculates a score representing correctness of the voice of the operator on the basis of the frequency values of the contexts acquired respectively corresponding to the phonemes having the time lengths of the output intervals that are determined as normal;

a frequency storage unit that stores therein occurrence frequencies of a plurality of contexts included in voices acquired in the past, as the frequency values;

an updating unit that updates the frequency values, which are stored in the frequency storage unit, of the contexts corresponding to the phonemes included in the voice of the operator reading aloud the text in accordance with the score; and

a text selector that selects, as the text, one piece of text from among a plurality of pieces of candidate text, wherein

the text selector selects the text on the basis of the frequency values of contexts corresponding to a plurality of phonemes included in the pieces of candidate text when the pieces of candidate text are read aloud.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2014
From: NAKATA, KOUTA
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 032321/0195 →