IP Library Granted Patent US 8,275,614
Granted Patent B2
US 8,275,614 · App. 12/428,907 · Granted Sep 25, 2012

Support device, program and support method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,275,614
App. No.
12/428,907
Granted
Sep 25, 2012
Kind
B2
Abstract

A support device, program and support method for supporting generation of text from speech data. The support device includes a confirmed rate calculator, a candidate obtaining unit and a selector. The confirmed rate calculator calculates a confirmed utterance rate which is an utterance rate of a confirmed part having already-confirmed text in the speech data. The candidate obtaining unit obtains multiple candidate character strings resulting from a speech recognition of an unconfirmed part having unconfirmed text in the speech data. The selector preferentially selects, from among the plurality of candidate character strings, a candidate character string whose utterance time consumed in uttering the candidate character string at the confirmed utterance rate is closest to an utterance time of the unconfirmed part of the speech data.

Claims (40)

1. A support device for supporting generation of text from speech data, comprising: a confirmed rate calculator for calculating a confirmed utterance rate which is an utterance rate of a confirmed part having already-confirmed text in the speech data; a candidate obtaining unit for obtaining a plurality of candidate character strings resulting from a speech recognition on an unconfirmed part having unconfirmed text in the speech data; and a selector for selecting one of the plurality of candidate character strings having an utterance time closest to an utterance time of the unconfirmed part in the speech data according to the utterance time consumed to utter the candidate character string at the confirmed utterance rate;

further comprising: a candidate time calculator for calculating, for each of the plurality of candidate character strings, an utterance time consumed to utter the candidate character string at the confirmed utterance rate, on the basis of the confirmed utterance rate and a number of moras or syllables in the candidate character string; wherein the confirmed rate calculator calculates, as the confirmed utterance rate, the number of moras or syllables uttered per unit time in the confirmed part having already-confirmed text in the speech data; and wherein the selector preferentially selects one of the plurality of candidate character strings having the utterance time closest to the utterance time of the unconfirmed part in the speech data according to the utterance time calculated by the candidate time calculator;

wherein the candidate time calculator comprises: a phoneme string generation unit for generating a phoneme string of the candidate character string; a correction factor calculation unit for calculating a correction factor based on a phoneme string of the candidate character string; and an utterance time calculator for calculating, as an utterance time consumed to utter the candidate character string at the confirmed utterance rate, a value obtained by a calculation where the number of moras in the candidate character string is multiplied by the correction factor and then the obtained value is divided by the confirmation utterance rate.

2. The support device according to claim 1 , further comprising:

a top setting unit for changing, when a part of text is confirmed as a confirmed character string, a top position of an unconfirmed part having unconfirmed text in the speech data, from a top position of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a position advanced from the top position by an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

3. The support device according to claim 2 , wherein:

the top position setting unit, when a part of text is confirmed as a confirmed character string, changes a first phoneme of an unconfirmed part having unconfirmed text in the speech data, from a first phoneme of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a phoneme right behind the last phoneme uttered within an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

4. The support device according to claim 3 , wherein: the top position setting unit, when a degree of coincidence between any one of the confirmed character strings and a phoneme string of the confirmed character string and any one of a character strings and a phoneme string of a speech recognition result at the top position of the unconfirmed part of the speech data is higher than a reference degree of coincidence, performs matching between any one of the character strings and the phoneme string of the speech recognition result and any one of the confirmed character strings and the phoneme string of the confirmed character string, and then sets a phoneme right behind the matched last phoneme to be a first phoneme of an unconfirmed part having unconfirmed text in the speech data; and

the top position setting unit, when the degree of coincidence is equal to or lower than the reference degree of coincidence, changes the first phoneme of an unconfirmed part having unconfirmed text in the speech data, from the first phoneme of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a phoneme right behind the last phoneme uttered within an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

5. The support device according to claim 2 , further comprising:

a replacement unit for replacing, in response to an instruction to replace speech of a confirmed part corresponding to the confirmed character string in the speech data, speech data corresponding to the confirmed character string by speech data in which the confirmed character string is read aloud.

6. The support device according to claim 1 , further comprising:

an input unit for receiving, from a user, at least a part of a character string corresponding to the unconfirmed part having unconfirmed text in the speech data;

wherein the candidate obtaining unit obtains the plurality of candidate character strings including a character string inputted by a user, from among a speech recognition result of the unconfirmed part having unconfirmed text in the speech data.

7. The support device according to claim 1 , wherein the selector preferentially selects, from the plurality of candidate character strings, a candidate character string included in a part where text is already confirmed.

8. The support device of claim 1 , wherein:

the confirmed rate calculator calculates a confirmed expression rate which is an expression rate of a confirmed part having already-confirmed text in moving image data;

the candidate obtaining unit obtains a plurality of candidate character strings resulting from an image recognition of an unconfirmed part having unconfirmed text in the moving image data; and

the selector selects, from among the plurality of candidate character strings, a candidate character string having an expression time closest to the expression time of the unconfirmed part of the moving image data, wherein the expression time is the time consumed to express the candidate character string at the confirmed expression rate.

9. A non-transitory computer readable medium storing computer readable instructions for causing a computer to execute the steps of: calculating a confirmed utterance rate which is an utterance rate of a confirmed part having already-confirmed text in the speech data; obtaining a plurality of candidate character strings which are a speech recognition result of an unconfirmed part having unconfirmed text in the speech data; and selecting one of the plurality of candidate character strings having the utterance time closest to the utterance time of the unconfirmed part in the speech data according to the utterance time consumed to utter the candidate character string at the confirmed utterance rate;

wherein: the calculating step calculates a confirmed expression rate which is an expression rate of a confirmed part having already-confirmed text in moving image data; the obtaining step obtains a plurality of candidate character strings resulting from an image recognition of an unconfirmed part having unconfirmed text in the moving image data; and the selecting step selects, from among the plurality of candidate character strings, a candidate character string having an expression time closest to the expression time of the unconfirmed part of the moving image data, wherein the expression time is the time consumed to express the candidate character string at the confirmed expression rate.

10. A support method for supporting generation of text from speech data, comprising the steps of: calculating a confirmed utterance rate, using a processor, which is an utterance rate of a confirmed part having already-confirmed text in the speech data; obtaining a plurality of candidate character strings which are a speech recognition result of an unconfirmed part having unconfirmed text in the speech data; and selecting one of the plurality of candidate character strings having the utterance time closest to the utterance time of the unconfirmed part in the speech data according to the utterance time consumed to utter the candidate character string at the confirmed utterance rate;

further comprising the step of: calculating a candidate utterance rate for each of the plurality of candidate character strings, wherein the candidate utterance rate is an utterance time consumed to utter the candidate character string at the confirmed utterance rate on the basis of the confirmed utterance rate and a number of moras or syllables in the candidate character string; wherein the step of calculating a confirmed utterance rate calculates, as the confirmed utterance rate, the number of moras or syllables uttered per unit time in the confirmed part having already-confirmed text in the speech data; and wherein the selecting step selects one of the plurality of candidate character strings having the utterance time closest to the utterance time of the unconfirmed part in the speech data according to the utterance time calculated by the candidate time calculator;

Wherein the step of calculating a candidate utterance rate comprises the steps of: generating a phoneme string of the candidate character string; calculating a correction factor based on a phoneme string of the candidate character string; and calculating, as an utterance time consumed to utter the candidate character string at the confirmed utterance rate, a value obtained by a calculation where the number of moras in the candidate character string is multiplied by the correction factor and then the obtained value is divided by the confirmation utterance rate.

11. The support method according to claim 10 , further comprising the step of:

changing, when a part of text is confirmed as a confirmed character string, a top position of an unconfirmed part having unconfirmed text in the speech data, from a top position of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a position advanced from the top position by an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

12. The support method according to claim 11 , wherein:

the step of changing, when a part of text is confirmed as a confirmed character string, changes a first phoneme of an unconfirmed part having unconfirmed text in the speech data, from a first phoneme of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a phoneme right behind the last phoneme uttered within an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

13. The support method according to claim 10 , further comprising the steps of: performing, when a degree of coincidence between any one of the confirmed character strings and a phoneme string of the confirmed character string and any one of a character strings and a phoneme string of a speech recognition result at the top position of the unconfirmed part of the speech data is higher than a reference degree of coincidence, matching between any one of the character strings and the phoneme string of the speech recognition result and any one of the confirmed character strings and the phoneme string of the confirmed character string, and then sets a phoneme right behind the matched last phoneme to be a first phoneme of an unconfirmed part having unconfirmed text in the speech data; and

changing, when the degree of coincidence is equal to or lower than the reference degree of coincidence, the first phoneme of an unconfirmed part having unconfirmed text in the speech data, from the first phoneme of an unconfirmed part having unconfirmed text in the speech data before confirmation, to a phoneme right behind the last phoneme uttered within an utterance time consumed to utter the confirmed character string at the confirmed utterance rate.

14. The support method according to claim 11 , further comprising the step of:

replacing, in response to an instruction to replace speech of a confirmed part corresponding to the confirmed character string in the speech data, speech data corresponding to the confirmed character string by speech data in which the confirmed character string is read aloud.

15. The support method according to claim 10 , further comprising the step of:

receiving, from a user, at least a part of a character string corresponding to the unconfirmed part having unconfirmed text in the speech data;

wherein the obtaining step obtains the plurality of candidate character strings including a character string inputted by a user, from among a speech recognition result of the unconfirmed part having unconfirmed text in the speech data.

16. The support method according to claim 10 , wherein the selecting step selects, from the plurality of candidate character strings, a candidate character string included in a part where text is already confirmed.

17. The support method of claim 10 , wherein:

the calculating step calculates a confirmed expression rate which is an expression rate of a confirmed part having already-confirmed text in the moving image data;

the obtaining step obtains a plurality of candidate character strings resulting from an image recognition of an unconfirmed part having unconfirmed text in the moving image data; and

the selecting step selects, from among the plurality of candidate character strings, a candidate character string having an expression time closest to the expression time of the unconfirmed part of the moving image data, wherein the expression time is the time consumed to express the candidate character string at the confirmed expression rate.

Assignments (7)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →