IP Library Granted Patent US 11,217,266
Granted Patent B2
US 11,217,266 · App. 16/089,174 · Granted Jan 4, 2022

Information processing device and information processing method

Inventors: Shinichi Kawano (Tokyo, JP); Yuhei Taki (Kanagawa, JP); Yusuke Nakagawa (Kanagawa, JP); Ayumi Kato (Kanagawa, JP)
Assignee: SONY CORPORATION
G10L25/51G10L15/04G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,217,266
App. No.
16/089,174
Granted
Jan 4, 2022
Kind
B2
Abstract

There is provided an information processing device to achieve more flexible correction of a recognized sentence, the information processing device including: a comparison unit configured to compare first sound-related information obtained from collected first utterance information with second sound-related information obtained from collected second utterance information; and a setting unit configured to set a new delimiter position different from a result of speech-to-text conversion associated with the first utterance information on a basis of a comparison result obtained by the comparison unit. There is also provided an information processing device including: a reception unit configured to receive information regarding a new delimiter position different from a result of speech-to-text conversion associated with collected first utterance information; and an output control unit configured to control output of a new conversion result obtained by performing speech-to-text conversion on a basis of the new delimiter position.

Claims (60)

1. An information processing device comprising:

at least one processor configured to

compare first sound-related information obtained from collected first utterance information with second sound-related information obtained from collected second utterance information,

set a new delimiter position different from a result of speech-to-text conversion associated with the first utterance information based on a comparison result indicating that the first utterance information is similar to the second utterance information, and

perform speech-to-text conversion based on the new delimiter position,

wherein the new delimiter position is set from initially delimiting one or more grouped phrases to delimiting one or more words grouped within each phrase of the initially delimited one or more grouped phrases, and

wherein the new delimiter position is set based on a confidence level determined for each variation of the new delimiter position.

2. The information processing device according to claim 1 ,

wherein the at least one processor performs speech-to-text conversion associated with the second utterance information on the basis of the new delimiter position.

3. The information processing device according to claim 1 ,

wherein the at least one processor performs the speech-to-text conversion associated with the first utterance information on the basis of the new delimiter position.

4. The information processing device according to claim 1 ,

wherein the at least one processor is further configured to receive the first utterance information and the second utterance information.

5. The information processing device according to claim 4 ,

wherein the at least one processor is further configured to

receive target information used to specify the first utterance information, and

compare the first sound-related information with the second sound-related information based on the target information.

6. The information processing device according to claim 1 ,

wherein the at least one processor is further configured to initiate transmission of information regarding the new delimiter position.

7. The information processing device according to claim 6 ,

wherein the at least one processor is further configured to transmit a result of the speech-to-text conversion based on the new delimiter position.

8. The information processing device according to claim 1 ,

wherein the at least one processor is further configured to perform speech recognition based on the first utterance information or the second utterance information.

9. An information processing device comprising:

at least one processor configured to

receive information regarding a new delimiter position different from a result of speech-to-text conversion associated with collected first utterance information,

control output of a new conversion result obtained by performing speech-to-text conversion based on the new delimiter position, and

initiate output of the new conversion result and the new delimiter position in association with each other,

wherein the new delimiter position is set based on a comparison result obtained by comparing first sound-related information obtained from the collected first utterance information with second sound-related information obtained from collected second utterance information, the comparison result indicating that the first sound-related information is similar to the second sound-related information,

wherein the new delimiter position is set from initially delimiting one or more grouped phrases to delimiting one or more words grouped within each phrase of the initially delimited one or more grouped phrases, and

wherein the new delimiter position is set based on a confidence level determined for each variation of the new delimiter position.

10. The information processing device according to claim 9 ,

wherein the at least one processor is further configured to transmit the first utterance information and the second utterance information.

11. The information processing device according to claim 10 ,

wherein the at least one processor is further configured to transmit target information used to specify the first utterance information, and

wherein the at least one processor receives the information regarding the new delimiter position set based on the target information.

12. The information processing device according to claim 11 ,

wherein the at least one processor is further configured to detect an input operation by a user and generate the target information based on the input operation.

13. The information processing device according to claim 9 ,

wherein the at least one processor is further configured to receive the new conversion result.

14. The information processing device according to claim 9 ,

wherein the at least one processor is further configured to perform the speech-to-text conversion on the basis of the new delimiter position.

15. The information processing device according to claim 9 , further comprising:

an output configured to output the new conversion result based on the initiated output.

16. The information processing device according to claim 9 ,

wherein the at least one processor is further configured to collect the first utterance information and the second utterance information,

wherein the second utterance information is acquired after acquisition of the first utterance information.

17. An information processing method comprising:

comparing, by a processor, first sound-related information obtained from collected first utterance information with second sound-related information obtained from collected second utterance information;

setting a new delimiter position different from a result of speech-to-text conversion associated with the first utterance information based on a comparison result indicating that the first sound-related information is similar to the second sound-related information; and

perform speech-to-text conversion based on the new delimiter position,

wherein the new delimiter position is set from initially delimiting one or more grouped phrases to delimiting one or more words grouped within each phrase of the initially delimited one or more grouped phrases, and

wherein the new delimiter position is set based on a confidence level determined for each variation of the new delimiter position.

18. An information processing method comprising:

receiving, by a processor, information regarding a new delimiter position different from a result of speech-to-text conversion associated with collected first utterance information;

controlling output of a new conversion result obtained by performing speech-to-text conversion based on the new delimiter position; and

outputting the new conversion result and the new delimiter position in association with each other,

wherein the new delimiter position is set based on a result obtained by comparing first sound-related information obtained from the collected first utterance information with second sound-related information obtained from collected second utterance information,

wherein the new delimiter position is set from initially delimiting one or more grouped phrases to delimiting one or more words grouped within each phrase of the initially delimited one or more grouped phrases, and

wherein the new delimiter position is set based on a confidence level determined for each variation of the new delimiter position.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2018
From: KAWANO, SHINICHI; TAKI, YUHEI; NAKAGAWA, YUSUKE; KATO, AYUMI
To: SONY CORPORATION
Reel/Frame 046997/0142 →
Priority Claims (1)
JP JP2016-122437 · Jun 21, 2016 · national
Continuity (1)
Related Publication 20200302950A1 · Sep 24, 2020