IP Library Granted Patent US 10,475,439
Granted Patent B2
US 10,475,439 · App. 15/536,299 · Granted Nov 12, 2019

Information processing system and information processing method

Inventors: Shinichi Kawano (Tokyo, JP); Yuhei Taki (Kanagawa, JP)
Assignee: SONY CORPORATION
G10L15/01G06F3/011G06F3/013G06F3/017G06F3/0425G06F3/0488G06F3/04817G06F3/16G06F3/167G06T7/20G10L15/22G10L15/28
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,475,439
App. No.
15/536,299
Granted
Nov 12, 2019
Kind
B2
Abstract

There is provided an information processing system enabling a user to provide easily an instruction on whether to continue speech recognition processing on sound information, the information processing system including: a recognition control portion configured to control a speech recognition portion so that the speech recognition portion performs speech recognition processing on sound information input from a sound collection portion. The recognition control portion controls whether to continue the speech recognition processing on the basis of a gesture of a user detected at predetermined timing.

Claims (41)

1. An information processing system comprising:

a recognition control portion configured to

control a speech recognition portion so that the speech recognition portion performs speech recognition processing on sound information input from a sound collection portion, and

detect a duration, in which a volume of the sound information continuously falls below a reference volume, reaching a predetermined target time after the speech recognition processing has started; and

an output control portion configured to cause an output portion to output a motion object that is displayed by a display after the detected duration,

wherein the recognition control portion controls whether to continue the speech recognition processing, based on a degree of coincidence between a trajectory of a viewpoint of a user detected after the output portion outputs the motion object and a trajectory of the motion object displayed by the display,

wherein the recognition control portion and the output control portion are each implemented via at least one processor.

2. The information processing system according to claim 1 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion continues the speech recognition processing in a case where the degree of coincidence exceeds a threshold.

3. The information processing system according to claim 2 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion executes a predetermined execution operation based on a result of the speech recognition processing in a case where the degree of coincidence falls below the threshold.

4. The information processing system according to claim 3 ,

wherein the predetermined execution operation includes at least one of an operation of outputting a search result corresponding to a result of the speech recognition processing, an operation of outputting the result of the speech recognition processing, an operation of outputting a processing result candidate obtained during the speech recognition processing, and an operation of outputting a string used to reply to utterance contents extracted from the result of the speech recognition processing.

5. The information processing system according to claim 1 ,

wherein the output control portion causes the output portion to output a predetermined first notification object in a case where the degree of coincidence exceeds a threshold.

6. The information processing system according to claim 5 ,

wherein the output control portion causes the output portion to output a predetermined second notification object different from the predetermined first notification object in a case where the degree of coincidence falls below the threshold.

7. The information processing system according to claim 1 ,

wherein the recognition control portion controls whether to continue the speech recognition processing, based on a tilt of a head of the user.

8. The information processing system according to claim 7 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion continues the speech recognition processing in a case where the tilt of the head of the user exceeds a predetermined reference value.

9. The information processing system according to claim 8 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion executes a predetermined execution operation based on a result of the speech recognition processing in a case where the tilt of the head of the user falls below the predetermined reference value.

10. The information processing system according to claim 1 ,

wherein the recognition control portion controls whether to continue the speech recognition processing, based on motion of a head of the user.

11. The information processing system according to claim 10 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion continues the speech recognition processing in a case where the motion of the head of the user indicates a predetermined motion.

12. The information processing system according to claim 11 ,

wherein the recognition control portion controls the speech recognition portion so that the speech recognition portion executes a predetermined execution operation based on a result of the speech recognition processing in a case where the motion of the head of the user fails to indicate the predetermined motion.

13. The information processing system according to claim 1 ,

wherein the recognition control portion causes the speech recognition portion to start the speech recognition processing in a case where an activation trigger of the speech recognition processing is detected.

14. An information processing method comprising:

performing speech recognition processing on sound information input from a sound collection portion;

detecting a duration, in which a volume of the sound information continuously falls below a reference volume, reaching a predetermined target time after the speech recognition processing has started;

outputting a motion object that is displayed by a display after the detected duration; and

controlling, by a processor, whether to continue the speech recognition processing, based on a degree of coincidence between a trajectory of a viewpoint of a user detected after the outputting of the motion object and a trajectory of the motion object displayed by the display.

15. A non-transitory computer-readable medium having embodied thereon a program, which when executed by a computer causes the computer to execute a method, the method comprising:

performing speech recognition processing on sound information input from a sound collection portion;

detecting a duration, in which a volume of the sound information continuously falls below a reference volume, reaching a predetermined target time after the speech recognition processing has started;

outputting a motion object that is displayed by a display after the detected duration; and

controlling, by a processor, whether to continue the speech recognition processing, based on a degree of coincidence between a trajectory of a viewpoint of a user detected after the outputting of the motion object and a trajectory of the motion object displayed by the display.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2017
From: KAWANO, SHINICHI; TAKI, YUHEI
To: SONY CORPORATION
Reel/Frame 042822/0141 →
Priority Claims (1)
JP 2015-059567 · Mar 23, 2015 · national
Continuity (1)
Related Publication 20170330555A1 · Nov 16, 2017