IP Library Granted Patent US 11,620,997
Granted Patent B2
US 11,620,997 · App. 16/982,201 · Granted Apr 4, 2023

Information processing device and information processing method

Inventors: Hiromi Kurasawa (Tokyo, JP); Kazumi Aoyama (Saitama, JP); Yasuharu Asano (Kanagawa, JP)
Assignee: SONY CORPORATION
G10L15/22G10L15/07G10L15/083G10L2015/088G10L2015/223G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,997
App. No.
16/982,201
Granted
Apr 4, 2023
Kind
B2
Abstract

Provided is an information processing device that includes a determination unit that determines whether an object that outputs voice is a dialogue target related to voice dialogue based on a result of recognition of an input image, and a dialogue function unit that performs control related to the voice dialogue based on the determination. The dialogue function unit provides a voice dialogue function to the object based on the determination that the object being the dialogue target. Further provided is a method that includes determining whether an object that outputs voice is a dialogue target related to voice dialogue based on a result of recognition of an input image, and performing control related to the voice dialogue based on a result of the determining. The performing of the control further includes providing a voice dialogue function to the object based on the determination that the object is the dialogue target.

Claims (46)

1. An information processing device, comprising:

a determination unit configured to receive a plurality of images continuously in a temporally sequential manner, wherein the determination unit includes:

a voice output object detection unit configured to detect, based on an image of the plurality of images, an object region related to an object that outputs voice;

a moving body region determination unit configured to determine a moving body region based on the plurality of images continuously received in the temporally sequential manner; and

a dialogue target region determination unit configured to:

determine whether the object that outputs the voice is a dialogue target related to voice dialogue, wherein the determination of the dialogue target is based on the image; and

specify a dialogue target region related to the dialogue target in the object region,

wherein the dialogue target region is specified based on the moving body region, the object region, and the determination that the object is the dialogue target; and

a dialogue function unit configured to:

execute control related to the voice dialogue based on the dialogue target region and the determination of whether the object is the dialogue target; and

output a voice dialogue function to the object in a case where the object is determined as the dialogue target.

2. The information processing device according to claim 1 , wherein the dialogue function unit does not perform active voice output to the object in a case where the object is different from the dialogue target.

3. The information processing device according to claim 1 , wherein the dialogue function unit does not respond to the voice output from the object in a case where the object is different from the dialogue target.

4. The information processing device according to claim 1 , wherein

the determination unit is further configured to determine whether the object is a reception target or a notification target related to the voice dialogue,

in a case where the object is determined as the reception target, the dialogue function unit is further configured to receive the voice output from the object, and

in a case where the object is determined as the notification target, the dialogue function unit is further configured to execute active voice output to the object.

5. The information processing device according to claim 1 , wherein the determination unit is further configured to determine that a person in a physical space of the information processing device is the dialogue target.

6. The information processing device according to claim 1 , wherein the determination unit is further configured to determine that a determined device specified in advance is the dialogue target.

7. The information processing device according to claim 6 , wherein, in a case where a voice output from the determined device is registered user voice, the determination unit is further configured to determine that the determined device is the dialogue target.

8. The information processing device according to claim 1 , wherein the dialogue target region determination unit is further configured to:

determine whether a moving body related to the moving body region is identical to the object related to the object region; and

specify the dialogue target region based on a result of the determination of whether the moving body is identical to the object.

9. The information processing device according to claim 8 , wherein the dialogue target region determination unit is further configured to:

determine that the moving body is identical to the object based on the moving body region that partially overlaps the object region; and

specify the moving body region as the dialogue target region based on the determination that the moving body is identical to the object.

10. The information processing device according to claim 9 , wherein, in a case where the moving body related to the moving body region is determined as the dialogue target in past, the dialogue target region determination unit is further configured to specify specifics the moving body region as the dialogue target region.

11. The information processing device according to claim 8 , wherein, in case where the dialogue target region determination unit determines that the moving body is identical to the object, the dialogue target region determination unit is further configured to specify specifics the moving body region as the dialogue target in a case where the object is the dialogue target.

12. The information processing device according to claim 8 , wherein, in a case where the dialogue target region determination unit determines that the moving body is identical to the object, the dialogue target region determination unit is further configured to specify the moving body region as the dialogue target based on the object that is a person.

13. The information processing device according to claim 1 , wherein the dialogue target region determination unit is further configured to specify the dialogue target region based on an operation range of the moving body region.

14. The information processing device according to claim 13 , wherein

the operation range exceeds a threshold, and

the dialogue target region determination unit is further configured to specify the moving body region as the dialogue target region based on the operation range that exceeds the threshold.

15. The information processing device according to claim 13 , wherein, in a case where the operation range does not exceed a region corresponding to the object, the dialogue target region determination unit is further configured to determine that the moving body region is different from the dialogue target region.

16. The information processing device according to claim 1 , wherein

the dialogue target region determination unit is further configured to exclude, from the dialogue target region, the moving body region corresponding to a following object, and

the following object follows the object determined as the dialogue target.

17. The information processing device according to claim 1 , wherein the determination unit is further configured to reset a determination result related to the dialogue target based on a change in disposition of the information processing device.

18. An information processing method comprising:

receiving a plurality of images continuously in a temporally sequential manner;

detecting, based on an image of the plurality of images, an object region related to an object that outputs voice;

determining a moving body region based on the plurality of images continuously received in the temporally sequential manner;

determining whether the object that outputs the voice is a dialogue target related to voice dialogue, wherein the determination of the dialogue target is based on the image;

specifying a dialogue target region related to the dialogue target in the object region, wherein the dialogue target region is specified based on the moving body region, the object region, and the determination that the object is the dialogue target;

executing control related to the voice dialogue based on the dialogue target region and the determination of whether the object is the dialogue target; and

outputting a voice dialogue function to the object in a case where the object is determined as the dialogue target.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 18, 2020
From: KURASAWA, HIROMI; AOYAMA, KAZUMI; ASANO, YASUHARU
To: SONY CORPORATION
Reel/Frame 053816/0448 →
Priority Claims (1)
JP JP2018-059203 · Mar 27, 2018 · national
Continuity (1)
Related Publication 20210027779A1 · Jan 28, 2021
Cited By (1)
US 12,373,027