IP Library › Granted Patent US 12,367,874
Granted Patent B2
US 12,367,874 · App. 17/309,993 · Granted Jul 22, 2025

Information processing apparatus and information processing method

Inventors: Katsutoshi Kanamori (Tokyo, JP); Masato Nishio (Tokyo, JP)
Assignee: SONY GROUP CORPORATION
G10L15/22G10L15/06G10L15/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,874
App. No.
17/309,993
Filed
Jul 8, 2021
Granted
Jul 22, 2025
Kind
B2
Art Unit
2657
USPC
704/275
Abstract

Provided is an information processing apparatus which includes a control section that controls a conversation with a user according to a recognized situation. The control section acquires knowledge elements related to the recognized situation in terms of knowledge from knowledge sets, and determines contents of an utterance on the basis of the knowledge elements and an utterance template. Further, provided is an information processing method that includes controlling, by a processor, a conversation with a user according to a recognized situation. The controlling further includes acquiring knowledge elements related to the recognized situation in terms of knowledge from knowledge sets, and determining contents of an utterance on the basis of the knowledge elements and an utterance template.

Claims (63)

1. An information processing apparatus, comprising:

a central processing unit (CPU) configured to:

receive a first utterance of a user;

recognize a situation associated with the user based on the first utterance of the user, wherein the recognized situation includes at least an intention analysis result of the first utterance of the user;

control a conversation with the user based on the recognized situation;

acquire one or more knowledge elements, from a first knowledge set of a plurality of knowledge sets, based on the recognized situation, wherein

the first knowledge set of the plurality of knowledge sets is based on a usage priority associated with the first knowledge set, and

the usage priority is set by the user based on a knowledge domain associated with the first knowledge set;

determine a type of utterance template of a plurality of types of utterance templates based on the intention analysis result of the first utterance of the user;

acquire an utterance template based on the determined type of utterance template;

determine contents of a second utterance based on the one or more knowledge elements and the utterance template; and

output the second utterance to the user.

2. The information processing apparatus according to claim 1 , wherein

the one or more knowledge elements include at least a first vocabulary, and

the first vocabulary corresponds to the knowledge domain of a plurality of knowledge domains.

3. The information processing apparatus according to claim 2 , wherein the first knowledge set of the plurality of knowledge sets contains a first description of the one or more knowledge elements and a second description of a relation between the one or more knowledge elements.

4. The information processing apparatus according to claim 3 , further comprising:

a memory configured to store the plurality of knowledge sets.

5. The information processing apparatus according to claim 1 ,

wherein the CPU is further configured to add one or more knowledge sets to the plurality of knowledge sets, based on an operation by the user.

6. The information processing apparatus according to claim 5 , wherein the CPU is further configured to download the plurality of knowledge sets from a server.

7. The information processing apparatus according to claim 1 ,

wherein the plurality of knowledge sets is based on a user input.

8. The information processing apparatus according to claim 1 ,

wherein the recognized situation includes at least a conversation history of the user, and

the CPU is further configured to acquire, from the first knowledge set, the one or more knowledge elements related to a second vocabulary included in the conversation history.

9. The information processing apparatus according to claim 1 , wherein the CPU is further configured to:

apply the one or more knowledge elements to the acquired utterance template.

10. The information processing apparatus according to claim 1 , wherein the recognized situation further includes at least one of an object recognition result, an environment recognition result, or location information.

11. The information processing apparatus according to claim 1 , wherein

the one or more knowledge elements are used in a specific period for the second utterance by a smaller number of times than a first threshold.

12. The information processing apparatus according to claim 1 , wherein the utterance template is used for a smaller number of times than a second threshold during a specific period.

13. The information processing apparatus according to claim 1 , wherein

the plurality of knowledge sets include a second knowledge set regarding an advertisement, and

the CPU is further configured to:

acquire the one or more knowledge elements from at least one of the second knowledge set or one or more advertisement-related words included in a conversation history of the user.

14. The information processing apparatus according to claim 1 , wherein the CPU is further configured to:

determine first contents of a third utterance based on a third vocabulary included in the first utterance of the user; and

recommend to the user, based on the third utterance, addition of a third knowledge set to the plurality of knowledge sets, wherein the third knowledge set is related to the third vocabulary.

15. The information processing apparatus according to claim 1 , further comprising:

a speaker configured to output a voice that corresponds to the contents of the second utterance.

16. An information processing method, comprising:

receiving a first utterance of a user;

recognizing a situation associated with the user based on the first utterance of the user, wherein the recognized situation includes at least an intention analysis result of the first utterance of the user;

controlling a conversation with the user based on the recognized situation;

acquiring one or more knowledge elements from a first knowledge set of a plurality of knowledge sets, based on the recognized situation, wherein

the first knowledge set of the plurality of knowledge sets is based on a usage priority associated with the first knowledge set, and

the usage priority is set by the user based on a knowledge domain associated with the first knowledge set;

determining a type of utterance template of a plurality of types of utterance templates based on the intention analysis result of the first utterance of the user;

acquiring an utterance template based on the determined type of utterance template;

determining contents of a second utterance based on the one or more knowledge elements and the utterance template; and

outputting the second utterance to the user.

17. A non-transitory computer-readable medium having stored thereon, computer-executable instructions which, when executed by a computer, cause the computer to execute operations, the operations comprising:

receiving a first utterance of a user;

recognizing a situation associated with the user based on the first utterance of the user, wherein the recognized situation includes at least an intention analysis result of the first utterance of the user;

controlling a conversation with the user based on the recognized situation;

acquiring one or more knowledge elements from a first knowledge set of a plurality of knowledge sets, based on the recognized situation, wherein

the first knowledge set of the plurality of knowledge sets is based on a usage priority associated with the first knowledge set, and

the usage priority is set by the user based on a knowledge domain associated with the first knowledge set;

determining a type of utterance template of a plurality of types of utterance templates based on the intention analysis result of the first utterance of the user;

acquiring an utterance template based on the determined type of utterance template;

determining contents of a second utterance based on the one or more knowledge elements and the utterance template; and

outputting the second utterance to the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2021
From: KANAMORI, KATSUTOSHI; NISHIO, MASATO
To: SONY GROUP CORPORATION
Reel/Frame 056795/0918 →
Priority Claims (1)
JP 2019-008621 · Jan 22, 2019 · national
Continuity (1)
Related Publication 20220076672A1 · Mar 10, 2022
References Cited (19)
US 20060036433A1 · Davis · 2006 [cited by examiner]
US 20110161829A1 · Kristensen · 2011 [cited by examiner]
US 20130151555A1 · Miyano · 2013 [cited by examiner]
US 20160293162A1 · Takahashi et al. · 2016 [cited by applicant]
US 20170270925A1 · Kennewick · 2017 [cited by examiner]
US 20180342007A1 · Brannigan · 2018 [cited by examiner]
US 20190236140A1 · Canim · 2019 [cited by examiner]
US 20190384855A1 · Bhattacharya · 2019 [cited by examiner]
CN 106055547A · 2016 [cited by applicant]
CN 106575503A · 2017 [cited by applicant]
JP 2003108376A · 2003 [cited by applicant]
JP 2003280683A · 2003 [cited by applicant]
JP 2016197227A · 2016 [cited by applicant]
JP 2017058318A · 2017 [cited by applicant]
KR 101677859B1 · 2016 [cited by applicant]
WO WO2018142686A1 · 2018 [cited by applicant]
WO 2018163646A1 · 2018 [cited by applicant]
WO WO2019011356A1 · 2019 [cited by applicant]
International Search Report and Written Opinion of PCT Application No. PCT/JP2019/048579, issued on Feb. 25, 2020, 09 pages of ISRWO. [cited by applicant]