IP Library Granted Patent US 11,322,141
Granted Patent B2
US 11,322,141 · App. 16/635,281 · Granted May 3, 2022

Information processing device and information processing method

Inventors: Takao Okuda (Tokyo, JP); Takashi Shibuya (Tokyo, JP)
Assignee: SONY CORPORATION
G10L15/22G06F40/30G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,322,141
App. No.
16/635,281
Granted
May 3, 2022
Kind
B2
Abstract

An information processing device includes a communication controller that performs communication control for receiving transmission data transmitted from a client, transmitting the transmission data to a first service providing server that performs a first service process, receiving a first service process result from the first service providing server, transmitting data according to the first service process result to a second service providing server that performs a second service process that is different from a first service, receiving a second service process result from the second service providing server, and transmitting the second service process result to the client. The first service process result is obtained by performing the first service process on the transmission data. The second service process result is obtained by performing the second service process on the data according to the first service process result.

Claims (62)

1. An information processing device, comprising:

a communication controller configured to:

control reception of transmission data transmitted from a client, wherein the transmission data is voice data;

control transmission of the voice data to a first service providing server that performs a first service process corresponding to a first service, wherein the first service providing server performs a voice recognition process to recognize voice as the first service process;

control reception of a first service process result from the first service providing server, wherein

the first service process result is a voice recognition result, and

the voice recognition result is obtained based on the voice recognition process performed on the voice data;

control transmission of the voice recognition result to a second service providing server, wherein

the second service providing server performs a second service process corresponding to a second service,

the second service is different from the first service, and

the second service providing server performs a semantic analysis process to analyze a meaning and content of a character string as the second service process;

control reception of a second service process result from the second service providing server, wherein

the second service process result is a semantic analysis result, and

the semantic analysis result is obtained based on the semantic analysis process performed on the voice recognition result;

control transmission of the semantic analysis result to a response generation server that performs a generation process to generate a response in an interaction;

control reception of the response from the response generation server, wherein the response is obtained based on the generation process performed on the semantic analysis result;

control transmission of the response to a voice synthesis server that performs a voice synthesis process to generate synthetic sound data;

control reception of the synthetic sound data from the voice synthesis server, wherein the synthetic sound data is obtained based on the voice synthesis process performed on the response; and

control transmission of the synthetic sound data to the client.

2. The information processing device according to claim 1 , wherein

the first service providing server includes a voice recognition server that performs the voice recognition process to recognize the voice as the first service process, and

the second service providing server includes a semantic analysis server that performs the semantic analysis process to analyze the meaning and content of the character string as the second service process.

3. The information processing device according to claim 2 , wherein the communication controller is further configured to:

control reception of auxiliary information from the client, wherein the auxiliary information assists the voice recognition process or the semantic analysis process; and

control transmission of the auxiliary information to the voice recognition server or the semantic analysis server.

4. The information processing device according to claim 1 , wherein the communication controller is further configured to control transmission of the voice recognition result to the client based on the voice data received from the client.

5. An information processing method, comprising:

receiving transmission data transmitted from a client, wherein the transmission data is voice data;

transmitting the voice data to a first service providing server that performs a first service process corresponding to a first service, wherein the first service providing server performs a voice recognition process to recognize voice as the first service process;

receiving a first service process result from the first service providing server, wherein

the first service process result is a voice recognition result, and

the voice recognition result is obtained based on the voice recognition process performed on the voice data;

transmitting the voice recognition result to a second service providing server, wherein

the second service providing server performs a second service process corresponding to a second service,

the second service is different from the first service, and

the second service providing server performs a semantic analysis process to analyze a meaning and content of a character string as the second service process;

receiving a second service process result from the second service providing server, wherein

the second service process result is a semantic analysis result, and

the semantic analysis result obtained based on the semantic analysis process performed on the voice recognition result;

transmitting the semantic analysis result to a response generation server that performs a generation process to generate a response in an interaction;

receiving the response from the response generation server, wherein the response is obtained based on the generation process performed on the semantic analysis result;

transmitting the response to a voice synthesis server that performs a voice synthesis process to generate synthetic sound data;

receiving the synthetic sound data from the voice synthesis server, wherein the synthetic sound data is obtained based on the voice synthesis process performed on the response; and

transmitting the synthetic sound data to the client.

6. A non-transitory computer-readable medium having stored thereon computer-executable instructions, that when executed by a processor, cause the processor to execute operations, the operations comprising:

receiving transmission data transmitted from a client, wherein the transmission data is voice data;

transmitting the voice data to a first service providing server that performs a first service process corresponding to a first service, wherein the first service providing server performs a voice recognition process to recognize voice as the first service process;

receiving a first service process result from the first service providing server, wherein

the first service process result is a voice recognition result, and

the voice recognition result is obtained based on the voice recognition process performed on the voice data;

transmitting the voice recognition result to a second service providing server, wherein

the second service providing server performs a second service process corresponding to a second service,

the second service is different from the first service, and

the second service providing server performs a semantic analysis process to analyze a meaning and content of a character string as the second service process;

receiving a second service process result from the second service providing server, wherein

the second service process result is a semantic analysis result, and

the semantic analysis result is obtained based on the semantic analysis process performed on the voice recognition result;

transmitting the semantic analysis result to a response generation server that performs a generation process to generate a response in an interaction;

receiving the response from the response generation server, wherein the response is obtained based on the generation process performed on the semantic analysis result;

transmitting the response to a voice synthesis server that performs a voice synthesis process to generate synthetic sound data;

receiving the synthetic sound data from the voice synthesis server, wherein the synthetic sound data is obtained based on the voice synthesis process performed on the response; and

transmitting the synthetic sound data to the client.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 30, 2020
From: OKUDA, TAKAO; SHIBUYA, TAKASHI
To: SONY CORPORATION
Reel/Frame 051758/0154 →
Priority Claims (1)
JP JP2017-157538 · Aug 17, 2017 · national
Continuity (1)
Related Publication 20200372910A1 · Nov 26, 2020