IP Library Granted Patent US 11,004,452
Granted Patent B2
US 11,004,452 · App. 16/598,449 · Granted May 11, 2021

Method and system for multimodal interaction with sound device connected to network

Inventors: Hyeoncheol Lee (Seongnam-si, KR); Jin Young Park (Seongnam-si, KR)
Assignees: NAVER CORPORATION; LINE CORPORATION
G10L15/22G06F3/167G10L15/30G10L25/90H04R1/406H04R3/005G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,004,452
App. No.
16/598,449
Granted
May 11, 2021
Kind
B2
Abstract

A method and a system for multimodal interaction with a sound device connected to a network are provided. The method for multimodal interaction comprises the steps of: outputting audio information for playing content through a voice-based interface included in an electronic device; receiving a speaker's voice input associated with the outputted audio information through the voice-based interface; generating location information associated with the speaker's voice input; and determining an operation associated with the playing of the content by using the voice input and the location information associated with the voice input.

Claims (50)

1. A multimodal interaction method of a multimodal interaction system including an electronic device, comprising:

outputting, by a processor of the electronic device, audio information for playing content through a voice-based interface included in the electronic device;

receiving, by the processor, a voice input of an utterer responsive to the output audio information through the voice-based interface;

generating, by the processor, location information associated with the voice input of the utterer;

determining, an operation associated with the playing of the content based on the voice input responsive to the output audio information and the location information associated with the voice input; and

performing, by the electronic device, the determined operation associated with the playing of the content.

2. The multimodal interaction method of claim 1 , wherein the location information associated with the voice input comprises at least one of a relative location or orientation of the utterer relative to the electronic device that is measured at a point in time or during a period of time associated with the reception of the voice input, whether the relative location or orientation is changed, a level of change in the relative location or orientation, and an orientation in which the relative location or orientation is changed.

3. The multimodal interaction method of claim 1 , wherein the generating of the location information comprises generating the location information associated with the voice input based on a phase shift of the voice input that is input to a plurality of microphones included in the voice-based interface.

4. The multimodal interaction method of claim 1 , wherein the electronic device comprises at least one of a camera and a sensor, and

the generating of the location information comprises generating the location information associated with the voice input based on an output value of at least one of the camera and the sensor in response to receiving the voice input.

5. The multimodal interaction method of claim 1 , wherein the determining of the operation associated with the playing of the content comprises integrating at least one of a tone of sound corresponding to the voice input, a pitch of the sound, and a command extracted by analyzing the voice input and the location information associated with the voice input.

6. The multimodal interaction method of claim 1 , further comprising:

receiving a measurement value that is measured by a sensor of a peripheral device interacting with the electronic device in association with the voice input, from the peripheral device,

wherein the determining of the operation associated with the playing of the content comprises using the received measurement value.

7. The multimodal interaction method of claim 1 , further comprising:

receiving a measurement value that is measured by a sensor of a peripheral device interacting with the electronic device regardless of the voice input, from the peripheral device; and

changing a setting associated with the playing of the content based on the received measurement value.

8. The multimodal interaction method of claim 1 , wherein the audio information comprises information that requires a change in a location of the utterer, and

the determining of the operation associated with the playing of the content depends on whether the voice input and the location information associated with the voice input meet a condition corresponding to the required information.

9. The multimodal interaction method of claim 1 , wherein the content is provided through an external server communicating with the electronic device over a network, and

the determining of the operation associated with the playing of the content comprises:

transmitting the voice input and the location information associated with the voice input to the external server over the network;

receiving, from the external server, operation information that is generated by the external server based on the voice input and the location information associated with the voice input; and

determining the operation associated with the playing of the content based on the received operation information.

10. A non-transitory computer-readable storage medium storing a program, which when executed by a processor, causing the processor to perform the multimodal interaction method of claim 1 .

11. A multimodal interaction system comprising:

a voice-based interface; and

at least one processor configured to execute computer-readable instructions,

wherein the at least one processor is configured to

output audio information for playing content through the voice-based interface,

receive a voice input of an utterer responsive to the output audio information through the voice-based interface,

generate location information associated with the voice input of the utterer,

determine an operation associated with the playing of the content based on the voice input responsive to the output audio information and the location information associated with the voice input, and

performing the determined operation associated with the playing of the content.

12. The multimodal interaction system of claim 11 , wherein the at least one processor is configured to generate the location information associated with the voice input based on a phase shift of the voice input that is input to a plurality of microphones included in the voice-based interface.

13. The multimodal interaction system of claim 11 , further comprising at least one of a camera and a sensor,

wherein the at least one processor is configured to generate the location information associated with the voice input based on an output value of at least one of the camera and the sensor in response to receiving the voice input.

14. The multimodal interaction system of claim 11 , wherein the at least one processor is configured to determine the operation associated with the playing of the content by integrating at least one of a tone of sound corresponding to the voice input, a pitch of the sound, and a command extracted by analyzing the voice input and the location information associated with the voice input.

15. The multimodal interaction system of claim 11 , wherein the at least one processor is configured to receive a measurement value that is measured by a sensor of a peripheral device interacting with the multimodal interaction system in association with the voice input, from the peripheral device, and

determine the operation associated with the playing of the content by using the received measurement value.

16. The multimodal interaction system of claim 11 , wherein the at least one processor is configured to

receive a measurement value that is measured by a sensor of a peripheral device interacting with the multimodal interaction system regardless of the voice input, from the peripheral device, and

change a setting associated with the playing of the content based on the measurement value.

17. The multimodal interaction system of claim 11 , wherein the audio information comprises information that requires an utterance of the utterer and a change in a location of the utterer, and

the at least one processor is configured to determine the operation associated with the playing of the content depending on whether the voice input and the location information associated with the voice input meet a condition corresponding to the required information.

18. The multimodal interaction system of claim 11 , wherein the content is provided through an external server performing communication over a network, and

to determine the operation associated with the playing of the content, the at least one processor is configured to

transmit the voice input and the location information associated with the voice input to the external server over the network,

receive, from the external server, operation information that is generated by the external server based on the voice input and the location information associated with the voice input, and

determine the operation associated with the playing of the content based on the received operation information.

Assignments (8)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 12, 2024
From: Z INTERMEDIATE GLOBAL CORPORATION
To: LY CORPORATION
Reel/Frame 067091/0109 →
CHANGE OF NAME Recorded Apr 10, 2024
From: LINE CORPORATION
To: Z INTERMEDIATE GLOBAL CORPORATION
Reel/Frame 067069/0467 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE CITY SHOULD BE SPELLED AS TOKYO PREVIOUSLY RECORDED AT REEL: 058597 FRAME: 0141. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 17, 2023
From: LINE CORPORATION
To: A HOLDINGS CORPORATION
Reel/Frame 062401/0328 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SPELLING OF THE ASSIGNEES CITY IN THE ADDRESS SHOULD BE TOKYO, JAPAN PREVIOUSLY RECORDED AT REEL: 058597 FRAME: 0303. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jan 17, 2023
From: A HOLDINGS CORPORATION
To: LINE CORPORATION
Reel/Frame 062401/0490 →
CHANGE OF NAME Recorded Dec 28, 2021
From: LINE CORPORATION
To: A HOLDINGS CORPORATION
Reel/Frame 058597/0141 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: A HOLDINGS CORPORATION
To: LINE CORPORATION
Reel/Frame 058597/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2019
From: NAVER CORPORATION
To: NAVER CORPORATION; LINE CORPORATION
Reel/Frame 051016/0910 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 10, 2019
From: LEE, HYEONCHEOL; PARK, JIN YOUNG
To: NAVER CORPORATION
Reel/Frame 050680/0903 →
Priority Claims (1)
KR 10-2017-0048304 · Apr 14, 2017 · national
Continuity (2)
Continuation PCTKR2018002075 · Feb 20, 2018
Related Publication 20200043491A1 · Feb 6, 2020