IP Library Granted Patent US 10,971,145
Granted Patent B2
US 10,971,145 · App. 16/179,436 · Granted Apr 6, 2021

Speech interaction feedback method for smart TV, system and computer readable medium

Inventors: Junnan Luo (Beijing, CN); Jing Li (Beijing, CN); Zhixi Chen (Beijing, CN)
Assignee: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
G10L15/22G06F3/167G10L15/1815H04N21/2401H04N21/4394H04N21/472G10L15/00G10L15/26G10L15/30G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,971,145
App. No.
16/179,436
Granted
Apr 6, 2021
Kind
B2
Abstract

The present disclosure provides a speech interaction feedback method for smart TV, a system and a computer readable medium. The method comprises: collecting audio stream of a speech query sent by a user and element information of a current interface of the smart TV; sending the audio stream and the element information of the current interface to a cloud server so that the cloud server generates an information response message carrying a target element, according to the audio stream and the element information of the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; receiving the response message returned by the cloud server; according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query. According to the technical solution of the present disclosure, it is possible to perform feedback for the user's speech query on the smart TV. As such, when the smart TV does not execute the control instruction, it is possible to accurately determine whether a reason for none execution of the control instruction is none recognition or blockage during execution.

Claims (81)

1. A speech interaction feedback method for smart TV, wherein the method comprises:

collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of the elements comprises a position, displayed words and hierarchical structure information of each element in the current interface;

sending the audio stream and the information of elements displayed on the current interface to a cloud server;

receiving, from the cloud server, an information response message carrying a target element, which is determined by the cloud server according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

2. The method according to claim 1 , wherein after collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, the method further comprises:

storing the information of elements displayed on the current interface in a buffer;

the step of, according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface specifically comprises:

looking up the information of elements displayed on the current interface stored in the buffer for the corresponding target element, according to information of the target element in the response message;

performing a preset effect display for the corresponding target element in the current interface.

3. A speech interaction feedback method for smart TV, wherein the method comprises:

receiving audio stream corresponding to a user's speech query and information of elements displayed on a current interface of the smart TV sent by the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of elements displayed on the current interface comprises a position, displayed words and hierarchical structure information of each element in the current interface;

generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

returning the response message to the smart TV so that the smart TV, according to information of the target element in the response message, performs a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

4. The method according to claim 3 , wherein the generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface specifically comprises:

according to audio stream and the information of elements displayed on the current interface, recognizing a target element in the current interface hit by an intention of the speech query corresponding to the audio stream;

generating the response message based on information of the target element.

5. The method according to claim 4 , wherein the step of, according to audio stream and the information of elements displayed on the current interface, recognizing a target element in the current interface hit by an intention of the speech query corresponding to the audio stream specifically comprises:

performing speech recognition for the audio stream to obtain a word instruction corresponding to the speech query corresponding to the audio stream;

performing natural language understanding processing for the word instruction and recognizing an intention of the speech query;

comparing the intention of the speech query with the information of elements displayed on the current interface, and recognizing the target element in the current interface hit by the intention of the speech query.

6. A computer device, wherein the device comprises:

one or more processors,

a storage for storing one or more programs,

the one or more programs, when executed by said one or more processors, enable said one or more processors to implement a speech interaction feedback method for smart TV, wherein the method comprises:

collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of the elements comprises a position, displayed words and hierarchical structure information of each element in the current interface;

sending the audio stream and the information of elements displayed on the current interface to a cloud server;

receiving, from the cloud server, an information response message carrying a target element, which is generated by the cloud server according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

7. The computer device according to claim 6 , wherein after collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, the method further comprises:

storing the information of elements displayed on the current interface in a buffer;

the step of, according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface specifically comprises:

looking up the information of elements displayed on the current interface stored in the buffer for the corresponding target element, according to information of the target element in the response message;

performing a preset effect display for the corresponding target element in the current interface.

8. A computer device, wherein the device comprises:

one or more processors,

a storage for storing one or more programs,

the one or more programs, when executed by said one or more processors, enable said one or more processors to implement a speech interaction feedback method for smart TV, wherein the method comprises:

receiving audio stream corresponding to a user's speech query and information of elements displayed on a current interface of the smart TV sent by the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of elements displayed on the current interface comprises a position, displayed words and hierarchical structure information of each element in the current interface;

generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

returning the response message to the smart TV so that the smart TV, according to information of the target element in the response message, performs a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

9. The computer device according to claim 8 , wherein the generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface specifically comprises:

according to audio stream and the information of elements displayed on the current interface, recognizing a target element in the current interface hit by an intention of the speech query corresponding to the audio stream;

generating the response message based on information of the target element.

10. The computer device according to claim 9 , wherein the step of, according to audio stream and the information of elements displayed on the current interface, recognizing a target element in the current interface hit by an intention of the speech query corresponding to the audio stream specifically comprises:

performing speech recognition for the audio stream to obtain a word instruction corresponding to the speech query corresponding to the audio stream;

performing natural language understanding processing for the word instruction and recognizing an intention of the speech query; and

comparing the intention of the speech query with the information of elements displayed on the current interface, and recognizing the target element in the current interface hit by the intention of the speech query.

11. A speech interaction system for smart TV, wherein the system comprises a smart TV apparatus and a cloud server, the smart TV apparatus is communicatively connected with the cloud server,

where in the smart TV apparatus comprises:

one or more first processors,

a first storage for storing one or more first programs,

the one or more first programs, when executed by said one or more first processors, enable said one or more first processors to implement the followings:

collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV apparatus, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of elements displayed on the current interface comprises a position, displayed words and hierarchical structure information of each element in the current interface;

sending the audio stream and the information of elements displayed on the current interface to the cloud server;

receiving, from the cloud server, an information response message carrying a target element, which is generated by the cloud server according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query,

and wherein the cloud server comprises:

one or more second processors,

a second storage for storing one or more second programs,

the one or more second programs, when executed by said one or more second processors, enable said one or more second processors to implement the followings:

receiving the audio stream corresponding to the user's speech query and the information of elements displayed on the current interface of the smart TV apparatus sent by the smart TV apparatus;

generating the information response message carrying the target element, according to the audio stream and the information of elements displayed on the current interface;

returning the response message to the smart TV apparatus so that the smart TV apparatus, according to information of the target element in the response message, performs the preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

12. A non-transitory computer readable medium on which a computer program is stored, wherein the program, when executed by a processor, implements a speech interaction feedback method for smart TV, wherein the method comprises:

collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of elements displayed on the current interface comprises a position, displayed words and hierarchical structure information of each element in the current interface;

sending the audio stream and the information of elements displayed on the current interface to a cloud server;

receiving, from the cloud server, an information response message carrying a target element, which is generated by the cloud server according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

13. The non-transitory computer readable medium according to claim 12 , wherein after collecting audio stream of a speech query sent by a user and information of elements displayed on a current interface of the smart TV, the method further comprises:

storing the information of elements displayed on the current interface in a buffer;

the step of, according to information of the target element in the response message, performing a preset effect display for the corresponding target element on the current interface specifically comprises:

looking up the information of elements displayed on the current interface stored in the buffer for the corresponding target element, according to information of the target element in the response message;

performing a preset effect display for the corresponding target element in the current interface.

14. A non-transitory computer readable medium on which a computer program is stored, wherein the program, when executed by a processor, implements a speech interaction feedback method for smart TV, wherein the method comprises:

receiving audio stream corresponding to a user's speech query and information of elements displayed on a current interface of the smart TV sent by the smart TV, wherein the elements are all elements explicitly displayed on the current interface corresponding to executable operations, and the information of elements displayed on the current interface comprises a position, displayed words and hierarchical structure information of each element in the current interface;

generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface; wherein the target element is an element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

returning the response message to the smart TV so that the smart TV, according to information of the target element in the response message, performs a preset effect display for the corresponding target element on the current interface, as an interaction feedback for the speech query.

15. The non-transitory computer readable medium according to claim 14 , wherein the generating an information response message carrying a target element, according to the audio stream and the information of elements displayed on the current interface specifically comprises:

according to audio stream and the information of elements displayed on the current interface, recognizing a target element in the current interface hit by an intention of the speech query corresponding to the audio stream; and

generating the response message based on information of the target element.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 30, 2021
From: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.; SHANGHAI XIAODU TECHNOLOGY CO. LTD.
Reel/Frame 056811/0772 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2019
From: LUO, JUNNAN; LI, JING; CHEN, ZHIXI
To: BAIDU ONLINE NETWORK TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 050331/0534 →
Priority Claims (1)
CN 201810195553.5 · Mar 9, 2018 · national
Continuity (1)
Related Publication 20190279628A1 · Sep 12, 2019