IP Library › Granted Patent US 12,223,957
Granted Patent B2
US 12,223,957 · App. 17/752,094 · Granted Feb 11, 2025

Voice interaction method, system, terminal device and medium

Inventor: Yingjie Li (Beijing, CN)
Assignee: BOE Technology Group Co., Ltd.
G10L15/22G10L25/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,223,957
App. No.
17/752,094
Granted
Feb 11, 2025
Kind
B2
Abstract

The present disclosure discloses a voice interaction method, system, terminal device and medium, with operations performing voice recognition on collected voice signals to acquire an input sentence; semantically matching the input sentence with cached sample sentences, determining whether there is a sample sentence having same or similar semantics as the input sentence among the cached sample sentences; if yes, acquiring cached response content having the same or similar semantics as the input sentence; if not, sending at least one of the input sentence or the collected voice signals to a server; receiving response content or the collected voice signals, as response content of the input sentence, acquired through semantic understanding according to a knowledge base, responding to the input sentence according to the response content of the input sentence, and updating at least one of the cached sample sentences or the response content of the cached sample sentences.

Claims (75)

1. A method performed by a terminal device, comprising:

performing voice recognition on collected voice signals to acquire an input sentence;

performing semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences;

in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquiring cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence;

in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, sending at least one of the input sentence or the collected voice signals to a server, and receiving transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, wherein the transmitted response content of the at least one of the input sentence or the collected voice signals from the server is acquired by the server through semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server; and

responding to the input sentence according to the first response content of the input sentence, wherein the first response content of the input sentence further comprises a control instruction and the responding to the input sentence according to the first response content of the input sentence includes performing a corresponding action according to the control instruction.

2. The method according to claim 1 , wherein:

the first response content of the input sentence further comprises voice response content; and

the responding to the input sentence according to the first response content of the input sentence further includes carrying out a voice broadcast on the voice response content.

3. The method according to claim 2 , wherein:

the control instruction comprises one or more execution instructions for controlling one or more applications to perform one or more operations.

4. The method according to claim 3 , wherein:

the control instruction is preset manually by a user of the terminal device or preset automatically by the terminal device based on a user profile of the user or by the server based on the user profile of the user.

5. The method according to claim 2 , wherein:

the control instruction comprises a plurality of execution instructions that are to be executed in a preset execution sequence.

6. The method according to claim 5 , wherein:

the execution sequence of the plurality of execution instructions is preset based on at least one of:

a manual configuration by a user of the terminal device;

configurations by the terminal device based on a user profile of the user or by the server based on the user profile of the user; or

big data obtained by the server.

7. The method according to claim 1 , wherein at least one of the cached sample sentences or response content of the cached sample sentences is pre-configured and/or updated based on at least one of:

a manual configuration by a user of the terminal device;

an initial system configuration;

configurations based on a user profile that are at the terminal device or at the server; or

big data obtained by the server.

8. The method according to claim 1 , further comprising:

updating at least one of the cached sample sentences or response content of the cached sample sentences according to the input sentence and the first response content of the input sentence.

9. The method according to claim 8 , wherein the updating at least one of the cached sample sentences or the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence further comprises:

determining an acquisition frequency of the input sentence;

comparing the acquisition frequency of the input sentence to a first preset threshold; and

in response to determining that the acquisition frequency of the input sentence is greater than the first preset threshold, update at least one of the cached sample sentences and the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence.

10. A terminal device, comprising:

a memory, storing computer instructions thereon; and

a processor coupled to the memory, wherein when the processor executes the computer instructions, the processor is configured to:

perform voice recognition on collected voice signals to acquire an input sentence;

performing semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences;

in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence;

in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, wherein the transmitted response content of the at least one of the input sentence or the collected voice signals from the server is acquired by the server through semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server; and

respond to the input sentence according to the first response content of the input sentence, wherein the first response content of the input sentence further comprises a control instruction and the responding to the input sentence according to the first response content of the input sentence includes performing a corresponding action according to the control instruction.

11. The terminal device according to claim 10 , wherein:

the first response content of the input sentence further comprises voice response content; and

when the processor executes the computer instructions, the processor is further configured to respond to the input sentence according to the first response content of the input sentence by

carrying out a voice broadcast on the voice response content.

12. The method according to claim 11 , wherein:

the control instruction comprises one or more execution instructions for controlling one or more applications to perform one or more operations.

13. The method according to claim 12 , wherein:

the control instruction is preset manually by a user of the terminal device or preset automatically by the terminal device based on a user profile of the user or by the server based on the user profile of the user.

14. The method according to claim 11 , wherein:

the control instruction comprises a plurality of execution instructions that are to be executed in a preset execution sequence.

15. The method according to claim 14 , wherein:

the execution sequence of the plurality of execution instructions is preset based on at least one of:

a manual configuration by a user of the terminal device;

configurations by the terminal device based on a user profile of the user or by the server based on the user profile of the user; or

big data obtained by the server.

16. The terminal device according to claim 10 , wherein at least one of the cached sample sentences or response content of the cached sample sentences are pre-configured and/or updated based on at least one of:

a manual configuration by a user of the terminal device;

an initial system configuration;

configurations based on a user profile that are at the terminal device or at the server; or

big data obtained by the server.

17. The terminal device according to claim 10 , wherein when the processor executes the computer instructions, the processor is further configured to:

update at least one of the cached sample sentences or response content of the cached sample sentences according to the input sentence and the first response content of the input sentence.

18. The terminal device according to claim 17 , wherein when the processor executes the computer instructions, the processor is further configured to update at least one of the cached sample sentences or the response content of the cached sample sentences according to the input sentence and the first response content of the input sentence by performing the following operations:

determining an acquisition frequency of the input sentence;

comparing the acquisition frequency of the input sentence to a first preset threshold; and

in response to determining that the acquisition frequency of the input sentence is greater than the first preset threshold, updating at least one of the cached sample sentences and the response content of the cached sample sentences using the input sentence and the first response content of the input sentence.

19. A voice interaction system, comprising:

a terminal device configured to:

perform voice recognition on collected voice signals to acquire an input sentence, perform semantic matching between the input sentence and cached sample sentences to determine whether there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences,

in response to determining that there is a sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, acquire cached response content corresponding to the sample sentence having the same or similar semantics as the input sentence as first response content of the input sentence,

in response to determining that there is no sample sentence having the same or similar semantics as the input sentence among the cached sample sentences, send at least one of the input sentence or the collected voice signals to a server, and receive transmitted response content of the at least one of the input sentence or the collected voice signals from the server as the first response content of the input sentence, and

respond to the input sentence according to the first response content of the input sentence, wherein the first response content of the input sentence further comprises a control instruction and the responding to the input sentence according to the first response content of the input sentence includes performing a corresponding action according to the control instruction; and

the server, wherein the server is configured to:

receive the at least one of the input sentence or the collected voice signals from the terminal device, and

perform semantic understanding of the at least one of the input sentence or the collected voice signals according to a knowledge base stored on the server to acquire the transmitted response content of the at least one of the input sentence or the collected voice signals, and send the transmitted response content of the at least one of the input sentence or the collected voice signals to the terminal device.

20. A non-transitory computer-readable storage medium having stored thereon a computer program, wherein the program, when executed by a processor, implements the method according to claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2022
From: LI, YINGJIE
To: BOE TECHNOLOGY GROUP CO., LTD.
Reel/Frame 060174/0723 →
Priority Claims (1)
CN 201910808807.0 · Aug 29, 2019 · national
Continuity (2)
Continuation In Part 16818145 · Mar 13, 2020
Related Publication 20220284900A1 · Sep 8, 2022
References Cited (64)
US 4403845A · Buelens et al. · 1983 [cited by applicant]
US 5668958A · Bendert et al. · 1997 [cited by applicant]
US 5682542A · Enomoto · 1997 [cited by examiner]
US 6377944B1 · Busey · 2002 [cited by examiner]
US 7315826B1 · Guheen et al. · 2008 [cited by applicant]
US 9990176B1 · Gray · 2018 [cited by examiner]
US 10558633B1 · Kim · 2020 [cited by applicant]
US 10713289B1 · Mishra et al. · 2020 [cited by applicant]
US 12093250B2 · Lai · 2024 [cited by examiner]
US 20020023144A1 · Linyard et al. · 2002 [cited by applicant]
US 20020046291A1 · O'Callaghan et al. · 2002 [cited by applicant]
US 20020198883A1 · Nishizawa et al. · 2002 [cited by applicant]
US 20030046434A1 · Flanagin et al. · 2003 [cited by applicant]
US 20040030556A1 · Bennett · 2004 [cited by examiner]
US 20040044516A1 · Kennewick · 2004 [cited by examiner]
US 20040230897A1 · Latzel · 2004 [cited by applicant]
US 20050240413A1 · Asano et al. · 2005 [cited by applicant]
US 20060020473A1 · Hiroe et al. · 2006 [cited by applicant]
US 20080104065A1 · Agarwal · 2008 [cited by examiner]
US 20090138477A1 · Piira et al. · 2009 [cited by applicant]
US 20090327234A1 · Coladonato et al. · 2009 [cited by applicant]
US 20100057723A1 · Rajaram · 2010 [cited by applicant]
US 20120015642A1 · Sec · 2012 [cited by applicant]
US 20120303358A1 · Ducatel · 2012 [cited by examiner]
US 20140200896A1 · Lee · 2014 [cited by examiner]
US 20150161996A1 · Petrov · 2015 [cited by examiner]
US 20160147873A1 · Henmi · 2016 [cited by examiner]
US 20160196258A1 · Ma et al. · 2016 [cited by applicant]
US 20160217129A1 · Lu · 2016 [cited by examiner]
US 20160294970A1 · Robinson · 2016 [cited by applicant]
US 20170295236A1 · Kulkarni et al. · 2017 [cited by applicant]
US 20180150739A1 · Wu · 2018 [cited by examiner]
US 20180166077A1 · Yamaguchi et al. · 2018 [cited by applicant]
US 20180190292A1 · Xu · 2018 [cited by applicant]
US 20180285348A1 · Shu et al. · 2018 [cited by applicant]
US 20180285928A1 · Kirmani et al. · 2018 [cited by applicant]
US 20180349256A1 · Fong · 2018 [cited by examiner]
US 20180373758A1 · Obradovic · 2018 [cited by examiner]
US 20190103102A1 · Tseretopoulos · 2019 [cited by examiner]
US 20190114082A1 · Pan et al. · 2019 [cited by applicant]
US 20190139544A1 · Chen et al. · 2019 [cited by applicant]
US 20190163735A1 · Xu et al. · 2019 [cited by applicant]
US 20190180756A1 · Yang · 2019 [cited by applicant]
US 20190354630A1 · Guo · 2019 [cited by examiner]
US 20190370273A1 · Frison · 2019 [cited by applicant]
US 20220114824A1 · Tomita · 2022 [cited by examiner]
CN 105206275A · 2015 [cited by applicant]
CN 106710596A · 2017 [cited by applicant]
CN 106875950A · 2017 [cited by applicant]
CN 106897263A · 2017 [cited by applicant]
CN 107102982A · 2017 [cited by applicant]
CN 107945798A · 2018 [cited by applicant]
CN 108376544A · 2018 [cited by applicant]
CN 108549637A · 2018 [cited by applicant]
CN 108932342A · 2018 [cited by applicant]
CN 109102809A · 2018 [cited by applicant]
CN 109147788A · 2019 [cited by applicant]
CN 109360555A · 2019 [cited by applicant]
CN 112185370A · 2021 [cited by examiner]
EP 1344148B1 · 2005 [cited by applicant]
Google Translation of CN 112185370 A , 2021, https://patents.google.com/patent/CN112185370A/en?oq=CN+112185370+A+ (Year: 2021). [cited by examiner]
Non-Final Office Action dated Oct. 5, 2021 for U.S. Appl. No. 16/818,145. [cited by applicant]
Examiner-Initiated Interview Summary dated Feb. 9, 2022 for U.S. Appl. No. 16/818,145. [cited by applicant]
Notice of Allowance and Fee(s) Due dated Feb. 24, 2022 for U.S. Appl. No. 16/818,145. [cited by applicant]