IP Library Granted Patent US 12,499,892
Granted Patent B2
US 12,499,892 · App. 18/744,076 · Granted Dec 16, 2025

Method and apparatus to provide comprehensive smart assistant services

Inventors: Hung Bun Choi (Hong Kong, CN); Kai Nang Pang (Hong Kong, CN); Yau Wai Ng (Hong Kong, CN); Chi Chung Liu (Hong Kong, CN); Tsz Kin Lee (Hong Kong, CN)
Assignee: Computime Ltd.
G10L15/22G10L15/08G10L15/30G10L2015/088G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,892
App. No.
18/744,076
Granted
Dec 16, 2025
Kind
B2
Abstract

An apparatus supports smart assistant services with a plurality of smart service providers. The apparatus includes an audio device that receives a speech signal having a user utterance, captures the user utterance when the user utterance includes a user wake word, and sends the captured utterance to a backend computing device. The backend computing device replaces the user wake word with specific wake words associated with different smart service providers. The processed utterances are then sent to selected smart service providers. The backend computing device subsequently constructs feedback to the user utterance based on voice responses from the different smart service providers. The backend computing device then passes a digital representation of the feedback to the audio device, and the audio device converts the digital representation to an audio reply to the user utterance.

Claims (86)

1 . An apparatus comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, configure the apparatus to:

detect a user trigger;

capture, based on the detected user trigger, a user utterance from a user;

include a first specific wake word in the captured user utterance to form a first processed utterance;

send the first processed utterance to a first service provider;

obtain, from the first service provider, a first response to the first processed utterance; and

construct, based on data associated with a history of processed utterances, first feedback to the user utterance based on the first response; and

generate, based on the first feedback, a reply to the user utterance.

2 . The apparatus of claim 1 , wherein the user trigger comprises one or more of a wake word, a sound, an audio condition, a hand gesture, a body gesture, a facial expression or biology signature.

3 . The apparatus of claim 1 , wherein the instructions, when executed by the one or more processors, further configure the apparatus to detect the user trigger based on a model trained on user input.

4 . The apparatus of claim 1 , wherein the data associated with the history of processed utterances comprises one or more of:

scoring metrics indicative of accuracies of feedbacks to the processed utterances;

a history of user utterances following processed utterances; or

data indicating response speeds of a plurality of service providers to one or more previous captured user utterances.

5 . The apparatus of claim 1 , wherein the instructions, when executed by the one or more processors, further configure the apparatus to:

include a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

send the second processed utterance to a second service provider;

obtain a second response from the second service provider, and

construct the first feedback by combining, based on the data associated with the history of processed utterances, the first response and the second response.

6 . The apparatus of claim 1 , wherein the instructions, when executed by the one or more processors, further configure the apparatus to:

include a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

send the second processed utterance to a second service provider;

obtain a second response from the second service provider,

construct second feedback based on the second response; and

update, based on the first feedback and on the second feedback, the data associated with the history of processed utterances.

7 . The apparatus of claim 1 , wherein the instructions, when executed by the one or more processors, further configure the apparatus to:

obtain a scoring function by training on the data associated with the history of processed utterances, wherein the scoring function measures a probabilistic prediction accuracy; and

construct the first feedback based on a scoring metric, obtained from the scoring function applied to the first response, satisfying a threshold.

8 . A method comprising:

detecting, by a computing device via a sensor, a user trigger;

capturing, based on the detected user trigger, a user utterance from a user;

including a first specific wake word in the captured user utterance to form a first processed utterance;

sending the first processed utterance to a first service provider;

obtaining, from the first service provider, a first response to the first processed utterance; and

constructing, based on data associated with a history of processed utterances, first feedback to the user utterance based on the first response; and

generating, based on the first feedback, a reply to the user utterance.

9 . The method of claim 8 , wherein the user trigger comprises one or more of a wake word, a sound, an audio condition, a hand gesture, a body gesture, a facial expression or biology signature.

10 . The method of claim 8 , further comprising detecting the user trigger based on a model trained on user input.

11 . The method of claim 8 , wherein the data associated with the history of processed utterances comprises one or more of:

scoring metrics indicative of accuracies of feedbacks to the processed utterances;

a history of user utterances following processed utterances; or

data indicating response speeds of a plurality of service providers to one or more previous captured user utterances.

12 . The method of claim 8 , further comprising:

including a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

sending the second processed utterance to a second service provider;

obtaining a second response from the second service provider, and

constructing the first feedback by combining, based on the data associated with the history of processed utterances, the first response and the second response.

13 . The method of claim 8 , further comprising

including a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

sending the second processed utterance to a second service provider;

obtaining a second response from the second service provider,

constructing second feedback based on the second response; and

updating, based on the first feedback and on the second feedback, the data associated with the history of processed utterances.

14 . The method of claim 8 , further comprising:

obtaining a scoring function by training on the data associated with the history of processed utterances, wherein the scoring function measures a probabilistic prediction accuracy; and

constructing the first feedback based on a scoring metric, obtained from the scoring function applied to the first response, satisfying a threshold.

15 . A non-transitory computer readable medium storing instructions that, when executed, cause:

detecting a user trigger;

capturing, based on the detected user trigger, a user utterance from a user;

including a first specific wake word in the captured user utterance to form a first processed utterance;

sending the first processed utterance to a first service provider;

obtaining, from the first service provider, a first response to the first processed utterance; and

constructing, based on data associated with a history of processed utterances, first feedback to the user utterance based on the first response; and

generating, based on the first feedback, a reply to the user utterance.

16 . The non-transitory computer readable medium of claim 15 , wherein the user trigger comprises one or more of a wake word, a sound, an audio condition, a hand gesture, a body gesture, a facial expression or biology signature.

17 . The non-transitory computer readable medium of claim 15 , further comprising detecting the user trigger based on a model trained on user input.

18 . The non-transitory computer readable medium of claim 15 , wherein the data associated with the history of processed utterances comprises one or more of:

scoring metrics indicative of accuracies of feedbacks to the processed utterances;

a history of user utterances following processed utterances; or

data indicating response speeds of a plurality of service providers to one or more previous captured user utterances.

19 . The non-transitory computer readable medium of claim 15 , further comprising:

including a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

sending the second processed utterance to a second service provider;

obtaining a second response from the second service provider, and

constructing the first feedback by combining, based on the data associated with the history of processed utterances, the first response and the second response.

20 . The non-transitory computer readable medium of claim 15 , further comprising

including a second specific wake word, different from the first specific wake word, in the user utterance to form a second processed utterance;

sending the second processed utterance to a second service provider;

obtaining a second response from the second service provider,

constructing second feedback based on the second response; and

updating, based on the first feedback and on the second feedback, the data associated with the history of processed utterances.

21 . The non-transitory computer readable medium of claim 15 , further comprising:

obtaining a scoring function by training on the data associated with the history of processed utterances, wherein the scoring function measures a probabilistic prediction accuracy; and

constructing the first feedback based on a scoring metric, obtained from the scoring function applied to the first response, satisfying a threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2025
From: CHOI, HUNG BUN; PANG, KAI NANG; NG, YAU WAI; LIU, CHI CHUNG; LEE, TSZ KIN
To: COMPUTIME LTD.
Reel/Frame 071602/0097 →
Continuity (4)
Continuation 17240340 · Apr 26, 2021
Continuation 16133899 · Sep 18, 2018
Provisional Application 62627929 · Feb 8, 2018
Related Publication 20240371374A1 · Nov 7, 2024
References Cited (17)
US 1073552A · Vaish · 1913 [cited by applicant]
US 9026441B2 · Ehsani · 2015 [cited by examiner]
US 10379808B1 · Angel et al. · 2019 [cited by applicant]
US 10504511B2 · Wang · 2019 [cited by examiner]
US 10540451B2 · Prendergast et al. · 2020 [cited by applicant]
US 10614799B2 · Kennewick, Jr. · 2020 [cited by examiner]
US 10614811B2 · Gabel et al. · 2020 [cited by applicant]
US 10735916B2 · Gilmartin et al. · 2020 [cited by applicant]
US 11024307B2 · Choi et al. · 2021 [cited by applicant]
US 12051410B2 · Choi · 2024 [cited by examiner]
US 20180096334A1 · Studnicka et al. · 2018 [cited by applicant]
US 20180210874A1 · Fuxman et al. · 2018 [cited by applicant]
US 20180341643A1 · Alders et al. · 2018 [cited by applicant]
WO 2005054230A1 · 2005 [cited by applicant]
WO 2016054230A1 · 2016 [cited by applicant]
WO 2017095476A1 · 2017 [cited by applicant]
Jul. 2, 2019—(EP)—Extended European Search Report—App 19154543.3. [cited by applicant]