IP Library Granted Patent US 12,488,787
Granted Patent B2
US 12,488,787 · App. 18/090,064 · Granted Dec 2, 2025

System for processing voice requests including wake word verification

Inventors: Daniel Bromand (Boston, MA); Björn Erik Roth (Stockholm, SE); Nick Priem (Rotterdam, NE)
Assignee: Spotify AB
G10L15/08G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,787
App. No.
18/090,064
Granted
Dec 2, 2025
Kind
B2
Abstract

A system for processing voice requests includes a voice assistant manager and a plurality of voice assistants. The voice assistant manager detects a wake word in an utterance and communicates the utterance to a voice assistant of the plurality of voice assistants. In some embodiments, the voice assistant may verify the detected wake word and communicate with a cloud service, which may also verify the detected wake word and generate a response to the utterance. In some embodiments, the voice assistant manager may activate or deactivate one or more of the voice assistants.

Claims (74)

1 . A system for processing voice requests, the system comprising:

a voice assistant manager;

a plurality of voice assistants;

a processor; and

memory coupled to the processor, wherein the memory stores instructions that, when executed by the processor, cause the voice assistant manager to:

receive an utterance from a user,

detect a wake word in the utterance using a first wake word detection model,

based on the wake word, identify a called assistant from the plurality of voice assistants, and

communicate the utterance to the called assistant,

wherein the instructions, when executed by the processor, cause the called assistant to:

receive the utterance from the voice assistant manager; and verify the wake word in the received utterance using a second wake word detection model,

wherein the second wake word detection model is trained to recognize one or more wake words,

wherein verifying the wake word comprises inputting the utterance into the second wake word detection model, and

wherein the system is further configured to generate a response to the utterance and transmit the response to the user.

2 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the voice assistant manager to:

deactivate the called assistant; and

listen for a second utterance.

3 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the voice assistant manager to:

determine whether the called assistant is active; and

in response to determining that the called assistant is not active, activate the called assistant before communicating the utterance to the called assistant.

4 . The system of claim 3 , wherein the instructions, when executed by the processor, further cause the voice assistant manager to, before activating the called assistant, deactivate a second assistant of the plurality of voice assistants.

5 . The system of claim 1 , wherein the called assistant is communicatively coupled to the voice assistant manager via a Matter network.

6 . The system of claim 1 , further comprising a cloud service associated with the called assistant;

wherein the instructions, when executed by the processor, further cause the called assistant to, in response to successfully verifying the wake word, communicate the utterance to the cloud service; and

wherein the cloud service verifies the wake word.

7 . The system of claim 6 , wherein the cloud service, in response to successfully verifying the wake word, processes a request of the utterance.

8 . The system of claim 6 , wherein the cloud service, in response to failing to verify the wake word, returns an error to the called assistant and deletes data associated with the utterance.

9 . The system of claim 6 , wherein the instructions, when executed by the processor, cause the called assistant to:

receive the response from the cloud service; and

communicate the response to one or more of the user or the voice assistant manager.

10 . The system of claim 6 , wherein one or more of communicating the utterance to the cloud service or communicating the utterance to the called assistant comprises sending a plurality of audio files, the plurality of audio files including an encrypted audio file and an unencrypted audio file;

wherein the unencrypted audio file includes the wake word; and

wherein the encrypted audio file includes one or more of the utterance or a request of the utterance.

11 . The system of claim 1 , wherein the instructions, when executed by the processor, further cause the voice assistant manager to:

receive a subscription request from the called assistant, the subscription request including the wake word;

associate the wake word with the called assistant; and

train a machine learning model to recognize the wake word.

12 . The system of claim 1 , further comprising a computing device, wherein the computing device includes the voice assistant manager, the plurality of voice assistants, the processor, and the memory.

13 . The system of claim 12 , wherein the computing device includes a screen displaying a user interface; and

wherein the user interface includes a plurality of voice assistant icons, wherein each icon of the plurality of voice assistant icons corresponds with a voice assistant of the plurality of voice assistants.

14 . A method for processing voice requests, the method comprising:

receiving an utterance from a user;

detecting, at a voice assistant manager, a wake word in the utterance using a first wake word detection model;

identifying, from a plurality of voice assistants, a called assistant associated with the wake word;

communicating the utterance to the called assistant;

detecting, at the called assistant, the wake word in the utterance using a second wake word detection model, wherein the second wake word detection model is trained to recognize one or more wake words, and wherein detecting the wake word at the called assistant comprises inputting the utterance into the second wake word detection model;

generating a response to the utterance; and

transmitting the response to the user.

15 . The method of claim 14 , further comprising:

transmitting the utterance to a cloud service; and

detecting, at the cloud service, the wake word in the utterance;

wherein generating the response to the utterance is performed at the cloud service.

16 . The method of claim 14 , further comprising:

prior to communicating the utterance to the called assistant, activating the called assistant;

determining that the called assistant finished processing the utterance; and

deactivating the called assistant.

17 . The method of claim 14 , further comprising:

subscribing the called assistant;

wherein subscribing the called assistant comprises (i) receiving one or more wake words associated with the called assistant, the one or more wake words associated with the called assistant including the wake word and (ii) training a machine learning model to detect the one or more wake words,

wherein detecting, at the voice assistant manager, the wake word in the utterance comprises inputting the utterance into the machine learning model.

18 . A device for processing voice commands, the device comprising:

a processor; and

memory coupled to the processor, the memory storing instructions that, when executed by the processor cause the device to:

receive an utterance;

detect a wake word in the utterance using a first wake word detection model;

identify, from a plurality of voice assistants, a called assistant associated with the wake word;

communicate the utterance to the called assistant;

detecting, at the called assistant, the wake word in the utterance using a second wake word detection model, wherein the second wake word detection model is trained to recognize one or more wake words, and wherein detecting the wake word at the called assistant comprises inputting the utterance into the second wake word detection model;

generate, at the called assistant, a response to the utterance; and

transmit the response to a user.

19 . The device of claim 18 , wherein detecting the wake word in the utterance using the first wake word detection model is performed at a voice assistant manager; and

wherein the instructions, when executed by the processor, further cause the device to:

prior to communicating the utterance to the called assistant, activate the called assistant; and

after generating the response to the utterance, deactivate the called assistant.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: BROMAND, DANIEL; ROTH, BJÖRN ERIK; PRIEM, NICK
To: SPOTIFY AB
Reel/Frame 063177/0459 →
Continuity (1)
Related Publication 20240221728A1 · Jul 4, 2024
References Cited (34)
US 10325599B1 · Naidu · 2019 [cited by examiner]
US 10748543B2 · Mixter et al. · 2020 [cited by applicant]
US 10957329B1 · Liu et al. · 2021 [cited by applicant]
US 11062702B2 · Wood et al. · 2021 [cited by applicant]
US 12230266B1 · Henry · 2025 [cited by examiner]
US 20160104480A1 · Sharifi · 2016 [cited by examiner]
US 20170025124A1 · Mixter · 2017 [cited by examiner]
US 20170076720A1 · Gopalan · 2017 [cited by examiner]
US 20180061420A1 · Patil et al. · 2018 [cited by applicant]
US 20180108351A1 · Beckhardt · 2018 [cited by examiner]
US 20180204569A1 · Nadkar et al. · 2018 [cited by applicant]
US 20180301147A1 · Kim · 2018 [cited by examiner]
US 20190341037A1 · Bromand et al. · 2019 [cited by applicant]
US 20200184964A1 · Myers · 2020 [cited by examiner]
US 20200302932A1 · Schramm · 2020 [cited by examiner]
US 20200320997A1 · Wagatsuma et al. · 2020 [cited by applicant]
US 20200365154A1 · Sindhwani · 2020 [cited by examiner]
US 20200372907A1 · Trufinescu et al. · 2020 [cited by applicant]
US 20220180867A1 · Bobboli et al. · 2022 [cited by applicant]
US 20220189470A1 · Sharifi · 2022 [cited by examiner]
US 20220230635A1 · Schillmoeller · 2022 [cited by examiner]
US 20220293097A1 · Jekeswaran · 2022 [cited by examiner]
US 20220351724A1 · Thomas · 2022 [cited by examiner]
US 20230090019A1 · Santhar · 2023 [cited by examiner]
US 20230097197A1 · Huang · 2023 [cited by examiner]
US 20230113883A1 · Carbune · 2023 [cited by examiner]
US 20240203413A1 · Sharifi · 2024 [cited by examiner]
TW I683306B · 2020 [cited by applicant]
WO 2021114852A1 · 2021 [cited by applicant]
WO 2022067345A1 · 2022 [cited by applicant]
Connectivity Standards Alliance, https://csa-iot.org/developer-resource/matter-network-transport/, published on Jun. 9, 2022 (Year: 2022). [cited by examiner]
Tizan Docs Webpage: Multi-Assistant located at: https://docs.tizen.org/application/native/guides/text-input/multi-assistant/, obtained on Mar. 27, 2023, 7 pages. [cited by applicant]
Sonos Voice Control User Guide, online at: https://www.sonos.com/en-gb/guides/sonosvoicecontrol, obtained Mar. 27, 2023, 4 pages. [cited by applicant]
Getting Started with Universal Search and Browse on Fire TV, from Amazon webpage: https://developer.amazon.com/docs/catalog/getting-started-universal-search-and-browse.html, last updated Jan. 20, 2022, 7 pages. [cited by applicant]