IP Library Granted Patent US 11,183,181
Granted Patent B2
US 11,183,181 · App. 15/936,177 · Granted Nov 23, 2021

Systems and methods of multiple voice services

Inventors: Klaus Hartung (Santa Barbara, CA); Daniele Giacobello (Los Angeles, CA)
Assignee: Sonos, Inc.
G10L15/22G06F3/167G10L15/08G10L15/30G10L25/51G10L15/14G10L15/32G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,183,181
App. No.
15/936,177
Granted
Nov 23, 2021
Kind
B2
Abstract

Disclosed herein are example techniques to identify a voice service to process a voice input. An example implementation may involve a network microphone device (NMD) receiving, via a microphone, voice data indicating a voice input. The NMD may identify, from among multiple voice services registered to a media playback system, a voice service to process the voice input and cause, via a network interface, the identified voice service to process the voice input.

Claims (63)

1. A network microphone device comprising:

one or more microphones;

a network interface;

one or more processors;

memory comprising tangible, non-transitory computer-readable media storing instructions executable by the one or more processors to cause the network microphone device to perform operations comprising:

receiving, via the one or more microphones, voice data indicating a voice input, wherein the received voice data includes a first portion representing an activation word corresponding to one of a plurality of voice services and a second portion representing a voice command, wherein the plurality of voice services are externally registered to a media playback system associated with the networked microphone device;

identifying, prior to performing speech recognition on the second portion of the received voice data representing the voice command, from among the plurality of voice services, a voice service to process the voice input, wherein the identifying comprises (i) determining a closest match of the first portion of the received voice data representing the activation word with corresponding activation word data stored in a recognition dataset on the network microphone device, (ii) determining a confidence score of the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data, and (iii) comparing the confidence score with a predetermined threshold score, wherein the predetermined threshold score has a first value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a first voice service, and the predetermined threshold score has a second, different value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a second voice service;

selecting, based on the determined closest match, the identified voice service and foregoing selection of another voice service;

transmitting, via the network interface, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the selected voice service only if the confidence score is greater than or equal to the predetermined threshold score;

receiving, from the identified voice service, an indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service; and

in response to the received indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service, updating the activation word data in the recognition dataset.

2. The network microphone device of claim 1 , wherein the network microphone device is a first device of the media playback system, and wherein the instructions stored on the memory further include instructions for:

transmitting, via the network interface, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to a second device of the media playback system, wherein the second device is configured to further analyze the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command if the confidence score is less than the predetermined threshold score.

3. The network microphone device of claim 2 , wherein the confidence score is a first confidence score, and wherein the instructions stored on the memory further include instructions for:

receiving, via the network interface from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

transmitting, via the network interface, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the identified voice service only if the second confidence score is greater than or equal to the predetermined threshold score.

4. The network microphone device of claim 2 , further comprising:

a transducer configured to output audio,

wherein the confidence score is a first confidence score, and wherein the instructions stored on the memory further include instructions for:

receiving, via the network interface from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

outputting, via the transducer, a request for additional user voice input if the second confidence score is less than the predetermined threshold score.

5. The network microphone device of claim 1 , wherein updating the activation word data in the recognition dataset comprises adjusting the predetermined threshold score.

6. The network microphone device of claim 1 , wherein updating the activation word data in the recognition dataset comprises adjusting the first value of the predetermined threshold score.

7. A tangible, non-transitory computer-readable medium storing instructions that, when executed by one or more processors of a network microphone device, cause the network microphone device to perform operations comprising:

receiving, via one or more microphones of the network microphone device, voice data indicating a voice input, wherein the received voice data includes a first portion representing an activation word corresponding to one of a plurality of voice services and a second portion representing a voice command, wherein the plurality of voice services are externally registered to a media playback system associated with the networked microphone device;

identifying, prior to performing speech recognition on the second portion of the received voice data representing the voice command, from among the plurality of voice services, a voice service to process the voice input, wherein the identifying comprises (i) determining a closest match of the first portion of the received voice data representing the activation word with corresponding activation word data stored in a recognition dataset on the network microphone device, (ii) determining a confidence score of the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data, and (iii) comparing the confidence score with a predetermined threshold score, wherein the predetermined threshold score has a first value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a first voice service, and the predetermined threshold score has a second, different value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a second voice service;

selecting, based on the determined closest match, the identified voice service and foregoing selection of another voice service;

transmitting, via a network interface of the network microphone device, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the selected voice service only if the confidence score is greater than or equal to the predetermined threshold score;

receiving, from the identified voice service, an indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service; and

in response to the received indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service, updating the activation word data in the recognition dataset.

8. The tangible, non-transitory computer-readable medium of claim 7 , wherein the network microphone device is a first device of a plurality of devices in a media playback system, the instructions further including instructions for:

transmitting, via the network interface, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to a second device in the media playback system, wherein the second device is configured to further analyze the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command if the confidence score is less than the predetermined threshold score.

9. The tangible, non-transitory computer-readable medium of claim 8 , wherein the confidence score is a first confidence score, the instructions further including instructions for:

receiving, via the network interface from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

transmitting, via the network interface, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the identified voice service only if the second confidence score is greater than or equal to the predetermined threshold score.

10. The tangible, non-transitory computer-readable medium of claim 8 , wherein the confidence score is a first confidence score, the instructions further including instructions for:

receiving, via the network interface from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

outputting, via a transducer of the network microphone device, a request for additional user voice input if the second confidence score is less than the predetermined threshold score.

11. The tangible, non-transitory computer-readable medium of claim 7 , wherein updating the recognition data set comprises adjusting the predetermined threshold score.

12. The tangible, non-transitory computer-readable medium of claim 7 , wherein updating the activation word data in the recognition dataset comprises adjusting the first value of the predetermined threshold score.

13. A method of operating a network microphone device, the method comprising:

receiving, via one or more microphones of the network microphone device, voice data indicating a voice input, wherein the received voice data includes a first portion representing an activation word corresponding to one of a plurality of voice services and a second portion representing a voice command, wherein the plurality of voice services are externally registered to a media playback system associated with the networked microphone device;

identifying, prior to performing speech recognition on the second portion of the received voice data representing the voice command, by the network microphone device from among the plurality of voice services, a voice service to process the voice input, wherein the identifying comprises (i) determining a closest match of the first portion of the received voice data representing the activation word with corresponding activation word data stored in a recognition dataset on the network microphone device, (ii) determining a confidence score of the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data, and (iii) comparing the confidence score with a predetermined threshold score, wherein the predetermined threshold score has a first value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a first voice service, and the predetermined threshold score has a second, different value if the closest match of the first portion of the received voice data representing the activation word with the corresponding activation word data is associated with a second voice service;

selecting, based on the determined closest match, the identified voice service and foregoing selection of another voice service;

transmitting, via a network interface of the network microphone device, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the selected voice service only if the confidence score is greater than or equal to the predetermined threshold score;

receiving, from the identified voice service, an indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service; and

in response to the received indication of whether the first portion of the received voice data representing the activation word was recognized by the identified voice service, updating the activation word data in the recognition dataset.

14. The method of claim 13 , wherein the network microphone device is a first device of the media playback system, and wherein the method further comprises:

transmitting, via the network interface of the network microphone device, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to a second device of the media playback system, wherein the second device is configured to further analyze the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command if the confidence score is less than the predetermined threshold score.

15. The method of claim 14 , wherein the confidence score is a first confidence score, and wherein the method further comprises:

receiving, via the network interface of the network microphone device from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

transmitting, via the network interface of the network microphone device, the first portion of the received voice data representing the activation word and the second portion of the received voice data representing the voice command to the identified voice service only if the second confidence score is greater than or equal to the predetermined threshold score.

16. The method of claim 14 , wherein the confidence score is a first confidence score, and wherein the method further comprises:

receiving, via the network interface of the network microphone device from the second device, an indication of a second confidence score, wherein the second confidence score is greater than the first confidence score;

comparing the second confidence score with the predetermined threshold score; and

outputting, via a transducer of the network microphone device, a request for additional user voice input if the second confidence score is less than the predetermined threshold score.

17. The method of claim 13 , wherein updating the recognition data set comprises adjusting the predetermined threshold score.

18. The method of claim 13 , wherein updating the activation word data in the recognition dataset comprises adjusting the first value of the predetermined threshold score.

Assignments (4)
RELEASE OF SECURITY INTEREST Recorded Oct 18, 2021
From: JPMORGAN CHASE BANK, N.A.
To: SONOS, INC.
Reel/Frame 058213/0597 →
SECURITY AGREEMENT Recorded Oct 15, 2021
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 058123/0206 →
SECURITY INTEREST Recorded Aug 30, 2018
From: SONOS, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 046991/0433 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 2, 2018
From: HARTUNG, KLAUS; GIACOBELLO, DANIELE
To: SONOS, INC.
Reel/Frame 045414/0568 →
Continuity (2)
Provisional Application 62477403 · Mar 27, 2017
Related Publication 20180277113A1 · Sep 27, 2018
Cited By (1)
US 12,431,125