IP Library Granted Patent US 11,600,265
Granted Patent B2
US 11,600,265 · App. 17/089,357 · Granted Mar 7, 2023

Systems and methods for determining whether to trigger a voice capable device based on speaking cadence

Inventors: Edison Lin (Los Altos Hills, CA); Rowena Young (Menlo Park, CA); Kanchan Sripathy (San Jose, CA); Reda Harb (Bellevue, WA)
Assignee: Rovi Guides, Inc.
G10L15/1807G10L15/22G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,600,265
App. No.
17/089,357
Granted
Mar 7, 2023
Kind
B2
Abstract

Systems and methods are described for determining whether to activate a voice activated device based on a speaking cadence of the user. When the user speaks with a first cadence the system may determine that the user does not intend to activate the device and may accordingly not to trigger a voice activated device. When the user speaks with a second cadence the system may determine that the user does wish to trigger the device and may accordingly trigger the voice activated device.

Claims (68)

1. A method for avoiding inadvertent triggering of a voice capable device, the method comprising:

detecting a voice input comprising a trigger word; and

in response to detecting the voice input comprising the trigger word:

determining a cadence of the voice input;

identifying a user associated with the voice input;

retrieving, from a profile of the user, a speaking cadence of the user;

determining whether a difference between the cadence of the voice input and the speaking cadence of the user is less than a threshold value;

in response to determining that the difference is less than the threshold value, refraining from triggering the voice capable device; and

in response to determining that the difference is not less than the threshold value, triggering the voice capable device.

2. The method of claim 1 , wherein the voice input is a first voice input, further comprising:

prompting the user to recite a sequence of utterances prior to detecting the voice input comprising the trigger word;

in response to prompting the user, receiving a second voice input from the user comprising the sequence of utterances;

in response to receiving the sequence of utterances, determining the speaking cadence of the user based on an amount of time between each utterance of the sequence of utterances; and

storing, in the profile of the user, the speaking cadence of the user.

3. The method of claim 1 , further comprising:

retrieving, from the profile of the user, demographic information corresponding to the user;

identifying, based on the demographic information corresponding to the user, a template speaking cadence; and

storing, in the profile of the user, the template speaking cadence as the speaking cadence of the user.

4. The method of claim 1 , wherein the voice input comprises a sequence of a plurality of words, further comprising:

determining a position of the trigger word in the sequence of the plurality of words; and

refraining from triggering the voice capable user device when the position of the trigger word is greater than a threshold maximum position.

5. The method of claim 4 , further comprising:

determining an average position of the trigger word based on a plurality of voice inputs from the user activating the voice capable device; and

determining the threshold maximum position based on the average position of the trigger word.

6. The method of claim 1 , wherein triggering the voice capable device comprises:

determining whether a portion of the voice input matches a word associated with a voice capable device function; and

in response to determining that the portion of the voice input matches the word associated with the voice capable device function, performing the voice capable device function.

7. The method of claim 5 , further comprising terminating the triggering of the voice capable device in response to determining that the portion of the voice input does not match the word associated with the voice capable device function.

8. The method of claim 1 , wherein identifying the user associated with the voice input, comprises:

identifying a plurality of features associated with the voice input;

comparing the plurality of features associated with the voice input to a voice fingerprint uniquely identifying features of the voice of the user; and

selecting the user in response to determining that at least a threshold number of the plurality of features associated with the voice input match the voice fingerprint uniquely identifying features of the voice of the user.

9. The method of claim 1 , further comprising triggering the voice capable device in response to determining that the difference is not less than the threshold value.

10. The method of claim 1 , further comprising detecting the trigger word based on a fingerprint associated with the trigger word.

11. A system for avoiding inadvertent triggering of a voice capable device, the system comprising control circuitry configured to:

detect a voice input comprising a trigger word; and

in response to detecting the voice input comprising the trigger word:

determine a cadence of the voice input;

identify a user associated with the voice input;

retrieve, from a profile of the user, a speaking cadence of the user;

determine whether a difference between the cadence of the voice input and the speaking cadence of the user is less than a threshold value;

in response to determining that the difference is less than the threshold value, refrain from triggering the voice capable device; and

in response to determining that the difference is not less than the threshold value, trigger the voice capable device.

12. The system of claim 11 , wherein the voice input is a first voice input, and wherein the control circuitry is further configured to:

prompt the user to recite a sequence of utterances prior to detecting the voice input comprising the trigger word;

in response to prompting the user, receive a second voice input from the user comprising the sequence of utterances;

in response to receiving the sequence of utterances, determine the speaking cadence of the user based on an amount of time between each utterance of the sequence of utterances; and

store, in the profile of the user, the speaking cadence of the user.

13. The system of claim 11 , wherein the control circuitry is further configured to:

retrieve, from the profile of the user, demographic information corresponding to the user;

identify, based on the demographic information corresponding to the user, a template speaking cadence; and

store, in the profile of the user, the template speaking cadence as the speaking cadence of the user.

14. The system of claim 11 , wherein the voice input comprises a sequence of a plurality of words, and wherein the control circuitry is further configured to:

determine a position of the trigger word in the sequence of the plurality of words; and

refrain from triggering the voice capable user device when the position of the trigger word is greater than a threshold maximum position.

15. The system of claim 14 , wherein the control circuitry is further configured to:

determine an average position of the trigger word based on a plurality of voice inputs from the user activating the voice capable device; and

determine the threshold maximum position based on the average position of the trigger word.

16. The system of claim 11 , wherein the control circuitry is further configured, when triggering the voice capable device, to:

determine whether a portion of the voice input matches a word associated with a voice capable device function; and

in response to determining that the portion of the voice input matches the word associated with the voice capable device function, perform the voice capable device function.

17. The system of claim 15 , wherein the control circuitry is further configured to terminate the triggering of the voice capable device in response to determining that the portion of the voice input does not match the word associated with the voice capable device function.

18. The system of claim 11 , wherein the control circuitry is further configured, when identifying the user associated with the voice input, to:

identify a plurality of features associated with the voice input;

compare the plurality of features associated with the voice input to a voice fingerprint uniquely identify features of the voice of the user; and

select the user in response to determining that at least a threshold number of the plurality of features associated with the voice input match the voice fingerprint uniquely identifying features of the voice of the user.

19. The system of claim 11 , wherein the control circuitry is further configured to trigger the voice capable device in response to determining that the difference is not less than the threshold value.

20. The system of claim 11 , wherein the control circuitry is further configured to detect the trigger word based on a fingerprint associated with the trigger word.

Assignments (3)
CHANGE OF NAME Recorded Oct 3, 2024
From: ROVI GUIDES, INC.
To: ADEIA GUIDES INC.
Reel/Frame 069106/0164 →
SECURITY INTEREST Recorded May 3, 2023
From: ADEIA GUIDES INC.; ADEIA IMAGING LLC; ADEIA MEDIA HOLDINGS LLC; ADEIA MEDIA SOLUTIONS INC.; ADEIA SEMICONDUCTOR ADVANCED TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR BONDING TECHNOLOGIES INC.; ADEIA SEMICONDUCTOR INC.; ADEIA SEMICONDUCTOR SOLUTIONS LLC; ADEIA SEMICONDUCTOR TECHNOLOGIES LLC; ADEIA SOLUTIONS LLC
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 063529/0272 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2020
From: LIN, EDISON; YOUNG, ROWENA; SRIPATHY, KANCHAN; HARB, REDA
To: ROVI GUIDES, INC.
Reel/Frame 054314/0162 →
Continuity (2)
Continuation 16139453 · Sep 24, 2018
Related Publication 20210125604A1 · Apr 29, 2021