IP Library Granted Patent US 12,579,972
Granted Patent B2
US 12,579,972 · App. 18/645,179 · Granted Mar 17, 2026

Method and apparatus for intelligent voice recognition

Inventors: Navdeep Jain (Philadelphia, PA); Hongcheng Wang (Arlington, VA)
Assignee: Comcast Cable Communications, LLC
G10L15/10G10L15/16G10L15/22G10L25/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,972
App. No.
18/645,179
Granted
Mar 17, 2026
Kind
B2
Abstract

Methods and systems are described for recognizing, based on a voice input, a user and/or a voice command. An algorithm is described herein that processes data associated with a voice input. The data may indicate characteristics of the voice such as a gender, an age, or accent associated with the voice and other metadata. For example, the system may process the data and determine the gender of a voice. The determined characteristics may be used as an input into a voice recognition engine to improve the accuracy of identifying the user who spoke the voice input and identifying a voice command associated with the voice input. For example, the determined gender may be used as a parameter to improve the accuracy of an identified user (e.g., the speaker) or command. The algorithm may adjust, based on gender, parameters such as confidence thresholds used to match voices and voice commands.

Claims (59)

1 . A device comprising:

one or more processors; and

memory storing instructions that, when executed by the one or more processors, cause the device to:

determine, based on data associated with a voice input, a gender associated with the voice input;

determine, based on the gender being associated with a female voice, a first confidence threshold;

determine, based on the voice input, a confidence value that the voice input was received from a user;

determine, based on the confidence value satisfying the first confidence threshold, to send an indication that the voice input was received from the user; and

send, based on determining to send the indication, the indication.

2 . The device of claim 1 , wherein the instructions, when executed by the one or more processors, further cause the device to:

determine, based on second data associated with a second voice input, a second gender associated with the second voice input;

determine, based on the second gender being associated with a male voice, a second confidence threshold, wherein the first confidence threshold is different than the second confidence threshold;

determine, based on the second voice input, a second confidence value that the second voice input was received from a second user;

determine, based on the second confidence value satisfying the second confidence threshold, to send a second indication that the second voice input was received from the second user; and

send, based on determining to send the second indication, the second indication.

3 . The device of claim 1 , wherein the first confidence threshold indicates a degree of accuracy needed to match the data with the user.

4 . The device of claim 1 , wherein the first confidence threshold minimizes at least one of: a false acceptance rate (FAR) associated with determining that the voice input was received from the user, or a false rejection rate (FRR) associated with determining that the voice input was received from the user.

5 . The device of claim 1 , wherein the first confidence threshold causes an equal error rate for a false acceptance rate (FAR) associated with determining that the voice input was received from the user and a false rejection rate (FRR) associated with determining that the voice input was received from the user.

6 . The device of claim 1 , wherein the determining the confidence value that the voice input was received from a user is further based on a comparison to stored information, wherein the stored information comprises an implicit user fingerprint comprising one or more features associated with the user.

7 . The device of claim 1 , wherein the instructions, when executed by the one or more processors, further cause the device to:

generate, based on a second voice input, an implicit user fingerprint comprising one or more features associated with the user, wherein the second voice input comprises at least one of: a previous voice input or an enrollment command comprising one or more words.

8 . The device of claim 1 , wherein the instructions, when executed by the one or more processors, further cause the device to:

determine, based at least in part on the first confidence threshold, a command.

9 . A non-transitory computer-readable medium storing instructions that, when executed, cause:

determining, based on data associated with a voice input, a gender associated with the voice input;

determining, based on the gender being associated with a female voice, a first confidence threshold;

determining, based on the voice input, a confidence value that the voice input was received from a user;

determining, based on the confidence value satisfying the first confidence threshold, to send an indication that the voice input was received from the user; and

sending, based on determining to send the indication, the indication.

10 . The non-transitory computer-readable medium of claim 9 , wherein the instructions, when executed, further cause:

determining, based on second data associated with a second voice input, a second gender associated with the second voice input;

determining, based on the second gender being associated with a male voice, a second confidence threshold, wherein the first confidence threshold is different than the second confidence threshold;

determining, based on the second voice input, a second confidence value that the second voice input was received from a second user;

determining, based on the second confidence value satisfying the second confidence threshold, to send a second indication that the second voice input was received from the second user; and

sending, based on determining to send the second indication, the second indication.

11 . The non-transitory computer-readable medium of claim 9 , wherein the first confidence threshold indicates a degree of accuracy needed to match the data with the user.

12 . The non-transitory computer-readable medium of claim 9 , wherein the first confidence threshold minimizes at least one of: a false acceptance rate (FAR) associated with determining that the voice input was received from the user, or a false rejection rate (FRR) associated with determining that the voice input was received from the user.

13 . The non-transitory computer-readable medium of claim 9 , wherein the first confidence threshold causes an equal error rate for a false acceptance rate (FAR) associated with determining that the voice input was received from the user and a false rejection rate (FRR) associated with determining that the voice input was received from the user.

14 . The non-transitory computer-readable medium of claim 9 , wherein the determining the confidence value that the voice input was received from a user is further based on a comparison to stored information, wherein the stored information comprises an implicit user fingerprint comprising one or more features associated with the user.

15 . The non-transitory computer-readable medium of claim 9 , wherein the instructions, when executed, further cause:

generating, based on a second voice input, an implicit user fingerprint comprising one or more features associated with the user, wherein the second voice input comprises at least one of: a previous voice input or an enrollment command comprising one or more words.

16 . The non-transitory computer-readable medium of claim 9 , wherein the instructions, when executed by the one or more processors, further cause the device to:

determining, based at least in part on the first confidence threshold, a command.

17 . A system comprising:

a first computing device configured to:

determine, based on data associated with a voice input, a gender associated with the voice input,

determine, based on the gender being associated with a female voice, a first confidence threshold,

determine, based on the voice input, a confidence value that the voice input was received from a user,

determine, based on the confidence value satisfying the first confidence threshold, to send an indication that the voice input was received from the user, and

send, based on determining to send the indication, the indication; and

a second computing device configured to:

receive the indication.

18 . The system of claim 17 , wherein the first computing device is further configured to:

determine, based on second data associated with a second voice input, a second gender associated with the second voice input;

determine, based on the second gender being associated with a male voice, a second confidence threshold, wherein the first confidence threshold is different than the second confidence threshold;

determine, based on the second voice input, a second confidence value that the second voice input was received from a second user;

determine, based on the second confidence value satisfying the second confidence threshold, to send a second indication that the second voice input was received from the second user; and

send, based on determining to send the second indication, the second indication.

19 . The system of claim 17 , wherein the first confidence threshold minimizes at least one of: a false acceptance rate (FAR) associated with determining that the voice input was received from the user, or a false rejection rate (FRR) associated with determining that the voice input was received from the user.

20 . The system of claim 17 , wherein the first confidence threshold causes an equal error rate for a false acceptance rate (FAR) associated with determining that the voice input was received from the user and a false rejection rate (FRR) associated with determining that the voice input was received from the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2024
From: JAIN, NAVDEEP; WANG, HONGCHENG
To: COMCAST CABLE COMMUNICATIONS, LLC
Reel/Frame 067967/0055 →
Continuity (2)
Continuation 17302386 · Apr 30, 2021
Related Publication 20240347046A1 · Oct 17, 2024
References Cited (36)
US 5953701A · Neti et al. · 1999 [cited by applicant]
US 6671669B1 · Garudadri et al. · 2003 [cited by applicant]
US 7401017B2 · Murveit et al. · 2008 [cited by applicant]
US 7475015B2 · Epstein et al. · 2009 [cited by applicant]
US 7502736B2 · Hong et al. · 2009 [cited by applicant]
US 7529665B2 · Kim et al. · 2009 [cited by applicant]
US 7865368B2 · Li-Chun Wang et al. · 2011 [cited by applicant]
US 7949526B2 · Ju et al. · 2011 [cited by applicant]
US 8010358B2 · Chen · 2011 [cited by applicant]
US 8296383B2 · Lindahl · 2012 [cited by applicant]
US 8831942B1 · Nucci · 2014 [cited by examiner]
US 8965764B2 · Ryu et al. · 2015 [cited by applicant]
US 9262612B2 · Cheyer · 2016 [cited by applicant]
US 9484030B1 · Meaney et al. · 2016 [cited by applicant]
US 9633660B2 · Haughay · 2017 [cited by applicant]
US 10013985B2 · Yue et al. · 2018 [cited by applicant]
US 10121471B2 · Hoffmeister et al. · 2018 [cited by applicant]
US 10685658B2 · Li et al. · 2020 [cited by applicant]
US 20030110038A1 · Sharma et al. · 2003 [cited by applicant]
US 20080195387A1 · Zigel et al. · 2008 [cited by applicant]
US 20090119103A1 · Gerl et al. · 2009 [cited by applicant]
US 20110153317A1 · Mao et al. · 2011 [cited by applicant]
US 20120209609A1 · Zhao et al. · 2012 [cited by applicant]
US 20130173267A1 · Washio · 2013 [cited by applicant]
US 20140172428A1 · Han · 2014 [cited by applicant]
US 20170270919A1 · Parthasarathi et al. · 2017 [cited by applicant]
US 20200045130A1 · Rastrow et al. · 2020 [cited by applicant]
CN 105513597A · 2016 [cited by applicant]
EP 2048656B1 · 2010 [cited by examiner]
EP 2216775A1 · 2010 [cited by applicant]
EP 1904347B1 · 2011 [cited by applicant]
EP 3477505A1 · 2019 [cited by applicant]
JP 6705008B2 · 2020 [cited by applicant]
KR 1020200012963A · 2020 [cited by applicant]
WO 2019022722A1 · 2019 [cited by applicant]
US Patent Application filed on Apr. 30, 2021, entitled “Method and Apparatus for Intelligent Voice Recognition”, U.S. Appl. No. 17/302,386. [cited by applicant]