IP Library Granted Patent US 11,514,920
Granted Patent B2
US 11,514,920 · App. 17/322,848 · Granted Nov 29, 2022

Method and system for determining speaker-user of voice-controllable device

Inventor: Ivan Aleksandrovich Karpukhin (Elektrostal, RU)
Assignee: YANDEX EUROPE AG
G10L17/00G06N20/00G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,920
App. No.
17/322,848
Granted
Nov 29, 2022
Kind
B2
Abstract

There are disclosed methods and systems for determining a speaker of a set of registered users associated with a voice-controllable device. The method is executable by an electronic device configured to execute a Machine Learning Algorithm (MLA). The method comprises executing the MLA to determine a first probability parameter indicative of the speaker of the user utterance being one of the set of registered users; executing a user frequency analysis to generate, for each given one of the set of registered users, a second probability parameter the being an apriori frequency based probability; generating, for the electronic device, for each given one of the set of registered users an amalgamated probability based on the first probability and the second probability associated therewith; selecting the given one of the set of registered users as the speaker of the user utterance based on the amalgamated probability value.

Claims (38)

1. A method of determining a speaker from a set of registered users associated with a voice-controllable device, the method executable by an electronic device configured to execute a Machine Learning Algorithm (MLA), the method comprising:

receiving an indication of a user utterance, wherein the user utterance was produced by the speaker;

executing the MLA to determine, for each registered user of the set of registered users, a first probability parameter indicating a predicted likelihood that the user utterance was produced by the respective registered user;

determining, for each registered user of the set of registered users, a second probability parameter indicating a frequency at which the respective registered user has interacted with the voice-controllable device;

generating, for each registered user of the set of registered users, an amalgamated probability value based on the first probability parameter and the second probability parameter associated with the respective registered user; and

selecting, based on the amalgamated probability values, one registered user of the set of registered users as the speaker.

2. The method of claim 1 , wherein determining the second probability parameter comprises executing a user frequency analysis.

3. The method of claim 2 , wherein the user frequency analysis weighs a sub-set of recorded frequency data, the sub-set including a pre-determined amount of recorded frequency data.

4. The method of claim 1 , further comprising updating, based on the selecting the speaker, an apriori frequency based probability associated with each registered user of the set of registered users.

5. The method of claim 1 , further comprising retrieving a user profile associated with the speaker and providing the speaker with a set of authorized voice-based actions.

6. The method of claim 1 , further comprising, before selecting the one registered user as the speaker, determining that the amalgamated probability value of the one registered user is greater than a pre-determined threshold value.

7. The method of claim 1 , wherein the indication of the user utterance comprises a plurality of parameters of the user utterance.

8. A method of determining a speaker from a set of registered users associated with a voice-controllable device, the method executable by the voice-controllable device, the method comprising:

receiving, by the voice-controllable device, an indication of a user utterance, wherein the user utterance was produced by the speaker;

executing, by the voice-controllable device, a Machine Learning Algorithm (MLA) to determine a first probability parameter indicative of the speaker of the user utterance being one of the set of registered users;

determining, by the voice-controllable device and for each registered user of the set of registered users, a second probability parameter indicating a frequency at which the respective registered user has interacted with the voice-controllable device;

generating, by the voice-controllable device and for each registered user of the set of registered users, an amalgamated probability value based on the first probability parameter and the second probability parameter associated with the respective registered user; and

selecting, by the voice-controllable device and based on the amalgamated probability values, one registered user of the set of registered users as the speaker.

9. The method of claim 8 , wherein determining the second probability parameter comprises executing a user frequency analysis.

10. The method of claim 9 , wherein the user frequency analysis weighs a sub-set of recorded frequency data, the sub-set including a pre-determined amount of recorded frequency data.

11. The method of claim 8 , further comprising updating, based on the selecting the speaker, an apriori frequency based probability associated with each registered user of the set of registered users.

12. The method of claim 8 , further comprising retrieving a user profile associated with the speaker and providing the speaker with a set of authorized voice-based actions.

13. The method of claim 8 , further comprising, before selecting the one registered user as the speaker, determining that the amalgamated probability value of the one registered user is greater than a pre-determined threshold value.

14. A system comprising a voice-controllable device and a server, wherein the voice-controllable device comprises at least one processor and memory storing a plurality of executable instructions which, when executed by the at least one processor of the voice-controllable device, cause the voice-controllable device to:

receive an indication of a user utterance, wherein the user utterance was produced by a speaker; and

send the indication of the user utterance to the server, and

wherein the server comprises at least one processor and memory storing a plurality of executable instructions which, when executed by the at least one processor of the server, cause the server to:

receive the indication of the user utterance;

execute a Machine Learning Algorithm (MLA) to determine a first probability parameter indicative of the speaker being one of a set of registered users;

determine, for each registered user of the set of registered users, a second probability parameter indicating a frequency at which the respective registered user has interacted with the voice-controllable device;

generate, for each registered user of the set of registered users, an amalgamated probability based on the first probability parameter and the second probability parameter associated with the respective registered user; and

after determining that each amalgamated probability is below a pre-determined threshold, select a guest user as the speaker.

15. The system of claim 14 , wherein the instructions, when executed by the at least one processor of the server, cause the server to, based on selecting the guest user as the speaker, update an apriori frequency based probability associated with the guest user.

16. The system of claim 14 , wherein the instructions, when executed by the at least one processor of the server, cause the server to retrieve a user profile associated with the guest user and provide a set of authorized voice-based actions for the guest user, wherein the set of voice-based actions associated with the guest user is smaller than a set of voice-based actions associated with the set of registered users.

17. The system of claim 14 , wherein the instructions, when executed by the at least one processor of the server, cause the server to, in response to a pre-determined number of past rendered determined identities of speakers being the guest speaker, execute a pre-determined guest scenario.

18. The system of claim 17 , wherein the pre-determined guest scenario comprises, during a future execution of the MLA, artificially decreasing an amount of time spent during generation of the first probability parameter.

19. The system of claim 14 , wherein the instructions that cause the server to determine the second probability parameter comprises instructions that, when executed by the at least one processor of the server, cause the server to execute a user frequency analysis.

20. The system of claim 19 , wherein the user frequency analysis weighs a sub-set of recorded frequency data, the sub-set including a pre-determined amount of recorded frequency data.

Assignments (6)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2024
From: DIRECT CURSUS TECHNOLOGY L.L.C
To: Y.E. HUB ARMENIA LLC
Reel/Frame 068534/0384 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2023
From: YANDEX EUROPE AG
To: DIRECT CURSUS TECHNOLOGY L.L.C
Reel/Frame 065692/0720 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: KARPUKHIN, IVAN ALEKSANDROVICH
To: YANDEX.TECHNOLOGIES LLC
Reel/Frame 056265/0955 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: YANDEX.TECHNOLOGIES LLC
To: YANDEX LLC
Reel/Frame 056265/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: YANDEX LLC
To: YANDEX EUROPE AG
Reel/Frame 056265/0971 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2021
From: KARPUKHIN, IVAN ALEKSANDROVICH
To: YANDEX.TECHNOLOGIES LLC
Reel/Frame 056265/0988 →