IP Library › Granted Patent US 11,748,057
Granted Patent B2
US 11,748,057 · App. 17/079,111 · Granted Sep 5, 2023

System and method for personalization in intelligent multi-modal personal assistants

Inventor: Caleb Ryan Phillips (Toronto, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06F3/167G06N20/00G06V10/80G06V20/10G06V40/172G10L17/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,748,057
App. No.
17/079,111
Granted
Sep 5, 2023
Kind
B2
Abstract

A method may include receiving, by a virtual assistant of a user device, an input from a user, the virtual assistant being based on software. The method may include obtaining, by the virtual assistant of the user device and via a sensor of the user device, audio information or video information of the user. The method may include determining, by the virtual assistant of the user device, an identity of the user based on the audio information or the video information of the user and a set of facial embeddings and speech embeddings that is correlated with the user, the set of facial embeddings and speech embeddings being generated using a facial embedding model, a speech embedding model, and a sound source localization model. The method may include performing, by the virtual assistant of the user device, an action based on the input and the identity of the user.

Claims (69)

1. A method, by a user device, comprising:

receiving an input from a user;

obtaining relation information of the user, audio information of the user obtained via a microphone of the user device, and video information of the user obtained via camera of the user device;

identifying the user based on the audio information and the video information of the user and a set of facial embeddings and speech embeddings that is correlated with the user, the set of facial embeddings and speech embeddings being generated using a facial embedding model, a speech embedding model, and a sound source localization model; and

performing an action based on the input and the relation information of the user,

wherein the sound source localization model is a model that is configured to determine the video information and the audio information that belongs to a same user.

2. The method of claim 1 , further comprising:

generating, using the facial embedding model, the facial embeddings.

3. The method of claim 1 , further comprising:

generating, using the speech embedding model, the speech embeddings.

4. The method of claim 1 , further comprising:

identifying a label associated with the user; and

correlating the label with the user, based on identifying the label.

5. The method of claim 4 , further comprising:

determining a confidence score associated with the label.

6. The method of claim 1 , further comprising:

generating, using the facial embedding model, a facial embedding of the user, based on the video information of the user;

comparing the facial embedding of the user and the set of facial embeddings that is correlated with the user; and

identifying the user, based on comparing the facial embedding of the user and the set of facial embeddings that is correlated with the user.

7. The method of claim 1 , further comprising:

generating, using the speech embedding model, a speech embedding of the user, based on the audio information of the user;

comparing the speech embedding of the user and the set of speech embeddings that is correlated with the user; and

identifying the user, based on comparing the speech embedding of the user and the set of speech embeddings that is correlated with the user.

8. A user device comprising:

a memory configured to store instructions; and

a processor configured to execute the instructions to:

receive an input from a user;

obtain relation information of the user, audio information of the user obtained via a microphone of the user device, and video information of the user obtained via camera of the user device;

identify the user based on the audio information and the video information of the user and a set of facial embeddings and speech embeddings that is correlated with the user, the set of facial embeddings and speech embeddings being generated using a facial embedding model, a speech embedding model, and a sound source localization model; and

perform an action based on the input and the relation information of the user,

wherein the sound source localization model is a model that is configured to determine the video information and the audio information that belongs to a same user.

9. The user device of claim 8 , wherein the processor is further configured to:

generate, using the facial embedding model, the facial embeddings.

10. The user device of claim 8 , wherein the processor is further configured to:

generate, using the speech embedding model, the speech embeddings.

11. The user device of claim 8 , wherein the processor is further configured to:

identify a label associated with the user; and

correlate the label with the user, based on identifying the label.

12. The user device of claim 11 , wherein the processor is further configured to:

determine a confidence score associated with the label.

13. The user device of claim 8 , wherein the processor is further configured to:

generate, using the facial embedding model, a facial embedding of the user, based on the video information of the user;

compare the facial embedding of the user and the set of facial embeddings that is correlated with the user; and

identify the user, based on comparing the facial embedding of the user and the set of facial embeddings that is correlated with the user.

14. The user device of claim 8 , wherein the processor is further configured to:

generate, using the speech embedding model, a speech embedding of the user, based on the audio information of the user;

compare the speech embedding of the user and the set of speech embeddings that is correlated with the user; and

identify identity of the user, based on comparing the speech embedding of the user and the set of speech embeddings that is correlated with the user.

15. A non-transitory computer-readable medium storing instructions that, when executed, cause at least one processor of a user device to:

receive an input from a user;

obtain relation information of the user, audio information of the user obtained via a microphone of the user device, and video information of the user obtained via camera of the user device;

identify the user based on the audio information and the video information of the user and a set of facial embeddings and speech embeddings that is correlated with the user, the set of facial embeddings and speech embeddings being generated using a facial embedding model, a speech embedding model, and a sound source localization model; and

perform an action based on the input and the relation information of the user,

wherein the sound source localization model is a model that is configured to determine the video information and the audio information that belongs to a same user.

16. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

generate, using the facial embedding model, the facial embeddings.

17. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

generate, using the speech embedding model, the speech embeddings.

18. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

identify a label associated with the user; and

correlate the label with the user, based on identifying the label.

19. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

generate, using the facial embedding model, a facial embedding of the user, based on the video information of the user;

compare the facial embedding of the user and the set of facial embeddings that is correlated with the user; and

identify the user, based on comparing the facial embedding of the user and the set of facial embeddings that is correlated with the user.

20. The non-transitory computer-readable medium of claim 15 , wherein the instructions further cause the at least one processor to:

generate, using the speech embedding model, a speech embedding of the user, based on the audio information of the user;

compare the speech embedding of the user and the set of speech embeddings that is correlated with the user; and

identify identity of the user, based on comparing the speech embedding of the user and the set of speech embeddings that is correlated with the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2020
From: PHILLIPS, CALEB RYAN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 054154/0430 →
Continuity (2)
Provisional Application 62981850 · Feb 26, 2020
Related Publication 20210264134A1 · Aug 26, 2021
Cited By (1)
US 12,260,183