IP Library Granted Patent US 12,189,752
Granted Patent B2
US 12,189,752 · App. 17/571,679 · Granted Jan 7, 2025

Multimodal authentication and liveliness detection

Inventor: Yuvaraj Nagarathnam (Karnataka, IN)
Assignee: ARRIS ENTERPRISES LLC
G06F21/40G06F21/32
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,189,752
App. No.
17/571,679
Granted
Jan 7, 2025
Kind
B2
Abstract

A robust and reliable multimodal authentication is provided by a multimodal authentication device. The multimodal authentication device utilizes an audio authentication, a video authentication, an audio liveliness authentication, and a video liveliness authentication to determine the authentication of a user. By including a liveliness component in the authentication determination reduces the risk of fraud by factoring in live movement and orientation into the authentication determines. For example, various image/location combinations are displayed to the user and the user is instructed to track and verbally identify the various images. In this way, the user is authenticated based not only on, for example, facial and audio recognition but also a liveliness associated with each.

Claims (99)

1. A multimodal authentication device to provide a multimodal authentication, comprising:

a memory storing one or more computer-readable instructions; and

a processor configured to execute the one or more computer-readable instructions to perform one or more operations to:

receive, from a user, a request for access to any of one or more devices, one or more services, one or more resources associated with the multimodal authentication device, or any combination thereof;

display one or more images at one or more locations of a display device associated with the multimodal authentication device;

receive one or more user images of the user in response to the display of the one or more images, wherein the user image is received from an image capture device associated with the multimodal authentication device;

receive one or more user audio inputs in response to the display of the one or more images, wherein the user audio is received from an audio capture device;

determine a visual authentication based on the one or more user images;

determine a visual liveliness authentication based on a tracking of the one or more user images to the display of the one or more images;

determine an audio authentication based on the one or more user audio inputs;

determine an audio liveliness authentication based on a comparison of a first timestamp associated with one or more lip movements of the user from the one or more user images and a second timestamp associated with a start of a command from the one or more user audio inputs; and

provide a multimodal authentication based on the visual authentication, the visual liveliness authentication, the audio authentication, and the audio liveliness authentication so that the user has the requested access.

2. The multimodal authentication device of claim 1 , wherein the processor is further configured to execute the one or more computer-readable instructions to perform one or more further operations to:

generate a location random number series, wherein the one or more locations are based on the location random number series; and

generate an image random number series, wherein the one or more images are based on the image random number series.

3. The multimodal authentication device of claim 1 , wherein the processor is further configured to execute the one or more computer-readable instructions to perform one or more further operations to:

set a visual authentication result based on the visual authentication;

set a visual liveliness authentication result based on the visual liveliness authentication;

set an audio authentication result based on the audio authentication; and

set an audio liveliness authentication result based on the audio liveliness authentication.

4. The multimodal authentication device of claim 3 , wherein the processor is further configured to execute the one or more computer-readable instructions to perform one or more further operations to:

compare the visual authentication result to a visual authentication result threshold;

compare the visual liveliness authentication result to a visual liveliness authentication threshold;

compare the audio authentication result to an audio authentication result threshold; and

compare the audio liveliness authentication result to an audio liveliness authentication result threshold, wherein providing the multimodal authentication is further based on each of the comparisons.

5. The multimodal authentication device of claim 1 , wherein the determining the visual liveliness authentication comprises at least one of:

determining a face angle associated with the one or more user images; and

determining a gaze angle associated with the one or more user images.

6. The multimodal authentication device of claim 5 , wherein the determining the visual liveliness authentication further comprises at least one of:

determining that the face angle tracks the displaying of the one or more images at each of the one or more locations; and

determining that the gaze angle tracks the displaying of the one or more images at each of the one or more locations.

7. The multimodal authentication device of claim 1 , wherein the processor is further configured to execute the one or more computer-readable instructions to perform one or more further operations to:

provide one or more instructions at the display device;

receive a user input in response to the displaying the one or more instructions; and

wherein the displaying the one or more images is based on the user input.

8. A method for a multimodal authentication device to provide a multimodal authentication, the method comprising:

receiving, from a user, a request for access to any of one or more devices, one or more services, one or more resources associated with the multimodal authentication device, or any combination thereof;

displaying one or more images at one or more locations of a display device associated with the multimodal authentication device;

receiving one or more user images of the user in response to the display of the one or more images, wherein the user image is received from an image capture device associated with the multimodal authentication device;

receiving one or more user audio inputs in response to the display of the one or more images, wherein the user audio is received from an audio capture device;

determining a visual authentication based on the one or more user images;

determining a visual liveliness authentication based on a tracking of the one or more user images to the display of the one or more images;

determining an audio authentication based on the one or more user audio inputs;

determining an audio liveliness authentication based on a comparison of a first timestamp associated with one or more lip movements of the user from the one or more user images and a second timestamp associated with a start of a command from the one or more user audio inputs; and

providing a multimodal authentication based on the visual authentication, the visual liveliness authentication, the audio authentication, and the audio liveliness authentication so that the user has the requested access.

9. The method of claim 8 , further comprising:

generating a location random number series, wherein the one or more locations are based on the location random number series; and

generating an image random number series, wherein the one or more images are based on the image random number series.

10. The method of claim 8 , further comprising:

setting a visual authentication result based on the visual authentication;

setting a visual liveliness authentication result based on the visual liveliness authentication;

setting an audio authentication result based on the audio authentication; and

setting an audio liveliness authentication result based on the audio liveliness authentication.

11. The method of claim 10 , wherein providing the multimodal authentication comprises:

comparing the visual authentication result to a visual authentication result threshold;

comparing the visual liveliness authentication result to a visual liveliness authentication threshold;

comparing the audio authentication result to an audio authentication result threshold; and

comparing the audio liveliness authentication result to an audio liveliness authentication result threshold.

12. The method of claim 8 , wherein determining the visual liveliness authentication comprises at least one of:

determining a face angle associated with the one or more user images; and

determining a gaze angle associated with the one or more user images.

13. The method of claim 12 , wherein determining the visual liveliness authentication further comprises at least one of:

determining that the face angle tracks the displaying of the one or more images at each of the one or more locations; and

determining that the gaze angle tracks the displaying of the one or more images at each of the one or more locations.

14. The method of claim 8 , further comprising:

providing one or more instructions at the display device;

receiving a user input in response to the displaying the one or more instructions; and

wherein the displaying the one or more images is based on the user input.

15. A non-transitory computer-readable medium of a multimodal authentication device storing one or more instructions for providing a multimodal authentication, which when executed by a processor of the multimodal authentication device, cause the multimodal authentication device to perform one or more operations comprising:

receive, from a user, a request for access to any of one or more devices, one or more services, one or more resources associated with the multimodal authentication device, or any combination thereof;

displaying one or more images at one or more locations of a display device associated with the multimodal authentication device;

receiving one or more user images of the user in response to the display of the one or more images, wherein the user image is received from an image capture device associated with the multimodal authentication device;

receiving one or more user audio inputs in response to the display of the one or more images, wherein the user audio is received from an audio capture device;

determining a visual authentication based on the one or more user images;

determining a visual liveliness authentication based on a tracking of the one or more user images to the display of the one or more images;

determining an audio authentication based on the one or more user audio inputs;

determining an audio liveliness authentication based on a comparison of a first timestamp associated with one or more lip movements of the user from the one or more user images and a second timestamp associated with a start of a command from the one or more user audio inputs; and

providing a multimodal authentication based on the visual authentication, the visual liveliness authentication, the audio authentication, and the audio liveliness authentication so that the user has the requested access.

16. The non-transitory computer-readable medium of claim 15 , the one or more instructions when executed by the processor further cause the multimodal authentication device to perform one or more further operations comprising:

generating a location random number series, wherein the one or more locations are based on the location random number series; and

generating an image random number series, wherein the one or more images are based on the image random number series.

17. The non-transitory computer-readable medium of claim 15 , the one or more instructions when executed by the processor further cause the multimodal authentication device to perform one or more further operations comprising:

setting a visual authentication result based on the visual authentication;

setting a visual liveliness authentication result based on the visual liveliness authentication;

setting an audio authentication result based on the audio authentication; and

setting an audio liveliness authentication result based on the audio liveliness authentication.

18. The non-transitory computer-readable medium of claim 17 , the one or more instructions when executed by the processor further cause the multimodal authentication device to perform one or more further operations comprising:

comparing the visual authentication result to a visual authentication result threshold;

comparing the visual liveliness authentication result to a visual liveliness authentication threshold;

comparing the audio authentication result to an audio authentication result threshold; and

comparing the audio liveliness authentication result to an audio liveliness authentication result threshold, wherein providing the multimodal authentication is further based on each of the comparisons.

19. The non-transitory computer-readable medium of claim 17 , the one or more instructions when executed by the processor further cause the multimodal authentication device to perform one or more further operations comprising:

providing one or more instructions at the display device;

receiving a user input in response to the displaying the one or more instructions; and

wherein the displaying the one or more images is based on the user input.

20. The non-transitory computer-readable medium of claim 15 , wherein determining the visual liveliness authentication comprises at least one of:

determining a face angle associated with the one or more user images;

determining a gaze angle associated with the one or more user images; and

determining that the face angle and the gaze angle tracks the displaying of the one or more images at each of the one or more locations.

Assignments (8)
RELEASE OF SECURITY INTEREST AT REEL/FRAME 059350/0743 Recorded Jan 12, 2026
From: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
To: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE NORTH CAROLINA, LLC (F/K/A COMMSCOPE, INC. OF NORTH CAROLINA)
Reel/Frame 074594/0156 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS AT REEL/FRAME NO. 59710/0506 Recorded Jan 9, 2026
From: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
To: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE NORTH CAROLINA, LLC (F/K/A COMMSCOPE, INC. OF NORTH CAROLINA)
Reel/Frame 074282/0522 →
RELEASE OF SECURITY INTEREST AT REEL/FRAME 059350/0921 Recorded Dec 19, 2024
From: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
To: ARRIS ENTERPRISES LLC (F/K/A ARRIS ENTERPRISES, INC.); COMMSCOPE, INC. OF NORTH CAROLINA; COMMSCOPE TECHNOLOGIES LLC
Reel/Frame 069743/0704 →
SECURITY INTEREST Recorded Dec 17, 2024
From: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE INC., OF NORTH CAROLINA; OUTDOOR WIRELESS NETWORKS LLC; RUCKUS IP HOLDINGS LLC
To: APOLLO ADMINISTRATIVE AGENCY LLC
Reel/Frame 069889/0114 →
SECURITY INTEREST Recorded Mar 9, 2022
From: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE, INC. OF NORTH CAROLINA
To: WILMINGTON TRUST
Reel/Frame 059710/0506 →
ABL SECURITY AGREEMENT Recorded Mar 8, 2022
From: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE, INC. OF NORTH CAROLINA
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 059350/0743 →
TERM LOAN SECURITY AGREEMENT Recorded Mar 8, 2022
From: ARRIS ENTERPRISES LLC; COMMSCOPE TECHNOLOGIES LLC; COMMSCOPE, INC. OF NORTH CAROLINA
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 059350/0921 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2022
From: NAGARATHNAM, YUVARAJ
To: ARRIS ENTERPRISES LLC
Reel/Frame 058595/0148 →
Continuity (2)
Provisional Application 63170062 · Apr 2, 2021
Related Publication 20220318362A1 · Oct 6, 2022
References Cited (10)
US 8856541B1 · Chaudhury · 2014 [cited by examiner]
US 9430629B1 · Ziraknejad et al. · 2016 [cited by applicant]
US 11256792B2 · Tussy · 2022 [cited by examiner]
US 20180232591A1 · Hicks · 2018 [cited by examiner]
US 20190080065A1 · Sheik-Nainar · 2019 [cited by examiner]
US 20190205680A1 · Miu · 2019 [cited by examiner]
US 20200097643A1 · Uzun · 2020 [cited by examiner]
US 20200134148A1 · Mortazavian et al. · 2020 [cited by applicant]
International Preliminary Report on Patentability and Written Opinion issued Oct. 12, 2023 in International Application No. PCT/US2022/011750. [cited by applicant]
International Search Report and the Written Opinion of the International Searching Authority dated Apr. 8, 2022 in International (PCT) Application No. PCT/US2022/011750. [cited by applicant]
Cited By (1)
US 12,645,777