IP Library › Granted Patent US 10,146,923
Granted Patent B2
US 10,146,923 · App. 15/559,793 · Granted Dec 4, 2018

Audiovisual associative authentication method, related system and device

Inventors: Martti Pitkänen (Helsinki, FI); Robert Parts (Saue, EE); Pirjo Huuhka-Pitkänen (Helsinki, FI)
Assignee: APLComp Oy
G06F21/32G10L17/22H04L63/0861H04W12/06G06Q20/3674G07C9/00158
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,146,923
App. No.
15/559,793
Granted
Dec 4, 2018
Kind
B2
Abstract

Electronic system for authenticating a user of an electronic service, said system preferably comprising at least one server apparatus, the system being configured to store, for a number of users, a plurality of personal voice-prints each of which being linked with a dedicated visual, audiovisual or audio cue, for challenge-response authentication of the users, wherein the cues are user-selected, user-provided or user-created, pick, upon receipt of an authentication request associated with a claim of an identity of an existing user of said number of users, a subset of cues for which there are voiceprints of the existing user stored, and provide the cues for representation to the user as a challenge, receive sound data indicative of the voice responses uttered by the user to the represented cues.

Claims (37)

1. An electronic system for authenticating a user of an electronic service, said system comprising:

at least one server apparatus; and

a hardware processor configured to store, for a number of users, a plurality of personal voiceprints each of which being linked with a dedicated visual, audiovisual or audio cue, for challenge-response authentication of the users, wherein the cues are user-selected, user-provided or user-created,

pick, upon receipt of an authentication request associated with a claim of an identity of an existing user of said number of users, a subset of cues for which there are voiceprints of the existing user stored, and provide the cues for representation to the user as a challenge,

receive sound data incorporating a contact microphone signal, comprising a throat microphone signal, in addition to a mouth microphone signal, comprising a close-speaking microphone signal, the sound data being indicative of voice response uttered by the user to the represented cues,

determine on the basis of the sound data, the represented cues and voiceprints linked therewith and the existing user, whether the response has been uttered by the existing user of said number of users, wherein the sound data indicative of the voice responses uttered to the represented cues are matched as concatenated against a concatenated voiceprint established based on the voiceprints linked with the represented cues and the existing user, and provided that this is the case,

elevate the authentication status of the user as the existing user, regarding at least the current communication session

wherein the hardware processor is further configured to utilize, based on the obtained data indicative of the location of the user, the estimated location of the user as an authentication factor.

2. The system of claim 1 , wherein at least one cue comprises a graphical image or video to be shown to the user via a display of a terminal device.

3. The system of claim 1 , wherein at least cue comprises an audio file, music or sound scenery file, to be audibly reproduced to the user.

4. The system of claim 1 , further configured to initially determine a personal voiceprint for a cue based on a voice response of the user to the cue.

5. The system of claim 1 , further comprising a first user terminal for accessing the service and reproducing the cues and a dynamic ID allocated by the system to the user.

6. The system of claim 5 , further comprising a second user terminal, a mobile device, comprising application for capturing the voice response by the user.

7. The system of claim 6 , wherein the second user terminal is further configured to obtain a dynamic ID allocated to the first terminal, browser thereat, and signal it to said at least one server of the system.

8. The system of claim 7 , wherein the second user terminal is configured to read a two-dimensional code representation of the ID shown on the display of the first terminal.

9. The system of claim 1 , configured to combine the contact and mouth microphone signals for obtaining authentic signal representing the uttered voice response with stop consonants, nasal cavity, tongue and lips-based sounds preserved.

10. The system of claim 1 , configured to apply the contact microphone signal in addition to the mouth microphone signal to reduce the effect of background noise.

11. An electronic device for authenticating a person, the electronic device comprising:

a voiceprint memory configured to store, for a number of users including at least one user, a plurality of personal voiceprints, each of which being linked with a dedicated visual, audiovisual or audio cue, for challenge-response authentication, wherein the cues are user-selected, user-provided or user-created,

an authentication hardware processor configured to pick, upon receipt of an authentication request associated with a claim of an identity of an existing user of said number of users, a subset of cues for which there are voiceprints of the existing user stored, and represent the cues in the subset to the person as a challenge, and

a response provision hardware processor for obtaining sound data incorporating a contact microphone signal, a throat microphone signal, in addition to a mouth microphone signal, the sound data being indicative of the voice response uttered by the person to the represented cues,

whereupon the authentication hardware processor is configured to determine, on the basis of the sound data, the represented cues and voiceprints linked therewith and the existing user, whether the response has been uttered by the existing user of said number of users, wherein the sound data indicative of the voice responses uttered to the represented cues are matched as concatenated against a concatenated voiceprint based on the voiceprints linked with the represented cues and the existing user, and provided that this is the case,

to elevate the authentication status of the person as the existing user,

wherein the electronic device is further configured to utilize, based on the obtained data indicative of the location of the user, the estimated location of the user as an authentication factor.

12. The electronic device of claim 11 , being or comprising at least one element selected from the group consisting of: portable communications-enabled user device, computer, desktop computer, laptop computer, personal digital assistant, mobile terminal, smartphone, tablet, wristop computer, access control terminal or panel, smart goggles, and wearable user device.

13. The electronic device of claim 11 , configured to control, responsive to the authentication status, access to a physical location or resource, via a controllable locking or unlocking mechanism, electrically controlled lock of a door, lid, or hatch.

14. The electronic device of claim 11 , configured to control, responsive to the authentication status, further access to the device itself or a feature, such as application feature or UI feature, thereof or at least accessible therethrough.

15. A method for authenticating a subject person to be executed by one or more electronic devices, comprising:

storing, by a hardware processor, for a number of users, a plurality of personal voiceprints each of which linked with a dedicated visual, audiovisual or audio cue, for challenge-response authentication of the users, cues being user-selected, user-provided or user-created:

picking, by the hardware processor, upon receipt of an authentication request associated with a claim of an identity of an existing user of said number of users, a subset of cues for which there are voiceprints of the existing user stored, to be represented as a challenge:

receiving, by the hardware processor, response incorporating sound data incorporating a contact microphone signal, a throat microphone signal, in addition to a mouth microphone signal, the sound data being indicative of the voice response uttered by the person to the represented cues, wherein the voice response is captured utilizing a throat microphone:

determining, by the hardware processor, on the basis of the sound data, the represented cues and voiceprints linked therewith and the existing user, whether the response has been uttered by the existing user, wherein the sound data indicative of the voice responses uttered to the represented cues are matched as concatenated against a concatenated voiceprint established based on the voiceprints linked with the represented cues and the existing user, and provided that this is the case: and

elevating, by the hardware processor, the authentication status of the person acknowledged as the existing user according to the determination,

wherein the estimated location of the user, based on the obtained data indicative of the location of the user, is further utilized as an authentication factor.

16. The method of claim 15 , further comprising controlling access based on the authentication status, wherein access to an electronic resource, such as electronic service, device, or feature accessible using the service or device, or to a physical location or resource, a space or container associated with an electric lock, door, or latch, is controlled.

17. The method of claim 15 , wherein a predefined normalization technique is utilized to compensate variation in signal characteristics arising from the used sound capturing equipment, environmental conditions and/or human voice generation mechanism itself.

18. The method of claim 15 , wherein said determining comprises utilizing at least one modeling, matching, or normalization element selected from the group consisting of: template matching, hidden Markov model, Gaussian mixture model, dynamic time warping, cepstral coefficients, likelihood ratio based scoring, linear prediction coefficients, blind equalization, H-norm normalization, Z-norm normalization, and T-norm normalization.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 044206 FRAME 0264. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 26, 2018
From: PITKÄNEN, MARTTI; PARTS, ROBERT; HUUHKA-PITKÄNEN, PIRJO
To: APLCOMP OY
Reel/Frame 045443/0955 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 24, 2017
From: PITKÄNEN, MARTTI; PARTS, ROBERT; HUUHKA-PITKÄNEN, PIRJO
To: ALCOMP OY
Reel/Frame 044206/0264 →
Priority Claims (2)
FI 20155197 · Mar 20, 2015 · national
FI 20154223 U · Dec 17, 2015 · national
Continuity (1)
Related Publication 20180068103A1 · Mar 8, 2018
Cited By (30)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,347,219 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,405,946 US 12,431,128 US 12,477,470 US 12,501,263 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452 US 12,659,162 US 12,748,568