IP Library Granted Patent US 12,499,548
Granted Patent B2
US 12,499,548 · App. 17/688,572 · Granted Dec 16, 2025

Methods and systems for authentication of a physical document

Inventors: Daniele Pizzocchero (London, GB); Jimmy Moore (London, GB); Zhiyuan Shi (London, GB); Christos Sagonas (London, GB); Mohan Mahadevan (London, GB); Yuanwei Li (London, GB)
Assignee: Onfido Ltd.
G06T7/11G06F2218/12G06T2207/10004G06T2207/10016G06T2207/20081G06T2207/20084G06T2207/30176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,548
App. No.
17/688,572
Granted
Dec 16, 2025
Kind
B2
Abstract

Described herein are computerized methods and systems for authentication of a physical document. An image capture device coupled to a mobile device captures a sequence of images of a physical document as at least one of the physical document or the image capture device is rotated, during which the mobile device tracks the physical document throughout the sequence of images, and adjusts operational parameters of the image capture device based upon imaging conditions associated with the physical document. The mobile device selects images from the sequence of images and classifies the physical document using the selected images. The mobile device identifies a region of interest in the physical document using the selected images and the classification. The mobile device reconstructs the region of interest, generates an authentication score for the document using the reconstructed region of interest, and determines whether the physical document is authentic based upon the authentication score.

Claims (110)

1 . A system for authenticating a physical document, the system comprising a mobile computing device coupled to an image capture device, the mobile computing device configured to:

capture, using the image capture device, a sequence of images of a physical document in a scene as at least one of the physical document or the image capture device is rotated or tilted along one or more axes, during which the mobile computing device:

tracks the physical document throughout the sequence of images by:

dynamically determining a minimum range of motion for the physical document based upon one or more of the imaging conditions or the operational parameters of the image capture device;

determining whether the rotation or tilt of the physical document or the image capture device satisfies the minimum range of motion; and

instructing a user of the mobile computing device to continue rotating or tilting the physical document or the image capture device until the minimum range of motion is satisfied, and

adjusts one or more operational parameters of the image capture device based upon one or more imaging conditions associated with the physical document, as detected in one or more images of the sequence of images;

select one or more images from the sequence of images and classify the physical document using the selected images;

identify a region of interest in the physical document using the selected images and the classification of the physical document;

reconstruct the region of interest using the selected images;

generate an authentication score for the physical document using the reconstructed region of interest; and

determine whether the physical document is authentic based upon the authentication score.

2 . The system of claim 1 , wherein the minimum range of motion comprises a rotation or tilt of at least a minimum number of degrees in each of one or more planes.

3 . The system of claim 1 , wherein the mobile computing device:

dynamically adjusts one or more lighting parameters of the image capture device during capture of the sequence of images and assesses a signal associated with a region of interest in the physical document; and

instructs the user of the mobile computing device to continue rotating or tilting the physical document or the image capture device until a minimum amount of signal associated with the region of interest is captured and the minimum range of motion is satisfied.

4 . The system of claim 3 , wherein the mobile computing device dynamically adjusts the one or more lighting parameters based upon one or more of: ambient lighting conditions, physical document characteristics, or amount of captured signal associated with the region of interest.

5 . The system of claim 1 , wherein tracking the physical document throughout the sequence of images comprises determining, for each image in the sequence of images, at least one of a location or a six-dimensional pose of the physical document in the image.

6 . The system of claim 1 , wherein the one or more imaging conditions comprise at least one or more of: lighting conditions, focus, or control attributes of the image capture device.

7 . The system of claim 6 , wherein the one or more operational parameters comprise at least one or more of: shutter speed, ISO speed, gain, aperture, flash intensity, flash duration, or light balance.

8 . The system of claim 1 , wherein selecting one or more images from the sequence of images comprises:

determining, for each image in the sequence of images, whether the image is usable or unusable for authentication; and

discarding the image when the image is determined as unusable.

9 . The system of claim 8 , wherein an image is determined to be unusable when: at least a portion of the physical document is occluded or missing, a viewing angle of the physical document exceeds a defined threshold, the image includes noise that exceeds a defined threshold, or at least a portion of the image is blurry.

10 . The system of claim 1 , wherein identifying a region of interest in the physical document using the selected images comprises:

for each image in the selected images:

detecting a location of the physical document in the image;

estimating a pose of the physical document in the image;

cropping a portion of the image based upon the detected location and the pose of the physical document;

estimating one or more characteristics of the physical document based upon the cropped portion of the image; and

aligning the cropped images based upon one or more of the estimated characteristics of the physical document in each cropped image.

11 . The system of claim 10 , wherein the mobile computing device identifies the region of interest in each of the aligned images based upon predefined coordinate values.

12 . The system of claim 1 , wherein the region of interest comprises an optical variable device (OVD).

13 . The system of claim 1 , wherein reconstructing the region of interest using the selected images comprises executing one or more of a robust principal component analysis (PCA) algorithm or a learned alternative mapping on the selected images to reconstruct the region of interest.

14 . The system of claim 1 , wherein the sequence of images of the physical document comprises a plurality of images of a front side of the physical document and a plurality of images of a back side of the physical document.

15 . The system of claim 1 , wherein generating an authentication score for the physical document using the reconstructed region of interest comprises executing one or more machine learning classification models using one or more features of the reconstructed region of interest as input to generate a classification value for the physical document.

16 . The system of claim 15 , wherein the one or more machine learning classification models comprise one or more of: deep learning models, Random Forest algorithms, Support Vector Machines, neural networks, or ensembles thereof.

17 . The system of claim 15 , wherein the classification value comprises at least one of a probability that the physical document is authentic, a confidence score that indicates whether the physical document is authentic, or a similarity metric that indicates whether the physical document is authentic.

18 . The system of claim 15 , wherein at least one of the one or more machine learning classification models is a convolutional neural network.

19 . The system of claim 15 , wherein the one or more machine learning classification models is an ensemble classifier comprised of a plurality of convolutional neural networks.

20 . The system of claim 15 , wherein one or more interpretable methods are used to validate the classification value.

21 . The system of claim 15 , wherein the one or more interpretable methods comprise occlusion of at least a portion of the physical document, perturbation of at least a portion of the physical document, or analysis of a heatmap of at least a portion of the physical document.

22 . The system of claim 21 , wherein an output of the one or more interpretable methods comprises an identification of the reconstructed region of interest that represents proof of the physical document being genuine or fraudulent.

23 . The system of claim 15 , wherein the one or more machine learning classification models are trained using a plurality of genuine documents, a plurality of fraudulent documents, or both.

24 . The system of claim 23 , wherein the classification value generated by the one or more machine learning classification models is a measure of similarity between one or more of the plurality of genuine documents, one or more of the plurality of fraudulent documents, or both.

25 . The system of claim 1 , wherein the mobile computing device preprocesses the sequence of images received from the image capture device prior to selecting the one or more images.

26 . The system of claim 25 , wherein preprocessing the sequence of images comprises one or more of: assessing video quality metrics for the entire sequence of images, detecting a location of the physical document in each image of the sequence of images, and determining one or more quality metrics for each image in the sequence of images.

27 . The system of claim 26 , wherein the video quality metrics comprise a length of the sequence of images, a frames-per-second (FPS) value associated with the sequence of images, and an image resolution associated with the sequence of images.

28 . The system of claim 26 , wherein the one or more quality metrics comprise (i) global image quality metrics including one or more of: glare, blur, white balance, or sensor noise characteristics, (ii) local image quality metrics including one or more of: blur, sharpness, text region confidence, character confidence, or edge detection, or (iii) both the global image quality metrics and the local image quality metrics.

29 . The system of claim 28 , wherein the sensor noise characteristics comprise one or more of: blooming, readout noise, or custom calibration variations.

30 . A computerized method of authenticating a physical document, the method comprising:

capturing, using an image capture device coupled to a mobile computing device, a sequence of images of a physical document in a scene as at least one of the physical document or the image capture device is rotated or tilted along one or more axes, during which the mobile computing device:

tracks the physical document throughout the sequence of images by:

dynamically determining a minimum range of motion for the physical document based upon one or more of the imaging conditions or the operational parameters of the image capture device;

determining whether the rotation or tilt of the physical document or the image capture device satisfies the minimum range of motion; and

instructing a user of the mobile computing device to continue rotating or tilting the physical document or the image capture device until the minimum range of motion is satisfied, and

adjusts one or more operational parameters of the image capture device based upon one or more imaging conditions associated with the physical document, as detected in one or more images of the sequence of images;

selecting, by the mobile computing device, one or more images from the sequence of images and classifying the physical document using the selected images;

identifying, by the mobile computing device, a region of interest in the physical document using the selected images and the classification of the physical document;

reconstructing, by the mobile computing device, the region of interest using the selected images;

generating, by the mobile computing device, an authentication score for the physical document using the reconstructed region of interest; and

determining, by the mobile computing device, whether the physical document is authentic based upon the authentication score.

31 . The method of claim 30 , wherein the minimum range of motion comprises a rotation or tilt of at least a minimum number of degrees in each of one or more planes.

32 . The method of claim 31 , further comprising:

dynamically adjusting one or more lighting parameters of the image capture device during capture of the sequence of images and assesses a signal associated with a region of interest in the physical document; and

instructing the user of the mobile computing device to continue rotating or tilting the physical document or the image capture device until a minimum amount of signal associated with the region of interest is captured and the minimum range of motion is satisfied.

33 . The method of claim 32 , wherein the mobile computing device dynamically adjusts the one or more lighting parameters based upon one or more of: ambient lighting conditions, physical document characteristics, or amount of captured signal associated with the region of interest.

34 . The method of claim 30 , wherein tracking the physical document throughout the sequence of images comprises determining, for each image in the sequence of images, at least one of a location or a six-dimensional pose of the physical document in the image.

35 . The method of claim 30 , wherein the one or more imaging conditions comprise at least one or more of: lighting conditions, focus, or control attributes of the image capture device.

36 . The method of claim 35 , wherein the one or more operational parameters comprise at least one or more of: shutter speed, ISO speed, gain, aperture, flash intensity, flash duration, or light balance.

37 . The method of claim 30 , wherein selecting one or more images from the sequence of images comprises:

determining, for each image in the sequence of images, whether the image is usable or unusable for authentication; and

discarding the image when the image is determined as unusable.

38 . The method of claim 37 , wherein an image is determined to be unusable when: at least a portion of the physical document is occluded or missing, a viewing angle of the physical document exceeds a defined threshold, the image includes noise that exceeds a defined threshold, or at least a portion of the image is blurry.

39 . The method of claim 30 , wherein identifying a region of interest in the physical document using the selected images comprises:

for each image in the selected images:

detecting a location of the physical document in the image;

estimating a pose of the physical document in the image;

cropping a portion of the image based upon the detected location and the pose of the physical document;

estimating one or more characteristics of the physical document based upon the cropped portion of the image; and

aligning the cropped images based upon one or more of the estimated characteristics of the physical document in each cropped image.

40 . The method of claim 39 , wherein the mobile computing device identifies the region of interest in each of the aligned images based upon predefined coordinate values.

41 . The method of claim 30 , wherein the region of interest comprises an optical variable device (OVD).

42 . The method of claim 30 , wherein reconstructing the region of interest using the selected images comprises executing one or more of a robust principal component analysis (PCA) algorithm or a learned alternative mapping on the selected images to reconstruct the region of interest.

43 . The method of claim 30 , wherein the sequence of images of the physical document comprises a plurality of images of a front side of the physical document and a plurality of images of a back side of the physical document.

44 . The method of claim 30 , wherein generating an authentication score for the physical document using the reconstructed region of interest comprises executing one or more machine learning classification models using one or more features of the reconstructed region of interest as input to generate a classification value for the physical document.

45 . The method of claim 44 , wherein the one or more machine learning classification models comprise one or more of: deep learning models, Random Forest algorithms, Support Vector Machines, neural networks, or ensembles thereof.

46 . The method of claim 44 , wherein the classification value comprises at least one of a probability that the physical document is authentic, a confidence score that indicates whether the physical document is authentic, or a similarity metric that indicates whether the physical document is authentic.

47 . The method of claim 44 , wherein at least one of the one or more machine learning classification models is a convolutional neural network.

48 . The method of claim 44 , wherein the one or more machine learning classification models is an ensemble classifier comprised of a plurality of convolutional neural networks.

49 . The method of claim 44 , wherein one or more interpretable methods are used to validate the classification value.

50 . The method of claim 49 , wherein the one or more interpretable methods comprise occlusion of at least a portion of the physical document, perturbation of at least a portion of the physical document, or analysis of a heatmap of at least a portion of the physical document.

51 . The method of claim 50 , wherein an output of the one or more interpretable methods comprises an identification of the reconstructed region of interest that represents proof of the physical document being genuine or fraudulent.

52 . The method of claim 44 , wherein the one or more machine learning classification models are trained using a plurality of genuine documents, a plurality of fraudulent documents, or both.

53 . The method of claim 52 , wherein the classification value generated by the one or more machine learning classification models is a measure of similarity between one or more of the plurality of genuine documents, one or more of the plurality of fraudulent documents, or both.

54 . The method of claim 30 , wherein the mobile computing device preprocesses the sequence of images received from the image capture device prior to selecting the one or more images.

55 . The method of claim 54 , wherein preprocessing the sequence of images comprises one or more of: assessing video quality metrics for the entire sequence of images, detecting a location of the physical document in each image of the sequence of images, and determining one or more quality metrics for each image in the sequence of images.

56 . The method of claim 55 , wherein the video quality metrics comprise a length of the sequence of images, a frames-per-second (FPS) value associated with the sequence of images, and an image resolution associated with the sequence of images.

57 . The method of claim 56 , wherein the one or more quality metrics comprise (i) global image quality metrics including one or more of: glare, blur, white balance, or sensor noise characteristics, (ii) local image quality metrics including one or more of: blur, sharpness, text region confidence, character confidence, or edge detection, or (iii) both the global image quality metrics and the local image quality metrics.

58 . The method of claim 57 , wherein the sensor noise characteristics comprise one or more of: blooming, readout noise, or custom calibration variations.

59 . A system for authenticating a physical document, the system comprising a mobile computing device coupled to an image capture device, the mobile computing device configured to:

capture, using the image capture device, a sequence of images of a physical document in a scene as at least one of the physical document or the image capture device is rotated, during which the mobile computing device:

tracks the physical document throughout the sequence of images, and

adjusts one or more operational parameters of the image capture device based upon one or more imaging conditions associated with the physical document, as detected in one or more images of the sequence of images;

preprocess the sequence of images received from the image capture device by performing one or more of: assessing video quality metrics for the entire sequence of images, detecting a location of the physical document in each image of the sequence of images, and determining one or more quality metrics for each image in the sequence of images, wherein the video quality metrics comprise a length of the sequence of images, a frames-per-second (FPS) value associated with the sequence of images, and an image resolution associated with the sequence of images;

select one or more images from the sequence of images and classify the physical document using the selected images;

identify a region of interest in the physical document using the selected images and the classification of the physical document;

reconstruct the region of interest using the selected images;

generate an authentication score for the physical document using the reconstructed region of interest; and

determine whether the physical document is authentic based upon the authentication score.

Assignments (5)
SECURITY INTEREST Recorded Jul 25, 2024
From: ONFIDO LTD
To: BMO BANK N.A., AS COLLATERAL AGENT
Reel/Frame 068079/0801 →
RELEASE OF SECURITY INTEREST Recorded Apr 9, 2024
From: HSBC INNOVATION BANK LIMITED (F/K/A SILICON VALLEY BANK UK LIMITED)
To: ONFIDO LTD
Reel/Frame 067053/0607 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 21, 2022
From: ONFIDO LTD
To: SILICON VALLEY BANK UK LIMITED
Reel/Frame 062200/0655 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2022
From: PIZZOCCHERO, DANIELE; MOORE, JIMMY; SHI, ZHIYUAN; SAGONAS, CHRISTOS; MAHADEVAN, MOHAN; LI, YUANWEI
To: ONFIDO LTD.
Reel/Frame 061624/0230 →
SUPPLEMENT NO. 1 TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 8, 2022
From: ONFIDO LTD
To: SILICON VALLEY BANK
Reel/Frame 060613/0293 →
Continuity (1)
Related Publication 20230281820A1 · Sep 7, 2023
References Cited (128)
US 7590344B2 · Petschnigg · 2009 [cited by applicant]
US 7672475B2 · O'Doherty et al. · 2010 [cited by applicant]
US 7925096B2 · Baxter et al. · 2011 [cited by applicant]
US 8025239B2 · Labrec et al. · 2011 [cited by applicant]
US 8156115B1 · Erol et al. · 2012 [cited by applicant]
US 8627939B1 · Jones et al. · 2014 [cited by applicant]
US 8786767B2 · Rihn et al. · 2014 [cited by applicant]
US 9058535B2 · Guigan · 2015 [cited by applicant]
US 9081988B2 · Dolev · 2015 [cited by applicant]
US 9153005B2 · Tremolada et al. · 2015 [cited by applicant]
US 9418282B2 · Rutz et al. · 2016 [cited by applicant]
US 9619701B2 · Ragnet · 2017 [cited by examiner]
US 9652690B2 · Eid · 2017 [cited by examiner]
US 9779296B1 · Ma et al. · 2017 [cited by applicant]
US 10019626B2 · Schilling et al. · 2018 [cited by applicant]
US 10019627B2 · Kutter et al. · 2018 [cited by applicant]
US 10127454B2 · Pau et al. · 2018 [cited by applicant]
US 10140511B2 · Macciola et al. · 2018 [cited by applicant]
US 10331291B1 · Poder et al. · 2019 [cited by applicant]
US 10354142B2 · Arlazarov et al. · 2019 [cited by applicant]
US 10402448B2 · Filgueiras de Araujo et al. · 2019 [cited by applicant]
US 10477087B2 · Rivard et al. · 2019 [cited by applicant]
US 10621426B2 · Gaubatz et al. · 2020 [cited by applicant]
US 10726256B2 · Wu et al. · 2020 [cited by applicant]
US 10789463B2 · Kutter et al. · 2020 [cited by applicant]
US 10872488B2 · Azanza Ladron et al. · 2020 [cited by applicant]
US 11144752B1 · Castelblanco Cruz · 2021 [cited by examiner]
US 11159763B2 · Donsbach · 2021 [cited by examiner]
US 11235613B2 · Bud et al. · 2022 [cited by applicant]
US 11527087B1 · Geusz et al. · 2022 [cited by applicant]
US 11546519B2 · Bhatia et al. · 2023 [cited by applicant]
US 11607901B2 · Cape et al. · 2023 [cited by applicant]
US 12026932B2 · Cheong et al. · 2024 [cited by applicant]
US 12056978B2 · Avitan et al. · 2024 [cited by applicant]
US 20070041628A1 · Fan · 2007 [cited by applicant]
US 20080158258A1 · Lazarus et al. · 2008 [cited by applicant]
US 20080231418A1 · Ophey et al. · 2008 [cited by applicant]
US 20120226600A1 · Dolev · 2012 [cited by applicant]
US 20130185618A1 · Macciola et al. · 2013 [cited by applicant]
US 20130262333A1 · Wicker et al. · 2013 [cited by applicant]
US 20140032406A1 · Roach et al. · 2014 [cited by applicant]
US 20140037196A1 · Blair · 2014 [cited by applicant]
US 20140055824A1 · Tremolada et al. · 2014 [cited by applicant]
US 20140105449A1 · Caton et al. · 2014 [cited by applicant]
US 20140355069A1 · Caton et al. · 2014 [cited by applicant]
US 20150093018A1 · Macciola et al. · 2015 [cited by applicant]
US 20150302421A1 · Caton et al. · 2015 [cited by applicant]
US 20150324390A1 · Macciola et al. · 2015 [cited by applicant]
US 20160307035A1 · Schilling et al. · 2016 [cited by applicant]
US 20160350592A1 · Ma et al. · 2016 [cited by applicant]
US 20170132465A1 · Kutter et al. · 2017 [cited by applicant]
US 20170132866A1 · Kuklinkski et al. · 2017 [cited by applicant]
US 20170200247A1 · Kuklinski et al. · 2017 [cited by applicant]
US 20190087942A1 · Ma · 2019 [cited by examiner]
US 20190164010A1 · Ma et al. · 2019 [cited by applicant]
US 20190213408A1 · Cali et al. · 2019 [cited by applicant]
US 20190220660A1 · Cali et al. · 2019 [cited by applicant]
US 20190372968A1 · Balogh et al. · 2019 [cited by applicant]
US 20190392196A1 · Sagonas et al. · 2019 [cited by applicant]
US 20200051059A1 · Filler · 2020 [cited by applicant]
US 20200210796A1 · Hsu et al. · 2020 [cited by applicant]
US 20200387700A1 · Wu · 2020 [cited by applicant]
US 20200394763A1 · Ma et al. · 2020 [cited by applicant]
US 20210117520A1 · Ross · 2021 [cited by applicant]
US 20210158036A1 · Huber, Jr. · 2021 [cited by applicant]
US 20210160687A1 · Ross et al. · 2021 [cited by applicant]
US 20210200855A1 · Guigan · 2021 [cited by applicant]
US 20220051423A1 · Du Preez · 2022 [cited by examiner]
US 20220343483A1 · Desai · 2022 [cited by applicant]
CN 112434727A · 2021 [cited by applicant]
DE 102018102015A1 · 2019 [cited by applicant]
EP 3526781 · 2018 [cited by applicant]
FR 2977349A3 · 2013 [cited by applicant]
WO 2021044082A1 · 2021 [cited by applicant]
Y. Yoon et al., “Online Multiple Pedestrians Tracking using Deep Temporal Appearance Matching Association,” arXiv:1907.00831v4 [cs.CV] Oct. 9, 2020, available at https://arxiv.org/pdf/1907.00831.pdf, 28 pages. [cited by applicant]
S. Mallick, “Object Tracking using OpenCV (C++ / Python),” Feb. 13, 2017, available at https://learnopencv.com/object-tracking-using-opencv-cpp-python/, 12 pages. [cited by applicant]
K. He et al., “Deep Residual Learning for Image Recognition,” arXiv:1512.03385v1 [cs.CV] Dec. 10, 2015, available at https://arxiv.org/pdf/1512.03385v1.pdf, 12 pages. [cited by applicant]
C. Szegedy et al., “Rethinking the Inception Architecture for Computer Vision,” arXiv:1500567v3 [cs.CV] Dec. 11, 2015, available at https://arxiv.org/pdf/1512.00567v3.pdf, 10 pages. [cited by applicant]
M. Tan & Q.V. Lee, “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks,” arXiv:1905.11946v5 [cs.LG] Sep. 11, 2020, available at https://arxiv.org/pdf/1905.11946.pdf, 11 pages. [cited by applicant]
C. Wang et al., “EfficientNet-eLite: Extremely Lightweight and Efficient CNN Models for Edge Devices by Network Candidate Search,” arXiv:2009.07409v1 [cs.CV] Sep. 16, 2020, available at https://arxiv.org/pdf/2009.07409v… [cited by applicant]
P. Simon and U. V., “Deep Learning based Feature Extraction for Texture Classification,” Third International Conference on Computing and Network Communications (CoCoNet'19), Procedia Computer Science 171 (2020), pp. 168… [cited by applicant]
T. Zhang et al., “Spatial-Temporal Recurrent Neural Network for Emotion Recognition,” arXiv:1705.0451v1 [cs.CV] May 12, 2017, available at https://arxiv.org/pdf/1705.04515.pdf, 8 pages. [cited by applicant]
Y. Dong et al., “A Hybrid Spatial-temporal Deep Learning Architecture for Lane Detection,” arXiv:2110.04079 [cs.CV] Oct. 14, 2021, available at https://arxiv.org/ftp/arxiv/papers/2110/2110.04079.pdf, 18 pages. [cited by applicant]
G. Balakrishnan et al., “VoxelMorph: A Learning Framework for Deformable Medical Image Registration,” arXiv:1809.05231v3 [cs.CV] Sep. 1, 2019, available at https://arxiv.org/pdf/1809.05231.pdf, 16 pages. [cited by applicant]
I. Rocco et al., “Convolutional neural network architecture for geometric matching,” arXiv:1703.05593v2 [cs.CV] Apr. 13, 2017, available at https://arxiv.org/pdf/1703.05593.pdf, 15 pages. [cited by applicant]
M. Jadenberg et al., “Spatial Transformer Networks,” arXiv:1506.02025v3 [cs.CV] Feb. 4, 2016, available at https://arxiv.org/pdf/1506.02025.pdf, 15 pages. [cited by applicant]
P. H. Seo et al., “Attentive Semantic Alignment with Offset-Aware Correlation Kernels,” arXiv:1808.02128v2 [cs.CV] Oct. 26, 2018, available at https://arxiv.org/pdf/1808.02128.pdf, 21 pages. [cited by applicant]
R. Chen et al., “Video Foreground Detection Algorithm Based on Fast Principal Component Pursuit and Motion Saliency,” Comput. Intell. Neurosci. 2019, doi: 10.1155/2019/4769185, published Feb. 3, 2019, available at https… [cited by applicant]
E. Candes et al., “Robust Principal Component Analysis?”, arXiv:0912.3599v1 [cs.IT] Dec. 18, 2009, available at https://arxiv.org/pdf/0912.3599.pdf, 39 pages. [cited by applicant]
E. Rosten and T. Drummond, “Machine learning for high-speed corner detection,” Computer Vision—ECCV 2006, Lecture Notes in Computer Science, vol. 3951, pp. 430-443, doi:10.1007/11744023_34, 14 pages. [cited by applicant]
E. Mair et al., “Adaptive and generic corner detection based on the accelerated segment test,” Computer Vision—ECCV 2010, Lecture Notes in Computer Science, vol. 6312, pp. 183-196, doi:10.1007/978-3-642-15552-9_14, 14 p… [cited by applicant]
K. Mikolajczyk and C. Schmid, “Scale & Affine Invariant Interest Point Detectors,” International Journal of Computer Vision 60(1), pp. 63-86, 2004, 24 pages. [cited by applicant]
M. Agarwal et al., “CenSurE: Center surround extremas for real time feature detection and matching,” Computer Vision—ECCV 2008, Lecture Notes in Computer Science, vol. 5305, pp. 102-115, doi: 10.1007/978-3-540-88693-8_8… [cited by applicant]
D.G. Lowe, “Distinctive Image Features from Scale-Invariant Keypoints,” International Journal of Computer Vision 60, pp. 99-110 (2004), doi: 10.1023/B:VISI.0000029664.99615.94, 28 pages. [cited by applicant]
E. Rublee et al., “ORB: an efficient alternative to Sift or Surf,” 2011 International Conference on Computer Vision, doi:10.1109/ICCV.2011.6126544, 8 pages. [cited by applicant]
S. Leutenegger et al., “BRISK: Binary Robust Invariant Scalable Keypoints,” 2011 International Conference on Computer Vision, pp. 2548-2555, doi:10.1109/ICCV.2011.6126542, 8 pages. [cited by applicant]
P.F. Alcantarilla et al., “Fast explicit diffusion for accelerated features in nonlinear scale spaces,” British Machine Vision Conf. (BMVC) 2013, doi: 10.5244/C.27.13, 11 pages. [cited by applicant]
P.F. Alcantarilla et al., “KAZE Features,” Computer Vision—ECCV 2012, Lecture Notes in Computer Science, vol. 7577, pp. 214-227, 14 pages. [cited by applicant]
A. Alahi et al., “Freak: Fast retina keypoint,” Proceedings of the 2012 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 510-517. [cited by applicant]
M. Calonder et al., “Brief: Computing a local binary descriptor very fast,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 34, No. 7, pp. 1281-1298 (2011), 29 pages. [cited by applicant]
E. Tola et al., “Daisy: An Efficient Dense Descriptor Applied to Wide Baseline Stereo,” IEEE Transactions on Pattern Matching and Machine Intelligence, 2010, vol. 32, No. 5, pp. 815-830, doi: 10.1109/TPAMI.2009.77, 17 p… [cited by applicant]
G. Levi and T. Hassner, “LATCH: Learned Arrangements of Three Patch Codes,” 2016 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Placid, NY, Mar. 7-10, 2016, available at https://talhassner.github… [cited by applicant]
K. Simonyan and A. Zisserman, “Very deep convolutional networks for large-scale image recognition,” arXiv:1409.1556v6 [cs.CV], Apr. 10, 2015, available at https://arxiv.org/pdf/1409.1556v6.pdf, 14 pages. [cited by applicant]
B. Lakshminarayanan et al., “Simple and Scalable Predictive Uncertainty Estimation using Deep Ensembles,” arXiv:1612.01474v3 [stat.ML] Nov. 4, 2017, available at https://arxiv.org/pdf/1612.01474v3.pdf, 15 pages. [cited by applicant]
R. Rahaman and A.H. Thiery, “Uncertainty Quantification and Deep Ensembles,” arXiv:2007.08792v4 [stat.ML] Nov. 2, 2021, available at https://arxiv.org/pdf/2007.08792.pdf, 16 pages. [cited by applicant]
R. Selvaraju et al., “Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization,” arXiv:1610.02391 [cs.CV] Dec. 3, 2019, available at https://arxiv.org/pdf/1610.02391.pdf, 23 pages. [cited by applicant]
G. Montavon et al., “Layer-Wise Relevance Propagation: An Overview,” Explainable AI: Interpreting, Explaining and Visualizing Deep Learning, Lecture Notes in Computer Science, vol. 11700, pp. 193-209, Sep. 10, 2019, Spr… [cited by applicant]
M. Sundararajan et al., “Axiomatic Attribution for Deep Networks,” arXiv:1703.01365v2 [cs.LG] Jun. 13, 2017, available at https://arxiv.org/pdf/1703.01365.pdf, 11 pages. [cited by applicant]
P. Kindermans et al., “Learning How to Explain Neural Networks: PatternNet and PatternAttribution,” arXiv:1705.05598v2 [stat.ML] Oct. 24, 2017, available at https://arxiv.org/pdf/1705.05598.pdf, 12 pages. [cited by applicant]
W. Lim et al., “The adoption of deep learning interpretability techniques on diabetic retinopathy analysis: a review,” Medical & Biological Engineering & Computing 60, pp. 633-642 (2022). [cited by applicant]
K. Simonyan et al., “Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps,” arXiv:1312.6034v2 [cs.CV] Apr. 19, 2014, available at https://arxiv.org/pdf/1312.6034.pdf, 8 pages. [cited by applicant]
M.D. Zeiler and R. Fergus, “Vizualizing and Understanding Convolutional Networks,” European Conference on Computer Vision (ECCV 2014), Lecture Notes in Computer Science, vol. 8689 (2014), Springer, pp. 818-833. [cited by applicant]
B. Gao and M.W. Spratling, “Robust Template Matching via Hierarchical Convolutional Features from a Shape Based CNN,” arXiv:2007.15817v3 [cs.CV] May 7, 2021, available at https://arxiv.org/pdf/2007.15817.pdf, 11 pages. [cited by applicant]
C. Kim et al., “End-to-end deep learning-based autonomous driving control for high-speed environment,” The Journal of Supercomputing 78, pp. 1961-1982 (2022), doi.org/10.1007/s11227-021-03929-8. [cited by applicant]
R. Polvara et al., “Toward End-to-End Control for UAV Autonomous Landing via Deep Reinforcement Learning,” 2018 International Conference on Unmanned Aircraft Systems (ICUAS), Jun. 12-15, 2018, DOI: 10.1109/ICUAS.2018.84… [cited by applicant]
J. Shi and Tomasi, “Good features to track,” 1994 Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, pp. 593-600, doi: 10.1109/CVPR.1994.323794. [cited by applicant]
A. Hartl et al., “Efficient Verification of Holograms Using Mobile Augmented Reality,” IEEE Transactions on Visualization and Computer Graphics, Nov. 2015, 10 pages. [cited by applicant]
J. Redmon et al., “You Only Look Once: Unified, Real-Time Object Detection,” arXiv:1506.02640v5 [cs.CV] May 9, 2016, available at https://arxiv.org/pdf/1506.02640.pdf, 10 pages. [cited by applicant]
W. Liu et al., “SSD: Single Shot MultiBox Detector,” arXiv:1512.02325v5 [cs.CV] Dec. 29, 2016, available at https://arxiv.org/pdf/1512.02325.pdf, 17 pages. [cited by applicant]
S. Ren et al., “Faster R-CNN: Toward Real-Time Object Detection with Region Proposal Networks,” arXiv:1506.01497v1 [cs.CV] Jun. 4, 2015, available at https://arxiv.org/pdf/1506.01497v1.pdf, 10 pages. [cited by applicant]
A. G. Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv:1704.0486v1 [cs.CV] Apr. 17, 2017, available at https://arxiv.org/pdf/1704.04861.pdf, 9 pages. [cited by applicant]
N. Wojke et al., “Simple Online and Realtime Tracking with a Deep Association Metric,” arXiv:1703.07402v1 [cs.CV] Mar. 21, 2017, available at https://arxiv.org/pdf/1703.07402.pdf, 5 pages. [cited by applicant]
P. Bergmann et al., “Tracking without bells and whistles,” arXiv:1903.05625v3 [cs.CV] Aug. 17, 2019, available at https://arxiv.org/pdf/1903.05625.pdf, 16 pages. [cited by applicant]
G. Ciaparrone et al., “Deep Learning in Video Multi-Object Tracking: A Survey,” arXiv:1907.12740v4 [cs.CV] Nov. 19, 2019, available at https://arxiv.org/pdf/1907.12740.pdf, 42 pages. [cited by applicant]
E. Bochinski et al., “Extending IOU Based Multi-Object Tracking by Visual Information” (2018), available at https://elvera.nue.tu-berlin.de/files/1547Bochinski2018.pdf, 6 pages. [cited by applicant]
X. Zhou et al., “Tracking Objects as Points,” arXiv:2004.01177v2 [cs.CV] Aug. 21, 2020, available at https://arxiv.org/pdf/2004.01177.pdf, 22 pages. [cited by applicant]
Hartl et al “AR-Based Hologram Detection on Security Documents Using a Mobile Phone” 18 [cited by applicant]
Soukup et al “Mobile Hologram Verification with Deep Learning” 15 [cited by applicant]