IP Library › Granted Patent US 12,505,343
Granted Patent B2
US 12,505,343 · App. 17/376,195 · Granted Dec 23, 2025

Method and apparatus for image recognition using dual-subnetwork architecture

Inventors: Ali Arslan (Kanagawa, JP); Matteo Testa (Kanagawa, JP); Lev Markhasin (Kanagawa, JP); Tiziano Bianchi (Kanagawa, JP); Enrico Magli (Kanagawa, JP)
Assignees: Sony Semiconductor Solutions Corporation; Politechnico di Torino
G06N3/08G06F18/2132G06F18/22G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,343
App. No.
17/376,195
Granted
Dec 23, 2025
Kind
B2
Abstract

The present disclosure relates to an apparatus for image recognition. The apparatus comprises a machine learning network configured to map first and second input image data to either a first or a second predefined target probability distribution, depending on whether the first and second input image data correspond to matching or non-matching images, wherein an output of the machine learning network matching the first target probability distribution is indicative of matching images and an output of the machine learning network matching the second target probability distribution is indicative of non-matching images. The present disclosure also relates to a method for training the apparatus for image recognition.

Claims (21)

1 . An apparatus for image recognition, the apparatus comprising:

a processor configured to

map, by a machine learning network, first and second input image data to either a first or a second predefined target probability distribution, depending on whether the first and second input image data correspond to matching or non-matching images,

wherein an output of the machine learning network matching the first target probability distribution is indicative of matching images and an output of the machine learning network matching the second target probability distribution is indicative of non-matching images,

wherein the machine learning network comprises a first machine learning subnetwork configured to extract respective discriminative image features from the first and second image data, and a second machine learning subnetwork, wherein the second machine learning subnetwork is configured to transform concatenated features formed from concatenating the extracted first and second discriminative image features by mapping the concatenated features to one of the first and second predefined target probability distributions,

wherein the first machine learning subnetwork comprises a Siamese neural network configured to process the first and second image data in tandem to compute the first and second discriminative image features, and

select, within mini-batches during online training, from outputs of the machine learning network, a subset of matching pairs whose output is outside a predetermined distance threshold from a mass center of the first target probability distribution and a subset of non-matching pairs whose output is outside the predetermined distance threshold from a mass center of the second target probability distribution for use in subsequent training epochs, wherein the selection is performed after each training epoch to identify pairs from the machine learning network outputs for use in the subsequent training epoch.

2 . The apparatus of claim 1 , wherein the first and the second target probability distribution correspond to a first and a second multivariate Gaussian distribution with distinct centers of mass.

3 . The apparatus of claim 1 , wherein the second machine learning subnetwork comprises a convolutional neural network comprising an input layer for the first and second discriminative image features, a plurality of fully connected layers to apply a previously trained metric on the first and second discriminative image features, and an output for an m-dimensional output.

4 . The apparatus of claim 1 , further comprising a preprocessor configured to preprocess the first and second input image data for alignment of corresponding first and second images based on a plurality of predefined image points.

5 . A method for training the apparatus of claim 1 , the method comprising:

feeding image data of pairs of matching or non-matching images into the machine learning network;

adjusting computational weights of the machine learning network to minimize a difference between the predefined target probability distributions and statistics of outputs generated by the machine learning network.

6 . The method of claim 5 , wherein adjusting the computational weights comprises minimizing a difference between the first predefined target probability distribution and a distribution of outputs of the machine learning network in response to pairs of matching images, and minimizing a difference between the second predefined target probability distribution and a distribution of outputs of the machine learning network in response to pairs of non-matching images.

7 . The method of claim 5 , wherein adjusting the computational weights comprises minimizing the Kullback-Leibler divergence between the target probability distributions and the statistics of outputs.

8 . A method for image recognition, the method comprising:

mapping, using a machine learning network, first and second input image data to either a first or a second predefined target probability distribution, depending on whether the first and second input image data correspond to matching or non-matching images;

deciding for matching images if an output of the machine learning network matches the first target probability distribution or deciding for non-matching images if the output of the machine learning network matches the second target probability distribution,

wherein the machine learning network comprises a first machine learning subnetwork configured to extract respective discriminative image features from the first and second image data, and a second machine learning subnetwork, wherein the second machine learning subnetwork is configured to transform concatenated features formed from concatenating the extracted first and second discriminative image features by mapping the concatenated features to one of the first and second predefined target probability distributions,

wherein the first machine learning subnetwork comprises a Siamese neural network configured to process the first and second image data in tandem to compute the first and second discriminative image features; and

selecting, within mini-batches during online training, from outputs of the machine learning network, a subset of matching pairs whose output is outside a predetermined distance threshold from a mass center of the first target probability distribution and a subset of non-matching pairs whose output is outside the predetermined distance threshold from a mass center of the second target probability distribution for use in subsequent training epochs, wherein the selection is performed after each training epoch to identify pairs from the machine learning network outputs for use in the subsequent training epoch.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 15, 2021
From: MARKHASIN, LEV
To: SONY SEMICONDUCTOR SOLUTIONS CORPORATION; POLITECHNICO DI TORINO
Reel/Frame 056862/0854 →
Priority Claims (1)
EP 20187576 · Jul 24, 2020 · regional
Continuity (1)
Related Publication 20220027732A1 · Jan 27, 2022
References Cited (37)
KR 102036957B1 · 2016 [cited by examiner]
WO WO2021147366A1 · 2021 [cited by examiner]
Face Verification Using Convolutional Neural Networks with Siamese Architecture, Zuzana Bukovciková et al., Research Gate Sep. 2017 DOI: 10.23919/ELMAR.2017.8124469 (Year: 2017). [cited by examiner]
Deep Metric Learning for Practical Person Re-Identification, Dong Yi et al., arXiv:1407.4979v1 [cs.CV] Jul. 18, 2014 (Year: 2014). [cited by examiner]
The History Began from AlexNet: A Comprehensive Survey on Deep Learning Approaches, Md Zahangir Alom et al. , https://doi.org/10.48550/arXiv.1803.01164 (Year: 2018). [cited by examiner]
Audio-Based Semantic Concept Classification for Consumer Video, Keansub Lee, IEEE Transactions on Audio, Speech, and Language Processing, vol. 18, No. 6, Aug. 2010 (Year: 2010). [cited by examiner]
Deep Face Recognition, Parkhi et al., Proceedings of the British Machine Vision Conference 2015 Discloses relevant information regarding facial recognition technologies, specifically CNNs, their training and validation.… [cited by examiner]
Siamese Neural Networks for One-shot Image Recognition, Gregory Koch et al., Proceedings of the 32nd International Conference on Machine Learning, Lille, France, 2015 (Year: 2015). [cited by examiner]
Visual fingerprinting for lobsters using deep learning, Mae L. Seto, 2019 IEEE International Conference on Systems, Man and Cybernetics (SMC) Bari, Italy. Oct. 6-9, 2019 (Year: 2019). [cited by examiner]
Learning Local Matching with Hard Sample Mining for Person Re-identification, Yumin Suh, https://s-space.snu.ac.kr/handle/10371/151855 (Year: 2019). [cited by examiner]
FaceNet: A Unified Embedding for Face Recognition and Clustering, Florian Schoff et al., Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, pp. 815-823 (Year: 2015). [cited by examiner]
End-to-End Incremental Learning, Francisco M. Castro et al., arXiv:1807.09536v2 [cs.CV] Sep. 3, 2018 (Year: 2018). [cited by examiner]
Training Region-based Object Detectors with Online Hard Example Mining, Shrivastava et al., Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 761-769 (Year: 2016). [cited by examiner]
Dong Yi et al: “Deep Metric Learning for Practical Person ReIdentification”, ARXIV 1407.4979v1, Jul. 18, 2014 (Jul. 18, 2014), pp. 1-11, XP055279061, Journal of Latex Class Files, vol. 11, No. 4, Dec. 2012. [cited by applicant]
Guo Zhenning: “Distributed Machine Learning over Directed Network with Fixed Communication Delays”, Machine Learning and Computing, ACM, Feb. 22, 2019 (Feb. 22, 2019), pp. 22-26, XP058435668, DOI: 10.1145/3318299.331834… [cited by applicant]
Testa Matteo et al: “Learning mappings onto regularized latent spaces for biometric authentication”, 2019 IEEE 21st International Workshop on Multimedia Signal Processing (MMSP), IEEE, Sep. 27, 2019 (Sep. 27, 2019), pp.… [cited by applicant]
Bukovcikova Zuzana et al: “Face verification using convolutional neural networks with Siamese architecture”, 2017 International Symposium Elmar, Croatian Society Electronics in Marine—Elmar, Sep. 18, 2017 (Sep. 18, 2017… [cited by applicant]
Arslan Ali et al: “BioMetricNet: deep unconstrained face verification through learning of metrics regularized onto Gaussian distributions”, arxiv.org, Cornell University Library, 201OLIN Library Cornell University Ithac… [cited by applicant]
Brilli Stefano: “BiometricNet—A deep learning-based approach for biometric authentication”, Master's Degree Thesis, Jul. 28, 2020 (Jul. 28, 2020), XP055864304, Torino, Italy Retrieved from the Internet: URL:https://webt… [cited by applicant]
Diederik P Kingma et al: “Auto-Encoding Variational Bayes”, Dec. 20, 2013 (Dec. 20, 2013), XP055452075, Retrieved from the Internet: URL:https://arxiv.0rg/pdf/1312.6114.pdf [retrieved on Feb. 16, 2018]. [cited by applicant]
Extended European Search Report issued Dec. 3, 2021, in European Patent Application 21185682.8. [cited by applicant]
Wen et al., “A Discriminative Feature Learning Approach for Deep Face Recognition”, Springer International Publishing AG, 2016, pp. 499-515. [cited by applicant]
Wang et al., “Additive Margin Softmax for Face Verification”, Available Online At: https://github.com/happynear/AMSoftmax, Workshop track—ICLR, 2018, pp. 1-9. [cited by applicant]
Deng et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition”, arXiv:1801.07698v3, Available Online At: https://github.com/deepinsight/insightface, Feb. 9, 2019, 11 pages. [cited by applicant]
“BioMetricNet: Deep Unconstrained Face Recognition Through Regularized Learning of Mappings Onto Target Distibutions in Latent Space”, Anonymous CVPR submission, 2019, pp. 1-10. [cited by applicant]
Sun et al., “Deep Learning Face Representation by Joint Identification-Verification”, NIPS'14: Proceedings of the 27th International Conference on Neural Information Processing Systems—vol. 2, Dec. 2014, pp. 1988-1996. [cited by applicant]
Sun et al., “DeepID3: Face Recognition with Very Deep Neural Networks”, arXiv:1502.00873v1, Feb. 3, 2015, pp. 1-5. [cited by applicant]
Sun et al., “Deeply Learned Face Representations are Sparse, Selective, and Robust”, arXiv:1412.1265v1, Dec. 3, 2014, pp. 1-12. [cited by applicant]
Qi et al., “Face Recognition via Centralized Coordinate Learning”, arXiv:1801.05678v1, Jan. 17, 2018, pp. 1-14. [cited by applicant]
Schroff et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering”, arXiv:1503.03832v3, Jun. 17, 2015, 10 pages. [cited by applicant]
Hasnat et al., “DeepVisage: Making Face Recognition Simple Yet With Powerful Generalization Skills”, Computer Vision Foundation, Mar. 2017, pp. 1682-1691. [cited by applicant]
Ranjan et al., “L2-constrained Softmax Loss for Discriminative Face Verification”, arXiv:1703.09507v3, Jun. 7, 2017, 10 pages. [cited by applicant]
Liu et al., “Large-Margin Softmax Loss for Convolutional Neural Networks Weiyang”, Proceedings of the 33 rd International Conference on Machine Learning, MLR: W&CP, vol. 48, arXiv:1612.02295v4, Nov. 17, 2017, 10 pages. [cited by applicant]
Liu et al., “SphereFace: Deep Hypersphere Embedding for Face Recognition”, Computer Vision Foundation, 2017, pp. 213-220. [cited by applicant]
Wang et al., “NormFace: L2 Hypersphere Embedding for Face Verification”, arXiv:1704.06369v4, Jul. 26, 2017, 11 pages. [cited by applicant]
Liu et al., “Rethinking Feature Discrimination and Polymerization for Large-scale Recognition”, 31st Conference on Neural Information Processing Systems (NIPS 2017) Deep Learning Workshop, arXiv: 1710.00870v2, Oct. 29, … [cited by applicant]
Wang et al., “CosFace: Large Margin Cosine Loss for Deep Face Recognition”, Computer Vision Foundation, Jan. 2018, pp. 5265-5274. [cited by applicant]