IP Library Granted Patent US 12,321,859
Granted Patent B2
US 12,321,859 · App. 18/332,230 · Granted Jun 3, 2025

Image classification and labeling

Inventors: Sandra Mau (Pittsburgh, PA); Sabesan Sivapalan (Kelvin Grove, AU)
Assignee: SEE-OUT PTY LTD
G06N3/08G06F16/55G06F16/583G06F18/2155G06F18/2415G06F18/254G06V10/454G06V10/764G06V10/774G06V10/776G06V10/82G06V20/10G06V20/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,859
App. No.
18/332,230
Granted
Jun 3, 2025
Kind
B2
Abstract

A method of training an image classification model includes obtaining training images associated with labels, where two or more labels of the labels are associated with each of the training images and where each label of the two or more labels corresponds to an image classification class. The method further includes classifying training images into one or more classes using a deep convolutional neural network, and comparing the classification of the training images against labels associated with the training images. The method also includes updating parameters of the deep convolutional neural network based on the comparison of the classification of the training images against the labels associated with the training images.

Claims (33)

1. A method, comprising:

obtaining, by processing circuitry, training images associated with labels, at least one training image of the training images being associated with two or more labels of the labels, each label corresponding to an image classification class, the labels having a hierarchical structure; and

classifying, by the processing circuitry, an input image into two or more classes based on trained at least two convolutional neural networks, the at least two convolutional neural networks having been trained using the training images and the hierarchically-structured labels associated with the training images, a separate convolutional neural network having been trained for each level of the hierarchical structure, the classifying the input image including

multiplying a probability score output from a first trained convolutional neural network for a label by a probability score output from a second trained convolutional neural network for the label.

2. The method according to claim 1 , wherein a classification layer of each respective convolutional neural network is based on soft-sigmoid activation, the soft-sigmoid activation being a combination of a softmax function and a sigmoid function.

3. The method according to claim 1 , wherein the training images and the input image include graphically-designed images.

4. The method according to claim 1 , wherein the labels are non-mutually exclusive labels.

5. The method according to claim 1 , wherein the labels are codes used by a trademark registration organization.

6. The method according to claim 1 , wherein the labels are codes used to classify design patent images or industrial design images.

7. The method according to claim 1 , wherein the labels are available as metadata of the training images associated with the labels.

8. The method according to claim 1 , wherein the classifying further includes

labelling the input image with two or more labels corresponding to the two or more classes.

9. The method according to claim 1 , further comprising

pre-processing, by the processing circuitry, the training images, the training of the at least two convolutional neural networks being based on the pre-processed training images and the labels associated with the training images.

10. A non-transitory computer-readable storage medium storing computer-readable instructions that, when executed by a computer, cause the computer to perform a method, the method comprising:

obtaining training images associated with labels, at least one training image of the training images being associated with two or more labels of the labels, each label corresponding to an image classification class, the labels having a hierarchical structure; and

classifying an input image into two or more classes based on trained at least two convolutional neural networks, the at least two convolutional neural networks having been trained using the training images and the hierarchically-structured labels associated with the training images, a separate convolutional neural network having been trained for each level of the hierarchical structure, the classifying the input image including

multiplying a probability score output from a first trained convolutional neural network for a label by a probability score output from a second trained convolutional neural network for the label.

11. The non-transitory computer-readable storage medium according to according to claim 10 , wherein a classification layer of each respective convolutional neural network is based on soft-sigmoid activation, the soft-sigmoid activation being a combination of a softmax function and a sigmoid function.

12. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the training images and the input images include graphically-designed images.

13. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the labels are non-mutually exclusive labels.

14. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the labels are codes used by a trademark registration organization.

15. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the labels are codes used to classify design patent images or industrial design images.

16. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the labels are available as metadata of the training images associated with the labels.

17. The non-transitory computer-readable storage medium according to according to claim 10 , wherein the classifying further includes

labelling the input image with two or more labels corresponding to the two or more classes.

18. The non-transitory computer-readable storage medium according to according to claim 10 , further comprising

pre-processing the training images, the training of the at least two convolutional neural networks being based on the pre-processed training images and the labels associated with the training images.

19. An apparatus, comprising:

processing circuitry configured to

obtain, by processing circuitry, training images associated with labels, at least one training image of the training images being associated with two or more labels of the labels, each label corresponding to an image classification class, the labels having a hierarchical structure, and

classify, by multiplying a probability score output from a first trained convolutional neural network for a label by a probability score output from a second trained convolutional neural network for the label, an input image into two or more classes based on trained at least two convolutional neural networks, the at least two convolutional neural networks having been trained using the training images and the hierarchically-structured labels associated with the training images, a separate convolutional neural network having been trained for each level of the hierarchical structure.

20. The apparatus according to claim 19 , wherein a classification layer of each respective convolutional neural network is based on soft-sigmoid activation, the soft-sigmoid activation being a combination of a softmax function and a sigmoid function.

Continuity (4)
Continuation 17308440 · May 5, 2021
Continuation 16074399
Provisional Application 62289902 · Feb 1, 2016
Related Publication 20230316079A1 · Oct 5, 2023
References Cited (48)
US 8873812B2 · Larlus-Larrondo · 2014 [cited by applicant]
US 10102454B2 · Merler · 2018 [cited by applicant]
US 10282677B2 · Merler · 2019 [cited by applicant]
US 20140086497A1 · Fei-Fei · 2014 [cited by applicant]
US 20150235073A1 · Hua · 2015 [cited by examiner]
US 20150254532A1 · Talathi et al. · 2015 [cited by applicant]
US 20170109615A1 · Yatviz · 2017 [cited by applicant]
US 20180089540A1 · Merler · 2018 [cited by applicant]
US 20180204111A1 · Zadeh · 2018 [cited by examiner]
US 20180373960A1 · Xu · 2018 [cited by applicant]
US 20190251392A1 · Brown · 2019 [cited by applicant]
US 20200401851A1 · Mau et al. · 2020 [cited by applicant]
WO 2015035477 · 2015 [cited by applicant]
Deng, Jia, et al. “Large-scale object classification using label relation graphs.” European conference on computer vision. Springer, Cham, 2014. (Year: 2014). [cited by applicant]
Hu, Hexiang, et al. “Learning Structured Inference Neural Networks with Label Relations.” 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2016. (Year: 2016). [cited by applicant]
Gong, Y. et al.—“Deep Convolutional Ranking for Multilabel Image Annotation”—Cornell University Library—Apr. 14, 2014, whole document. [cited by applicant]
Simony AN, K et al.—“Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps”—Cornell University Library—Apr. 19, 2014, whole document. [cited by applicant]
Krizhevsky, A. et al.—“ImageNet Classification with Deep Convolutional Neural Networks”, Advances in Neural Information Processing Systems 25 (NIPS 2012) whole document. [cited by applicant]
Wei, Y. et al.—“CNN: Single-label to Multi-label”,—Journal of Latex Class Files, vol. 6, No. 1, Jan. 201 whole document. [cited by applicant]
International Search Report issued in PCT/IB2017/000134, May 23, 2017, 4 pages. [cited by applicant]
European Extended Search Report issued in EP 17747065.5, Aug. 22, 2019, 6 pages. [cited by applicant]
Min-Ling Zhang et al: “Multilabel Neural Networks with Applications to Functional Genomics and Text Categorization”, IEEE Trans on Knowledge and Data Engineering, Oct. 2006, pp. 1338-1351, XP055612826, DOI: 10.1109/TKDE… [cited by applicant]
Search Report and Written Opinion issued in application No. I 1201806541R, Jan. 28, 2020, pages, Intellectual Property Office of Singapore. [cited by applicant]
Huang Y. et al., Multi-Task Deep Neural Network for Multi-Label Learning. 2013 IEEE International Conference on Image Processing, Sep. 15, 2013, pp. 2897-2900. [cited by applicant]
Vallet A. et al., A Multi-Label Convolutional Neural Network for Automatic Image Annotation. Journal ofInformation Processing, Jul. 1, 2015, vol. 23, No. 6, pp. 767-775. [cited by applicant]
Zhang M.-L. et al., Multi-Label Neural Networks with Applications to Functional Genomics and Text Categorization. IEEE Transactions on Knowledge and Data Engineering, Oct. 1, 2006, pp. 1338-1351. [cited by applicant]
Cerri R. et al., Hierarchical multi-label classification using local neural networks. Journal of Computer and System Sciences, Mar. 22, 2013, pp. 39-56. [cited by applicant]
Examination Report issued Sep. 23, 2020 in corresponding Australian Patent Application No. 2017214619. [cited by applicant]
Vallet, Alexis, and Hiroyasu Sakamoto. “A multi-label convolutional neural network for automatic image annotation.” Journal of information processing 23.6 (2015): 767-775. (Year: 2015). [cited by applicant]
Everingham, Mark, et al. “The pascal visual object classes (voe) challenge.” International journal of computer vision 88.2 (2010): 303-338. (Year: 2010). [cited by applicant]
Leskovec, Jure, et al. Mining of massive datasets. 3rd ed., Cambridge University Press, 2014, pp. 509-556. (Year: 2014). [cited by applicant]
Gulcehre, Caglar, et al. “Noisy Activation Functions.” arXiv preprint arXiv: 1603.00391v1 (2016). (Year: 2016). [cited by applicant]
Office Action issued Dec. 22, 2020 in corresponding Japanese Patent Application No. 2018-558501. [cited by applicant]
M. Saito et al., “Illustration2Vec: A Semantic Vector Representation of Illustrations”, SIGGRAPH Asia 2015 Technical Briefs, ACM, 2015, pp. 5:1-5:4, URL, https://dl.acm.org/doi/pdf/10.1145/2820903.2820907. [cited by applicant]
Bannour, Hichem, and Celine Hudelot. “Hierarchical image annotation using semantic hierarchies.” Proceedings of the 21st ACM international conference on Information and knowledge management. 2012. (Year: 2012). [cited by applicant]
Deng, Jia, et al. “Imagenet: A large-scale hierarchical image database.” 2009 IEEE conference on computer vision and pattern recognition. Ieee, 2009. (Year: 2009). [cited by applicant]
Examination report No. 4 for standard patent application issued May 20, 2021 in corresponding Australian Patent Application No. 2017214619 citing documents AA and AX therein, 6 pages. [cited by applicant]
Dumais, S. et al., “SIGIR '00: Proceedings of the 23rd annual international ACM SIGIR conference on Research and development in information retrieval”, Jul. 2000, pp. 256-263, https://doi.org/10.1145/345508.345593. [cited by applicant]
Office Action issued Apr. 30, 2021 in corresponding European Patent Application No. 17 747 065.5 (with English Translation), citing document AZ therein, 4 pages. [cited by applicant]
Celine Yens et al: “Decision Trees for Hierarchical Multi-Label Classification”, Machine Learning, vol. 73, No. 2, XP019646600, Aug. 1, 2008, pp. 185-214, ISSN: 1573-0565, DOI: 10.1007/S10994-008-5077-3. [cited by applicant]
European Office Action issued on Dec. 22, 2022 in European Patent Application No. 17 747 065.5, citing reference 24 therein, 5 pages. [cited by applicant]
Yan et al., “HD-CNN: Hierarchical Deep Convolutional Neural Networks for Large Scale Visual Recognition”, 2015 IEEE International Conference on Computer Vision (ICCV), 2015, pp. 2740-2748, XP032866619. [cited by applicant]
Office Action issued Aug. 2, 2022, in corresponding Japanese Patent Application No. 2021-105527 (with English Translation), 6 pages. [cited by applicant]
Examination report No. 1 for standard patent application issued Jul. 21, 2022, in corresponding Australian Patent Application No. 2021203831, 6 pages. [cited by applicant]
Examination Report issued Oct. 1, 2024 in Australian Patent Application No. 2023263508, citing documents 24-26 therein. [cited by applicant]
Asaduzzaman, M. et al., ‘Faster training using fusion of activation functions for feed forward neural networks’, International journal of neural systems, 2009, vol. 19, No. 6, pp. 437-448. [cited by applicant]
Manessi, F. et al., ‘Learning Combinations of Activation Functions’, Jan. 29, 2018, pp. 1-6. [cited by applicant]
Jie, R. et al., ‘Combined flexible activation functions for deep neural networks’, ICLR 2020 Conference, 2020, pp. 1-17. [cited by applicant]