IP Library Granted Patent US 12,315,053
Granted Patent B2
US 12,315,053 · App. 18/544,288 · Granted May 27, 2025

Image classification through label progression

Inventors: Hessam Bagherinezhad (Seattle, WA); Maxwell Horton (Los Angeles, CA); Mohammad Rastegari (Kirkland, WA); Ali Farhadi (Seattle, WA)
Assignee: Apple Inc.
G06T11/60G06F18/2148G06F18/241G06N3/045G06N3/08G06V10/454G06V10/764G06V10/82G06V20/52G06V40/10G06T2210/22G06V20/68
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,315,053
App. No.
18/544,288
Granted
May 27, 2025
Kind
B2
Abstract

Systems and methods are disclosed for training neural networks using labels for training data that are dynamically refined using neural networks and using these trained neural networks to perform detection and/or classification of one or more objects appearing in an image. Particular embodiments may generate a set of crops of images from a corpus of images, then apply a first neural network to the set of crops to obtain a set of respective outputs. A second neural network may then be trained using the set of crops as training examples. The set of respective outputs may be applied as labels for the set of crops.

Claims (33)

1. A method, comprising:

obtaining an image from an image sensor of a computing device;

obtaining, by the computing device, person detection data from the image by applying, to the image, a neural network that has been trained using training data that includes a set of crops of a training image and a set of respective outputs from another neural network that are applied as labels for the set of crops; and

generating an alert signal with the computing device based on the person detection data.

2. The method of claim 1 , wherein the image is an image frame of a video from a camera.

3. The method of claim 1 , wherein the alert signal comprises a bounding box that is overlaid on a portion of the image.

4. The method of claim 3 , wherein the bounding box is a color-coded bounding box.

5. The method of claim 3 , further comprising storing the image having the bounding box at the computing device that applies the neural network to the image.

6. The method of claim 3 , further comprising transmitting the image having the bounding box from the computing device that applies the neural network to the image to a remote system.

7. The method of claim 1 , wherein the alert signal comprises an intruder alert.

8. The method of claim 1 , further comprising:

modifying, by the computing device, at least one parameter for operating the image sensor based on the person detection data; and

obtaining another image with the image sensor using the at least one modified parameter.

9. A device, comprising:

an image sensor; and

at least one processor configured to:

obtain an image from the image sensor;

obtain person detection data from the image by applying, to the image, a neural network that has been trained using training data that includes a set of crops of a training image and a set of respective outputs from another neural network that are applied as labels for the set of crops; and

generate an alert signal based on the person detection data.

10. The device of claim 9 , wherein the image is an image frame of a video from a camera.

11. The device of claim 9 , wherein the alert signal comprises a bounding box that is overlaid on a portion of the image.

12. The device of claim 11 , wherein the bounding box is a color-coded bounding box.

13. The device of claim 11 , further comprising storing the image having the bounding box at the device that applies the neural network to the image.

14. The device of claim 11 , further comprising transmitting the image having the bounding box from the device that applies the neural network to the image to a remote system.

15. A non-transitory machine-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:

obtaining an image from an image sensor of a computing device;

obtaining, by the computing device, person detection data from the image by applying, to the image, a neural network that has been trained using training data that includes a set of crops of a training image and a set of respective outputs from another neural network that are applied as labels for the set of crops; and

generating an alert signal with the computing device based on the person detection data.

16. The non-transitory machine-readable medium of claim 15 , wherein the image is an image frame of a video from a camera.

17. The non-transitory machine-readable medium of claim 15 , wherein the alert signal comprises a bounding box that is overlaid on a portion of the image.

18. The non-transitory machine-readable medium of claim 17 , wherein the bounding box is a color-coded bounding box.

19. The non-transitory machine-readable medium of claim 17 , wherein the operations further comprise storing the image having the bounding box at the computing device that applies the neural network to the image.

20. The non-transitory machine-readable medium of claim 17 , wherein the operations further comprise transmitting the image having the bounding box from the computing device that applies the neural network to the image to a remote system.

Continuity (4)
Division 17308032 · May 4, 2021
Continuation 16386151 · Apr 16, 2019
Provisional Application 62660901 · Apr 20, 2018
Related Publication 20240144566A1 · May 2, 2024
References Cited (55)
US 10497122B2 · Zhang · 2019 [cited by examiner]
US 20180268307A1 · Kobayashi · 2018 [cited by applicant]
US 20190205625A1 · Luo · 2019 [cited by examiner]
CN 105354595A · 2016 [cited by applicant]
CN 106446927A · 2017 [cited by applicant]
CN 107085585A · 2017 [cited by applicant]
CN 107358157A · 2017 [cited by applicant]
CN 107609587A · 2018 [cited by applicant]
JP 2013114596A · 2013 [cited by applicant]
Ba, et al., “Do Deep Nets Really Need to be Deep?” Advances in Neural Information Processing Systems, 2014, pp. 2654-2662. [cited by applicant]
Bagherinezhad, et al., “Label Refinery: Improving ImageNet Classification Through Label Progression,” University of Washington, May 2018. [cited by applicant]
Bucila, et al., “Model Compression,” Proceedings of the 12th ACM SIGKKD International Conference on Knowledge Discovery and Data Mining, 2006, pp. 535-541. [cited by applicant]
Cai, et al., “Exploiting Known Taxonomies in Learning Overlapping Concepts,” IJCAI, 2007, vol. 7, pp. 708-713. [cited by applicant]
Deng, et al., “ImageNet: A Large-Scale Hierarchical Image Database,” IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248-255. [cited by applicant]
He, et al., “Deep Residual Learning for Image Recognition,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
Hinton, et al., Distilling the Knowledge in a Neural Network, Mar. 2015. [cited by applicant]
Howard, et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” Apr. 2017. [cited by applicant]
Huang, et al., “Transfer Learning with Deep Convolutional Neural Network for SAR Target Classification with Limited Labeled Data,” Remote Sensing, Aug. 2017, vol. 9, No. 9, pp. 1-21. [cited by applicant]
Ioffe, et al., “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift,” International Conference on Machine Learning, 2015, pp. 118-156. [cited by applicant]
Kingma, et at., “Adam: A Method for Stochastic Optimization,” ICLR Conference, 2015. [cited by applicant]
Krizhevsky, et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems, 2012, pp. 1097-1105. [cited by applicant]
Li, et al., “Learning from Noisy Labels with Distillation,” Open Access Version from ICCV, provided by the Computer Vision Foundation, 2017. [cited by applicant]
Li, et al., “Learning Small-Size DNN with OUTPUT-Distribution-Based Criteria,” 15th Annual Conference of the International Speech Communication Association, 2014. [cited by applicant]
McAuley, et al., “Optimization of Robust Loss Functions for Weakly-Labeled Image Taxonomies,” Sep. 2012. [cited by applicant]
Miller, et al., “Introduction to WordNet: An On-Line Lexical Database,” International Journal of Lexicography, 1990, vol. 3, No. 4, pp. 235-244. [cited by applicant]
Miyato, et al., “Distributional Smoothing by Virtual Adversarial Training,” ICLR Conference, 2016. [cited by applicant]
Nguyen, et al., “Deep Neural Networks are Easily Fooled: high Confidence Predictions for Unrecognizable Images,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2015, pp. 427-436. [cited by applicant]
Paszke, et al., “Automatic Differentiation in PyTorch,” 31st Conference on Neural Information Processing Systems, 2017. [cited by applicant]
Pereyra, et al., “Regularizing Neural Networks by Penalizing Confident Output Distributions,” Under review as a conference paper at ICLR 2017. [cited by applicant]
Rastegari, et al., “XNOR-Net: ImageNet Classification Using Binary Convolutional Neural Networks,” European Conference on Computer Vision, 2016, pp. 525-542. [cited by applicant]
Redmon, et al., “Yolo9000: Better, Faster, Stronger,” IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 6517-6525. [cited by applicant]
Reed, et al., “Training Deep Neural Networks on Noisy Labels with Bootstrapping,” Accepted as a workshop contribution at ICLR 2015. [cited by applicant]
Romero, et al., “FitNets: Hints for Thin Deep Nets,” ICLR Conference, 2015. [cited by applicant]
Shen, et al., “In Teacher We Trust: Learning Compressed Models for Pedestrian Detection,” Dec. 2016. [cited by applicant]
Shrivastava, et al., “Learning from Simulated and Unsupervised Images through Adversarial Training,” IEEE Conference on Computer Vision and Pattern Recognition, 2017. [cited by applicant]
Simonyan, et al., “Very Deep Convolutional Networks for Large-Scale Image Recognition,” ICLR Conference, 2015. [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning,” AAAI vol. 4, 2017. [cited by applicant]
Szegedy, et al., “Going Deeper with Convolutions,” Open Access version from ICCV, provided by the Computer Vision Foundation, 2015. [cited by applicant]
Szegedy, et al., “Rethinking the Inception Architecture for Computer Vision,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 2818-2826. [cited by applicant]
Tuytelaars, et al., “The NBNN Kernel,” IEEE International Conference on Computer Vision, Nov. 2011. [cited by applicant]
Wang, et al., “The Effectiveness of Data Augmentation in Image Classification Using Deep Learning,” 2017. [cited by applicant]
Wei, et al., “HCP: A Flexible CNN Framework for Multi-Label Image,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Sep. 2016, vol. 38, No. 9, pp. 1907-1907. [cited by applicant]
Wong, et al., “Understanding Data Augmentation for Classification: When to Warp?” IEEE International Conference on Digital Image Computing: Techniques and Applications, 2016, pp. 1-6. [cited by applicant]
Wu, et al., “MI-MG Multi-label Learning with Missing Labels Using a Mixed Graph,” Proceedings of the IEEE International Conference on Computer Vision, 2015, pp. 4157-4165. [cited by applicant]
Wu, et al., “Verb Semantics and Lexical Selection,” Jun. 1994. [cited by applicant]
Wu, et al., “Hierarchical Loss for Classification,” 2017. [cited by applicant]
Xie, et al., “DisturbLabel: Regularizing CNN on the Loss Layer,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 4753-4762. [cited by applicant]
Xu, et al., “Improved Relation Classification by Deep Recurrent Neural Networks with Data Augmentation,” Oct. 2016. [cited by applicant]
International Search Report and Written Opinion from PCT/US2019/028570, Sep. 11, 2019. [cited by applicant]
Liu, et al., “Active contour driven by region-scalable fitting and Kullback-Leibler divergence for image segmentation,” Journal of Harbin Institute of Technology, May 2016, vol. 48, No. 5, 9 pages including English lang… [cited by applicant]
Wei, et al., “HCP: A Flexible CNN Framework for Multi-Label Image Classification,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2015, 8 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980027189.4, dated Mar. 20, 2024, 14 pages including machine-generated English language translation. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980027189.4, dated Sep. 27, 2024, 11 pages including English language translation. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980027189.4, dated Jan. 1, 2025, 11 pages with English language translation. [cited by applicant]
Apple Inc., R2-1817467, 3GPP TSG-RAN WG2 Meeting #104, Nov. 16, 2018, 3 pages. [cited by applicant]