IP Library Granted Patent US 12,614,062
Granted Patent B2
US 12,614,062 · App. 16/905,252 · Granted Apr 28, 2026

Techniques for classification with neural networks

Inventors: Pablo Ribalta (Tychy, PL); Michal Marcinkiewicz (Katowice, PL); Dong Yang (Pocatello, ID); Daguang Xu (Potomac, MD); Przemyslaw Miroslaw Strzelczyk (Warsaw, PL)
Assignee: NVIDIA Corporation
G06N3/063G06N3/045G06N3/08G06V10/764G06V10/774G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,062
App. No.
16/905,252
Filed
Jun 18, 2020
Granted
Apr 28, 2026
Kind
B2
Art Unit
2124
USPC
706/20
Abstract

Apparatuses, systems, and techniques to train neural networks to perform classification. In at least one embodiment, one or more neural networks are trained to perform classification based on, for example, using one or more compressed representations of one or more class labels, where the one or more compressed representations have fewer bits than a representation of the one or more class labels.

Claims (40)

1 . One or more processors comprising: circuitry to train one or more neural networks using one or more binary encodings of one or more classification labels of one or more image segments, wherein the one or more binary encodings comprises fewer bits than the one or more classification labels of the one or more image segments.

2 . The one or more processors of claim 1 , wherein the one or more classification labels are segmentation labels for a three-dimensional (3D) image.

3 . The one or more processors of claim 1 , wherein the one or more binary encodings comprise fewer bits than a vector representation of the one or more classification labels, where components of the vector correspond to respective categories of labels.

4 . The one or more processors of claim 1 , wherein the one or more binary encodings are binary encoded representations of values assigned to the one or more classification labels.

5 . The one or more processors of claim 4 , wherein the one or more neural networks include a three-dimensional (3D) U-Net.

6 . The one or more processors of claim 1 , wherein the circuitry is also to train the one or more neural networks using at least one of: one or more weights, one or more activations, or one or more gradients stored in a half-precision format, and a single-precision copy of the weights to accumulate gradients after one or more optimizer steps.

7 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

train one or more neural networks using one or more binary encodings of one or more classification labels of one or more image segments, wherein the one or more binary encodings comprises fewer bits than the one or more classification labels of the one or more image segments.

8 . The non-transitory machine-readable medium of claim 7 , wherein the one or more classification labels are segmentation labels for an image.

9 . The non-transitory machine-readable medium of claim 7 , wherein the one or more binary encodings comprise fewer bits than a vector representation of the one or more classification labels, where components of the vector correspond to respective categories of labels.

10 . The non-transitory machine-readable medium of claim 7 , wherein the one or more binary encodings are binary encoded representations of values assigned to the one or more classification labels.

11 . The non-transitory machine-readable medium of claim 7 , wherein the one or more neural networks include a convolutional neural network with a contraction path and an expansion path.

12 . The non-transitory machine-readable medium of claim 7 , wherein the instructions, which if performed by the one or more processors, further cause the one or more processors to:

train the one or more neural networks using mixed-precision training.

13 . A method comprising:

training one or more neural networks using one or more binary encodings of one or more classification labels of one or more image segments, wherein each of the one or more binary encodings comprises fewer bits than the one or more binary encodings of the one or more classification labels.

14 . The method of claim 13 , wherein the one or more class labels are image segmentation class labels.

15 . The method of claim 13 , wherein the one or more classification labels are three-dimensional (3D) whole brain image segmentation classification labels.

16 . The method of claim 13 , further comprising:

assigning a set of values to a set of classes associated with a set of training data; and

converting the assigned set of values to the binary encodings of the one or more class labels.

17 . The method of claim 16 , wherein the binary encodings of the one or more classification labels are binary encoded representations of values in the set of values, and wherein assigning the set of values to the set of classes is performed such that when the set of values is converted to the binary encoded representations, a probability of each bit of the binary encoded representation being a one is within a predetermined threshold of probability of every other bit of the binary encoded representation being a one.

18 . The method of claim 17 , further comprising:

reducing a difference in probability of each bit of the binary encoded representations being a one for the set of training data.

19 . A system comprising:

one or more processors to train one or more neural networks using one or more binary encodings of one or more classification labels of one or more image segments, wherein the one or more binary encodings comprises fewer bits than the one or more classification labels of the one or more image segments; and

one or more memories to store the one or more neural networks.

20 . The system of claim 19 , wherein the one or more classification labels indicate respective objects in a three-dimensional (3D) image.

21 . The system of claim 19 , wherein the one or more binary encodings comprise fewer bits than a vector representation of the one or more classification labels, where components of the vector correspond to respective categories.

22 . The system of claim 19 , wherein the one or more binary encodings are binary encoded representations of values assigned to the one or more classification labels.

23 . The system of claim 22 , wherein the one or more neural networks include a three-dimensional (3D) convolutional neural network with a contraction path and an expansion path.

24 . The system of claim 19 , wherein the one or more processors are also to train the one or more neural networks using weights, activations, and gradients stored in a half-precision format, and a single-precision copy of the weights to accumulate gradients after one or more optimizer steps.

25 . A medical imaging system comprising:

one or more processors comprising circuitry to perform image segmentation using one or more neural networks trained, at least in part, by using one or more binary encodings of one or more classification labels of one or more image segments, wherein the one or more binary encodings fewer bits than the one or more classification labels of the one or more image segments; and

one or more memories to store an output of the performed image segmentation.

26 . The medical imaging system of claim 25 , wherein the one or more classification labels are segmentation labels for an image.

27 . The medical imaging system of claim 25 , wherein the one or more binary encodings comprise fewer bits than a vector representation of the one or more classification labels, where components of the vector correspond to respective categories of labels.

28 . The medical imaging system of claim 25 , wherein the one or more binary encodings are binary encoded representations of values assigned to the one or more classification labels.

29 . The medical imaging system of claim 28 , wherein the one or more neural networks includes a convolutional neural network with a contraction path and an expansion path.

30 . The medical imaging system of claim 25 , wherein the one or more processors are to perform image segmentation based, at least in part, on translating a set of binary encodings of classification labels output by the one or more neural networks to a set of image segmentation classes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2020
From: RIBALTA, PABLO; MARCINKIEWICZ, MICHAL; YANG, DONG; XU, DAGUANG; STRZELCZYK, PRZEMYSLAW MIROSLAW
To: NVIDIA CORPORATION
Reel/Frame 053400/0937 →
Continuity (1)
Related Publication 20210397943A1 · Dec 23, 2021
References Cited (41)
US 5701302A · Geiger · 1997 [cited by examiner]
US 20200021865A1 · Topiwala · 2020 [cited by examiner]
US 20200302289A1 · Ren · 2020 [cited by examiner]
CN 109635866A · 2019 [cited by applicant]
Myronenko “3D MRI brain tumor segmentation using autoencoder regularization”, 2018, pp. 10, arXiv:1810.11654v3. [cited by examiner]
Cicek et al. “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation”, 2016, pp. 8, arXiv:1606.06650v1. [cited by examiner]
Mishchenko et al. “Low-bit quantization and quantization-aware training for small-footprint keyword spotting”, ICMLA, 2019, p. 706--711. [cited by examiner]
Alexander et al., “Desikan-Killiany-Tourville Atlas Compatible Version of M-CRIB Neonatal Parcellated Whole Brain Atlas: The M-CRIB 2.0,” Frontiers in Neuroscience, 13(34): Feb. 5, 2019, 9 pages. [cited by applicant]
Anonymous, “U-Net—Wikipedia,” retrieved from https://en.wikipedia.org/w/index.php?title=U-Net&oldid=951207160, retrieved on Sep. 17, 2021, Apr. 16, 2020, 3 pages. [cited by applicant]
Bankier et al., “Consensus Interpretation in Imaging Research: Is There a Better Way?” Radiology, 257(1): 2010, 4 pages. [cited by applicant]
Brébisson et al., “Deep Neural Networks for Anatomical Brain Segmentation,” IEEE Conference on Computer Vision and Pattern Recognition Workshops, 2015, 9 pages. [cited by applicant]
Chen et al., “Training Deep Nets with Sublinear Memory Cost,” Apr. 22, 2016, 12 pages. [cited by applicant]
Cole et al., “Prediction of Brain Age Suggests Accelerated Atrophy After Traumatic Brain Injury,” Annals of Neurology, 77(4): 2015, 11 pages. [cited by applicant]
Coupé et al., Assemblynet: A Large Ensemble of CNNs for 3D Whole Brain MRI Segmentation, 2019, 16 pages. [cited by applicant]
Coupé et al., “Lifespan Changes of the Human Brain in Alzheimer's Disease,” Scientific Reports 9, 2019, 12 pages. [cited by applicant]
Fischl, “FreeSurfer,” Neuroimage, 62(2): Aug. 2012, 8 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” CVPR, 2016, 9 pages. [cited by applicant]
Huo et al., “Spatially Localized Atlas Network Tiles Enables 3D Whole Brain Segmentation from Limited Data,” MICCAI, Jun. 5, 2018, 8 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2021/037507, mailed Sep. 28, 2021, filed Jun. 15, 2021, 15 pages. [cited by applicant]
Jarvik et al., “Moderate Versus Mediocre: The Reliability of Spine MR Data Interpretations,” Radiology, 250(1): 2009, 3 pages. [cited by applicant]
Kaku et al., “DARTS: Denseunet-based Automatic Rapid Tool for Brain Segmentation,” Nov. 14, 2019, 27 pages. [cited by applicant]
Klein et al., “101 Labeled Brain Images and a Consistent Human Cortical Labeling Protocol,” Frontiers in Neuroscience 6, Dec. 5, 2012, 12 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” In Advances in Neural Information Processing Systems, 2012, 9 pages. [cited by applicant]
Li et al., “On the Compactness, Efficiency, and Representation of 3D Convolutional Networks: Brain Parcellation as a Pretext Task,” International Conference on Information Processing in Medical Imaging, Jul. 6, 2017, 13… [cited by applicant]
Litjens et al., “A Survey on Deep Learning in Medical Image Analysis,” Medical Image Analysis 42, 2017, 38 pages. [cited by applicant]
Mehta et al., “BrainSegNet: A Convolutional Neural Network Architecture for Automated Segmentation of Human Brain Structures,” Journal of Medical Imaging, 4(2): Apr.-Jun. 2017, 11 pages. [cited by applicant]
Micikevicius et al., “Mixed Precision Training,” Oct. 12, 2017, 14 pages. [cited by applicant]
Miotto et al., “Deep Learning for Healthcare: Review, Opportunities and Challenges,” Briefings in Bioinformatics, 19(6): 2018, 11 pages. [cited by applicant]
Roy et al., “Quicknat: A Fully Convolutional Network for Quick and Accurate Segmentation of Neuroanatomy,” NeuroImage 186, 2019, 20 pages. [cited by applicant]
Shao et al., “Microwave Imaging by Deep Learning Network: Feasibility and Training Method,” IEEE Transactions on Antennas and Propagation, 68(7): Jul. 2020, 10 pages. [cited by applicant]
Shen et al., “Deep Learning in Medical Image Analysis,” Annual Review of Biomedical Engineering 19, 2017, 30 pages. [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201609, issued Jan… [cited by applicant]
Society of Automotive Engineers On-Road Automated Vehicle Standards Committee, “Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles,” Standard No. J3016-201806, issued Jan… [cited by applicant]
Stefano et al., “Clinical Relevance of Brain Volume Measures in Multiple Sclerosis,” CNS Drugs, 28(2): Feb. 2014, 10 pages. [cited by applicant]
Wachinger et al., “DeepNAT: Deep Convolutional Neural Network for Segmenting Neuroanatomy,” NeuroImage 170, 2018, 15 pages. [cited by applicant]
Zhou et al., “A review: Deep Learning for Medical Image Segmentation using Multi-Modality Fusion,” Science Direct, 2019, 11 pages. [cited by applicant]
Çiçek et al., “3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation,” International Conference on Medical Image Computing and Computer-Assisted Intervention, 2016, 8 pages. [cited by applicant]
Narang et al., “Mixed Predecision Training,” ICLR, 2018, 12 pages. [cited by applicant]
Office Action for Chinese Application No. 202180006026.5, mailed Apr. 30, 2025, 23 pages. [cited by applicant]
Office Action for Chinese Application No. 202180006026.5, mailed Dec. 25, 2025, 19 pages. [cited by applicant]