IP Library › Granted Patent US 12,548,309
Granted Patent B2
US 12,548,309 · App. 17/569,030 · Granted Feb 10, 2026

Label inheritance for soft label generation in information processing system

Inventors: Zijia Wang (WeiFang, CN); Jiacheng Ni (Shanghai, CN); Wenbin Yang (Shanghai, CN); Kenneth Durazzo (Morgan Hill, CA); Zhen Jia (Shanghai, CN)
Assignee: Dell Products L.P.
G06V10/82G06F16/55G06F16/583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,309
App. No.
17/569,030
Filed
Jan 5, 2022
Granted
Feb 10, 2026
Kind
B2
Art Unit
2159
USPC
707/736
Abstract

Label inheritance techniques are disclosed for soft label generation in an information processing system that uses machine learning. For example, a method generates at least one label for a given data instance from a training data set useable to train a machine learning-based model. The at least one label is generated by assigning one or more labels associated with one or more ancestors of the data instance such that the data instance inherits the one or more labels associated with the one or more ancestors as the at least one label.

Claims (61)

1 . A method, comprising:

training a capsule neural network that incorporates one or more dynamic routing algorithms to estimate features of objects using a training data set comprising a plurality of image data instances to generate one or more mutually orthogonal capsules corresponding to one or more data instances of the training data set, each mutually orthogonal capsule comprising a vector representing an estimate of a local feature of a corresponding image data instance;

extracting the one or more mutually orthogonal capsules inside the trained capsule neural network;

computing a probabilistic bag of images representation for each of the one or more mutually orthogonal capsules based on a relation between each mutually orthogonal capsule and the training data set, the computing comprising:

generating a first matrix associated with the training data set and the one or more mutually orthogonal capsules;

generating a second matrix associated with the one or more mutually orthogonal capsules and corresponding dimensions of the one or more mutually orthogonal capsules; and

multiplying the first matrix and the second matrix;

identifying one or more ancestors of a given data instance based on the probabilistic bag of images representation of a given mutually orthogonal capsule associated with the given data instance;

generating at least one soft label for the given data instance from the training data set, wherein the at least one soft label is generated by:

comparing the probabilistic bag of images representation to the identified one or more ancestors of the given data instance to compute a similarity value for each of the identified one or more ancestors; and

assigning one or more labels associated with one or more ancestors of the given data instance based on the computed similarity values such that the given data instance inherits the one or more labels associated with at least one of the identified one or more ancestors of the given data instance as the at least one soft label; and

iteratively re-training the capsule neural network by inputting one or more new data instances to generate at least one or more new soft labels for the one or more new data instances, the generating comprising generating a third matrix associated with the one or more new data instances and the one or more mutually orthogonal capsules to identify one or more ancestors of each of the one or more new data instances and to compute a similarity value for each of the identified one or more ancestors of each of the one or more new data instances, the one or more new soft labels further being associated with the one or more ancestors of the one or more new data instances;

wherein the method is performed by at least one processor and at least one memory storing executable computer program instructions.

2 . The method of claim 1 , wherein the step of generating the at least one soft label for the given data instance further comprises selecting one or more similar samples from the probabilistic bag of images representation for each mutually orthogonal capsule, and assigning one or more probabilities associated with the one or more similar samples as a soft label for each mutually orthogonal capsule.

3 . The method of claim 1 , wherein each mutually orthogonal capsule represents a topic associated with the corresponding image data instance.

4 . The method of claim 1 , further comprising using each mutually orthogonal capsule as distilled data in a proxy data generation application.

5 . The method of claim 1 , further comprising using each mutually orthogonal capsule as distilled data in a neural network self-training application.

6 . The method of claim 1 , wherein identifying the one or more ancestors of the given data instance further comprises:

normalizing the probabilistic bag of images representation; and

selecting as the one or more ancestors a predetermined number of samples from the training data set having the highest computed similarity values relative to the given mutually orthogonal capsule.

7 . The method of claim 1 , wherein assigning the one or more labels comprises performing a weighted summation of labels associated with the identified one or more ancestors, wherein a weight for the label of each ancestor is based on its computed similarity value.

8 . The method of claim 1 , wherein each value in the probabilistic bag of images representation represents a similarity between a corresponding image data instance in the training data set and the given mutually orthogonal capsule.

9 . The method of claim 1 , wherein assigning the one or more labels comprises performing a weighted summation of labels associated with the identified one or more ancestors, wherein a weight for the label of each ancestor is based on its computed similarity value.

10 . The method of claim 1 , wherein each value in the probabilistic bag of images representation represents a similarity between a corresponding image data instance in the training data set and the given mutually orthogonal capsule.

11 . An apparatus, comprising:

at least one processor and at least one memory storing computer program instructions wherein, when the at least one processor executes the computer program instructions, the apparatus is configured to:

train a capsule neural network that incorporates one or more dynamic routing algorithms to estimate features of objects using a training data set comprising a plurality of image data instances to generate one or more mutually orthogonal capsules corresponding to one or more data instances of the training data set, each mutually orthogonal capsule comprising a vector representing an estimate of a local feature of a corresponding image data instance;

extract the one or more mutually orthogonal capsules inside the trained capsule neural network;

compute a probabilistic bag of images representation for each of the one or more mutually orthogonal capsules based on a relation between each mutually orthogonal capsule and the training data set, the computing comprising:

generating a first matrix associated with the training data set and the one or more mutually orthogonal capsules;

generating a second matrix associated with the one or more mutually orthogonal capsules and corresponding dimensions of the one or more mutually orthogonal capsules; and

multiplying the first matrix and the second matrix;

identify one or more ancestors of a given data instance based on the probabilistic bag of images representation of a given mutually orthogonal capsule associated with the given data instance;

generate at least one soft label for the given data instance from the training data set, wherein the at least one soft label is generated by:

comparing the probabilistic bag of images representation to the identified one or more ancestors of the given data instance to compute a similarity value for each of the identified one or more ancestors; and

assigning one or more labels associated with one or more ancestors of the given data instance based on the computed similarity values such that the given data instance inherits the one or more labels associated with at least one of the identified one or more ancestors of the given data instance as the at least one soft label; and

iteratively re-train the capsule neural network by inputting one or more new data instances to generate one or more new soft labels for the one or more new data instances, the generating comprising generating a third matrix associated with the one or more new data instances and the one or more mutually orthogonal capsules to identify one or more ancestors of each of the one or more new data instances and to compute a similarity value for each of the identified one or more ancestors of each of the one or more new data instances, the one or more new labels further being associated with the one or more ancestors of the one or more new soft data instances.

12 . The apparatus of claim 11 , wherein generating the at least one soft label for the given data instance further comprises selecting one or more similar samples from the probabilistic bag of images representation for each mutually orthogonal capsule, and assigning one or more probabilities associated with the one or more similar samples as a soft label for each mutually orthogonal capsule.

13 . The apparatus of claim 11 , wherein each mutually orthogonal capsule represents a topic associated with the corresponding image data instance.

14 . The apparatus of claim 11 , wherein the apparatus is further configured to use each mutually orthogonal capsule as distilled data in a proxy data generation application.

15 . The apparatus of claim 11 , wherein the apparatus is further configured to use each mutually orthogonal capsule as distilled data in a neural network self-training application.

16 . The apparatus of claim 11 , wherein identifying the one or more ancestors of the given data instance further comprises:

normalizing the probabilistic bag of images representation; and

selecting as the one or more ancestors a predetermined number of samples from the training data set having the highest computed similarity values relative to the given mutually orthogonal capsule.

17 . A computer program product stored on a non-transitory computer-readable medium and comprising machine executable instructions, the machine executable instructions, when executed, causing a processing device to:

train a capsule neural network that incorporates one or more dynamic routing algorithms to estimate features of objects using a training data set comprising a plurality of image data instances to generate one or more mutually orthogonal capsules corresponding to one or more data instances of the training data set, each mutually orthogonal capsule comprising a vector representing an estimate of a local feature of a corresponding image data instance;

extract the one or more mutually orthogonal capsules inside the trained capsule neural network;

compute a probabilistic bag of images representation for each of the one or more mutually orthogonal capsules based on a relation between each mutually orthogonal capsule and the training data set, the computing comprising:

generating a first matrix associated with the training data set and the one or more mutually orthogonal capsules;

generating a second matrix associated with the one or more mutually orthogonal capsules and corresponding dimensions of the one or more mutually orthogonal capsules; and

multiplying the first matrix and the second matrix;

identify one or more ancestors of a given data instance based on the probabilistic bag of images representation of a given mutually orthogonal capsule associated with the given data instance;

generate at least one soft label for the given data instance from the training data set, wherein the at least one soft label is generated by:

comparing the probabilistic bag of images representation to the identified one or more ancestors of the given data instance to compute a similarity value for each of the identified one or more ancestors; and

assigning one or more labels associated with one or more ancestors of the given data instance based on the computed similarity values such that the given data instance inherits the one or more labels associated with at least one of the identified one or more ancestors of the given data instance as the at least one soft label; and

iteratively re-train the capsule neural network by inputting one or more new data instances to generate one or more new soft labels for the one or more new data instances, the generating comprising generating a third matrix associated with the one or more new data instances and the one or more mutually orthogonal capsules to identify one or more ancestors of each of the one or more new data instances and to compute a similarity value for each of the identified one or more ancestors of each of the one or more new data instances, the one or more new soft labels further being associated with the one or more ancestors of the one or more new data instances.

18 . The computer program product of claim 17 , wherein assigning the one or more labels comprises performing a weighted summation of labels associated with the identified one or more ancestors, wherein a weight for the label of each ancestor is based on its computed similarity value.

19 . The computer program product of claim 17 , wherein each value in the probabilistic bag of images representation represents a similarity between a corresponding image data instance in the training data set and the given mutually orthogonal capsule.

20 . The computer program product of claim 17 , wherein identifying the one or more ancestors of the given data instance further comprises:

normalizing the probabilistic bag of images representation; and

selecting as the one or more ancestors a predetermined number of samples from the training data set having the highest computed similarity values relative to the given mutually orthogonal capsule.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2022
From: WANG, ZIJIA; NI, JIACHENG; YANG, WENBIN; DURAZZO, KENNETH; JIA, ZHEN
To: DELL PRODUCTS L.P.
Reel/Frame 058631/0852 →
Continuity (1)
Related Publication 20230215155A1 · Jul 6, 2023
References Cited (82)
US 8358856B2 · Loui · 2013 [cited by examiner]
US 9075824B2 · Gordo · 2015 [cited by examiner]
US 9990687B1 · Kaufhold · 2018 [cited by examiner]
US 10929757B2 · Baker · 2021 [cited by examiner]
US 11068782B2 · Lyske · 2021 [cited by examiner]
US 11461415B2 · Lu · 2022 [cited by examiner]
US 11657340B2 · Cella · 2023 [cited by examiner]
US 11684241B2 · Crosby · 2023 [cited by examiner]
US 11748414B2 · Mohanty · 2023 [cited by examiner]
US 11791914B2 · Cella · 2023 [cited by examiner]
US 11868428B2 · Rhee · 2024 [cited by examiner]
US 12165311B2 · Hao · 2024 [cited by examiner]
US 12299070B2 · Wang · 2025 [cited by examiner]
US 12327206B2 · Jia · 2025 [cited by examiner]
US 12374089B2 · Wang · 2025 [cited by examiner]
US 20060253418A1 · Charnock · 2006 [cited by examiner]
US 20090080731A1 · Krishnapuram · 2009 [cited by examiner]
US 20130290222A1 · Gordo · 2013 [cited by examiner]
US 20140143251A1 · Wang · 2014 [cited by examiner]
US 20140307958A1 · Wang · 2014 [cited by examiner]
US 20170083608A1 · Ye · 2017 [cited by examiner]
US 20180204111A1 · Zadeh · 2018 [cited by examiner]
US 20190095788A1 · Yazdani · 2019 [cited by examiner]
US 20190205748A1 · Fukuda · 2019 [cited by examiner]
US 20190213503A1 · Navratil · 2019 [cited by examiner]
US 20190303742A1 · Bonnell · 2019 [cited by examiner]
US 20190325269A1 · Bagherinezhad · 2019 [cited by examiner]
US 20190355115A1 · Niebauer · 2019 [cited by examiner]
US 20200110982A1 · Gou · 2020 [cited by examiner]
US 20200174433A1 · Hughes · 2020 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20200250971A1 · Zhao · 2020 [cited by examiner]
US 20200257976A1 · Polanía Cabrera · 2020 [cited by examiner]
US 20200401929A1 · Duerig · 2020 [cited by examiner]
US 20210034985A1 · Vongkulbhisal · 2021 [cited by examiner]
US 20210142177A1 · Mallya · 2021 [cited by examiner]
US 20210201003A1 · Banerjee · 2021 [cited by examiner]
US 20210350176A1 · Klaiman · 2021 [cited by examiner]
US 20210358101A1 · Neumann · 2021 [cited by examiner]
US 20210374504A1 · Kurasawa · 2021 [cited by examiner]
US 20210383306A1 · Somashekairah · 2021 [cited by examiner]
US 20210390270A1 · Fei · 2021 [cited by examiner]
US 20220035867A1 · Tambi · 2022 [cited by examiner]
US 20220121884A1 · Zadeh · 2022 [cited by examiner]
US 20220164714A1 · Hron, II · 2022 [cited by examiner]
US 20220180065A1 · Liu · 2022 [cited by examiner]
US 20220198274A1 · Shindin · 2022 [cited by examiner]
US 20220230425A1 · Kosiorek · 2022 [cited by examiner]
US 20220237436A1 · Shin · 2022 [cited by examiner]
US 20220237788A1 · Shaul · 2022 [cited by examiner]
US 20220254190A1 · Gallagher · 2022 [cited by examiner]
US 20220300761A1 · Zhang · 2022 [cited by examiner]
US 20220360515A1 · Vaina · 2022 [cited by examiner]
US 20220391433A1 · Maheshwari · 2022 [cited by examiner]
US 20230020886A1 · Mahapatra · 2023 [cited by examiner]
US 20230022845A1 · Meng · 2023 [cited by examiner]
US 20230083724A1 · Cella · 2023 [cited by examiner]
US 20230162005A1 · Cheng · 2023 [cited by examiner]
US 20230169331A1 · Vasilev · 2023 [cited by examiner]
US 20230215155A1 · Wang · 2023 [cited by examiner]
US 20230274422A1 · Peleg · 2023 [cited by examiner]
US 20230319099A1 · Karimibiuki · 2023 [cited by examiner]
US 20230342364A1 · Pavlovic · 2023 [cited by examiner]
US 20230394387A1 · Jia · 2023 [cited by examiner]
US 20230401274A1 · Denninghoff · 2023 [cited by examiner]
US 20240020526A1 · Bondi · 2024 [cited by examiner]
US 20240029416A1 · Wang · 2024 [cited by examiner]
US 20240037131A1 · Magureanu · 2024 [cited by examiner]
US 20240119260A1 · Wang · 2024 [cited by examiner]
US 20240185564A1 · Wang · 2024 [cited by examiner]
US 20240202494A1 · Ming Chang · 2024 [cited by examiner]
US 20240205140A1 · Lokhandwala · 2024 [cited by examiner]
US 20240338532A1 · Pauli · 2024 [cited by examiner]
US 20240419873A1 · Yu · 2024 [cited by examiner]
US 20250037429A1 · Wang · 2025 [cited by examiner]
US 20250037430A1 · Wang · 2025 [cited by examiner]
US 20250165544A1 · Shou · 2025 [cited by examiner]
E. Strubell et al., “Energy and Policy Considerations for Deep Learning in NLP,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 28-Aug. 2, 2019, pp. 3645-3650. [cited by applicant]
G. Hinton et al., “Distilling the Knowledge in a Neural Network,” arXiv:1503.02531v1, Mar. 9, 2015, 9 pages. [cited by applicant]
T. Wang et al., “Dataset Distillation,” arXiv:1811.10959v3, Feb. 24, 2020, 14 pages. [cited by applicant]
S. Sabour et al., “Dynamic Routing Between Capsules,” arXiv:1710.09829v2, Nov. 7, 2017, 11 pages. [cited by applicant]
Y. Lecun et al., “Gradient-Based Learning Applied to Document Recognition,” Proceedings of the IEEE, Nov. 1998, 46 pages. [cited by applicant]