IP Library Granted Patent US 12,314,236
Granted Patent B2
US 12,314,236 · App. 17/599,200 · Granted May 27, 2025

Automatic generation of labeled data in IoT systems

Inventors: Quang Ly (North Wales, PA); Lu Liu (Conshohocken, PA); Dale N. Seed (Allentown, PA); Zhuo Chen (Claymont, DE); William Robert Flynn, IV (Schwenksville, PA); Catalina Mihaela Mladin (Hatboro, PA); Jiwan L. Ninglekhu (Royersford, PA); Hongkun Li (Malvern, PA); Rocco Di Girolamo (Laval, CA)
Assignee: Convida Wireless, LLC
G06F16/215G06F16/2365
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,236
App. No.
17/599,200
Granted
May 27, 2025
Kind
B2
Abstract

A labeled data generation service provides an Internet-of-Things (IoT) system with a capability whereby users may configure how the system gathers, processes, and generates labeled data instances by: collecting and processing the data into a format required by supervised learning algorithms; generating expected outputs from data available in the IoT system; supporting the linking of collected inputs with generated expected outputs; forming labeled data instances; cleaning the labeled data set appropriately; sending the labeled data set to target nodes; and/or communicating with target nodes regarding improving the data processing and labeling processes, as required.

Claims (48)

1. An apparatus comprising one or more processors and memory storing computer-executable instructions which, when executed by the one or more processors of the apparatus, cause the apparatus to perform operations comprising:

maintaining a configuration, the configuration comprising design information for a labeled data set, the labeled data set comprising a plurality of labeled data instances, wherein each labeled data instance comprises a plurality of data values, the data values comprising one or more data inputs and one or more expected data outputs associated with the one or more data inputs;

acquiring a plurality of raw data inputs from data sources;

processing, according to the configuration, the raw data inputs to create processed data inputs, wherein the processing of the raw data inputs comprises pre-processing based on first parameters indicated in the configuration and data transformation based on second parameters indicated in the configuration;

generating, according to the configuration, labeled data instances, wherein a labeled data instance comprises one or more processed data inputs and one or more expected data output values;

storing the labeled instances in a labeled data set; and

sending the labeled data set to a repository.

2. The apparatus of claim 1 , wherein, for one or more raw data inputs, the processing of the raw data inputs comprises converting or scaling each of the raw data inputs to create a processed data input value for each raw data input.

3. The apparatus of claim 1 , wherein, for one or more raw data inputs, the processing of the raw data inputs comprises scaling a processed data input value for each of the raw data inputs in accordance with one or more statistical observations of the plurality of raw data inputs.

4. The apparatus of claim 1 , wherein, for one or more sets of raw data inputs, the processing of the raw data inputs comprises deriving a processed data input value for each plurality of raw data inputs in accordance with one or more statistical observations of the plurality of raw data inputs.

5. The apparatus of claim 1 , wherein the labeled data instances are generated with data cleaning based on one or more cleaning rules indicated by the configuration.

6. The apparatus of claim 5 , wherein the data cleaning comprises one or more of:

identifying duplicate labeled data instances in the labeled data set;

removing the identified duplicate labeled data instances from the labeled data set;

verifying data is valid from the labeled data set;

monitoring for mandatory data in the labeled data set;

detecting conflicts with data instances in the labeled data set; and

informing the repository of the identified duplicate labeled data instances.

7. The apparatus of claim 1 , wherein:

the configuration comprises an output time requirement parameter; and

the operations further comprise acquiring an expected data output in accordance with the output time requirement parameter.

8. The apparatus of claim 1 , wherein the pre-processing comprises one or more of: measurement unit conversion, data type conversion, or data aggregation.

9. The apparatus of claim 8 , wherein the data aggregation comprises one or more of: a sum, an average, a minimum, a maximum, or a count.

10. The apparatus of claim 1 , wherein the data transformation comprises one or more of: normalization, standardization, or binning.

11. A method comprising:

maintaining a configuration, the configuration comprising design information for a labeled data set, the labeled data set comprising a plurality of labeled data instances, wherein each labeled data instance comprises a plurality of data values, the data values comprising one or more data inputs and one or more expected data outputs associated with the one or more data inputs;

acquiring a plurality of raw data inputs from data sources;

processing, according to the configuration, the raw data inputs to create processed data inputs, wherein the processing of the raw data inputs comprises pre-processing based on first parameters indicated in the configuration and data transformation based on second parameters indicated in the configuration;

generating, according to the configuration, labeled data instances, wherein a labeled data instance comprises one or more processed data inputs and one or more expected data output values;

storing the labeled data instances in a labeled data set; and

sending the labeled data set to a repository.

12. The method of claim 11 , wherein, for one or more raw data inputs, the processing of the raw data inputs comprises converting or scaling each of the raw data inputs to create a processed data input value for each raw data input.

13. The method of claim 11 , wherein, for one or more raw data inputs, the processing of the raw data inputs comprises scaling a processed data input value for each of the raw data inputs in accordance with one or more statistical observations of the plurality of raw data inputs.

14. The method of claim 11 , wherein, for one or more sets of raw data inputs, the processing of the raw data inputs comprises deriving a processed data input value for each plurality of raw data inputs in accordance with one or more statistical observations of the plurality of raw data inputs.

15. The method of claim 11 , wherein the labeled data instances are generated with data cleaning based on one or more cleaning rules indicated by the configuration.

16. The method of claim 15 , wherein the data cleaning comprises one or more of:

identifying duplicate labeled data instances in the labeled data set;

removing the identified duplicate labeled data instances from the labeled data set;

verifying data is valid from the labeled data set;

monitoring for mandatory data in the labeled data set;

detecting conflicts with data instances in the labeled data set; and

informing the repository of the identified duplicate labeled data instances.

17. The method of claim 11 , wherein:

the configuration comprises an output time requirement parameter; and

the operations further comprise acquiring an expected data output in accordance with the output time requirement parameter.

18. The method of claim 11 , wherein the pre-processing comprises one or more of: measurement unit conversion, data type conversion, or data aggregation.

19. The method of claim 18 , wherein the data aggregation comprises one or more of: a sum, an average, a minimum, a maximum, or a count.

20. The method of claim 11 , wherein the data transformation comprises one or more of: normalization, standardization, or binning.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2025
From: CONVIDA WIRELESS,LLC
To: IPLA HOLDINGS INC.
Reel/Frame 073903/0733 →
CORRECTIVE ASSIGNMENT TO CORRECT THE SECOND ASSIGNOR'S NAME PREVIOUSLY RECORDED AT REEL: 057625 FRAME: 0826. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT . Recorded May 16, 2022
From: LY, QUANG; LIU, LU; SEED, DALE N.; CHEN, ZHUO; FLYNN, WILLIAM ROBERT, IV; MLADIN, CATALINA MIHAELA; NINGLEKHU, JIWAN L.; LI, HONGKUN; DI GIROLAMO, ROCCO
To: CONVIDA WIRELESS, LLC
Reel/Frame 060072/0335 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2021
From: LY, QUANG; LU, LIU; SEED, DALE N.; CHEN, ZHUO; FLYNN, WILLIAM ROBERT, IV; MLADIN, CATALINA MIHAELA; NINGLEKHU, JIWAN L.; LI, HONGKUN; DI GIROLAMO, ROCCO
To: CONVIDA WIRELESS, LLC
Reel/Frame 057625/0826 →
Continuity (2)
Provisional Application 62827475 · Apr 1, 2019
Related Publication 20220156235A1 · May 19, 2022
References Cited (22)
US 20130262406A1 · Kumar · 2013 [cited by examiner]
US 20160170980A1 · Stadnisky · 2016 [cited by examiner]
US 20170032277A1 · Klinger · 2017 [cited by examiner]
US 20170032281A1 · Hsu · 2017 [cited by applicant]
US 20170359362A1 · Kashi et al. · 2017 [cited by applicant]
US 20180314914A1 · Kuriyama · 2018 [cited by examiner]
US 20190370687A1 · Pezzillo · 2019 [cited by examiner]
US 20190378619A1 · Meyer · 2019 [cited by examiner]
US 20200158810A1 · Zhang · 2020 [cited by examiner]
US 20220101337A1 · Kashibuchi · 2022 [cited by examiner]
CN 107809766A · 2018 [cited by applicant]
CN 108027911A · 2018 [cited by applicant]
CN 109328448A · 2019 [cited by applicant]
Eklund, Martin. “Comparing feature extraction methods and effects of pre-processing methods for multi-label classification of textual data.” (2018). [cited by examiner]
Kotsiantis, Sotiris B., Dimitris Kanellopoulos, and Panagiotis E. Pintelas. “Data preprocessing for supervised leaning.” International journal of computer science 1.2 (2006): 111-117. [cited by examiner]
Maharana, Kiran, Surajit Mondal, and Bhushankumar Nemade. “A review: Data pre-processing and data augmentation techniques.” Global Transitions Proceedings 3.1 (2022): 91-99. [cited by examiner]
Daniel Gutierrez, “Labeled Training Sets for Machine Learning”, insideBIGDATA, Mar. 7, 2019, 4 pages. [cited by applicant]
Mahdavinejad et al., “Machine learning for Internet of Things data analysis: A survey”, Digital Communications and Networks, 2017, 56 pages. [cited by applicant]
Sezer et al., “Context-Aware Computing, Learning, and Big Data in Internet of Things: A Survey”, IEEE Internet of Things Journal, vol. 5(1), Feb. 2018, pp. 1-27. [cited by applicant]
Wikipedia et al., “Labeled data”, available online at <https://en.wikipedia.org/w/index.php?t%20itle=Labeled%20data&oldid=857071684>, Aug. 29, 2018, 1 page. [cited by applicant]
OneM2M TS-0001, Functional Architecture, V3.12.0, 2018. [cited by applicant]
“Introduction to Machine Learning Contents”, Retrieved from https://www.datascienceassn.org/sites/default/files/Introduction%20to%20Machine%20Learning.pdf, Jul. 13, 2015, pp. 1-146. [cited by applicant]