IP Library Granted Patent US 12,242,542
Granted Patent B2
US 12,242,542 · App. 17/408,852 · Granted Mar 4, 2025

Ordinal time series classification with missing information

Inventors: Cristian Lumezanu (Princeton Junction, NJ); Yuncong Chen (Plainsboro, NJ); Takehiko Mizoguchi (West Windsor, NJ); Dongjin Song (Princeton, NJ); Haifeng Chen (West Windsor, NJ); Jurijs Nazarovs (Madison, WI)
Assignee: NEC Corporation
G06F16/906
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,242,542
App. No.
17/408,852
Granted
Mar 4, 2025
Kind
B2
Abstract

A method classifies missing labels. The method computes, using a neural network model trained on training data, rank-based statistics of a feature of a time series segment to attempt to select two candidate labels from the training data that the segment most likely belongs to. The method classifies the segment using k-NN-based classification applied to the training data, responsive to the two candidate labels being present in the training data. The method classifies the segment by hypothesis testing, responsive to only one candidate label being present in the training data. The method classifies the segment into a class with higher values of the rank-based statistics from among a plurality of classes with different values of the rank-based statistics, responsive to no candidate labels being present in the training data. The method corrects a prediction by an applicable one of the classifying steps by majority voting with time windows.

Claims (40)

1. A computer-implemented method for time series classification of missing labels, comprising:

extracting a feature of an incoming time series segment to be classified during an inference stage;

computing, by a hardware processor using a neural network model trained on training data, rank-based statistics of the feature to attempt to select two candidate labels from the training data that the incoming time series segment most likely belongs to;

classifying the incoming time series segment using k-NN-based classification applied to the training data, responsive to the two candidate labels being present in the training data;

classifying the incoming time series segment by hypothesis testing, responsive to only one of the two candidate labels being present in the training data;

classifying the incoming time series segment into a class with higher values of the rank-based statistics from among a plurality of classes with different values of the rank-based statistics, responsive to none of the two candidate labels being present in the training data; and

correcting a prediction by an applicable one of the classifying steps by majority voting with time windows.

2. The computer-implemented method of claim 1 , further comprising training a neural network model on the training data based on ordinal quadruplet loss.

3. The computer-implemented method of claim 2 , wherein the ordinal quadruplet loss comprises a similarity component, a discrimination component, and a feature order component with respect to a feature space.

4. The computer-implemented method of claim 2 , wherein the ordinal quadruplet loss comprises a triplet loss and a log ratio loss.

5. The computer-implemented method of claim 1 , further comprising computing, for each of the two candidate labels a label retrieval vector from the each of the two candidate labels to each of present labels of the original data.

6. The computer-implemented method of claim 1 , further comprising computing, for each incoming times series to be classified during the inference stage, a test retrieval vector from each of the two candidate labels to a feature center of each of present labels in a feature space.

7. The computer-implemented method of claim 6 , wherein the feature center is computed as an average of features of incoming time series segments having a same label.

8. The computer-implemented method of claim 1 , wherein the hypothesis testing comprises determining outliers for a missing class using statistical methods.

9. The computer-implemented method of claim 8 , wherein determining outliers comprises determining a distribution of distances from the training data in present classes to a class center of an evaluated class with respect to a threshold distance.

10. The computer-implemented method of claim 1 , wherein said detecting step comprises making label predictions within a window of time using a majority voting scheme that selects a most predicted class during the window of time.

11. The computer-implemented method of claim 1 , further comprising replacing an impending failing workplace machine with a backup workplace machine responsive to the prediction to avoid system downtime.

12. A computer program product for time series classification of missing labels, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

extracting, by a hardware processor, a feature of an incoming time series segment to be classified during an inference stage;

computing, by the hardware processor using a neural network model trained on training data, rank-based statistics of the feature to attempt to select two candidate labels from the training data that the incoming time series segment most likely belongs to;

classifying, by the hardware processor, the incoming time series segment using k-NN-based classification applied to the training data, responsive to the two candidate labels being present in the training data;

classifying the incoming time series segment by hypothesis testing, responsive to only one of the two candidate labels being present in the training data;

classifying the incoming time series segment into a class with higher values of the rank-based statistics from among a plurality of classes with different values of the rank-based statistics, responsive to none of the two candidate labels being present in the training data; and

correcting a prediction by an applicable one of the classifying steps by majority voting with time windows.

13. The computer program product of claim 12 , wherein the method further comprises training a neural network model on the training data based on ordinal quadruplet loss.

14. The computer program product of claim 13 , wherein the ordinal quadruplet loss comprises a similarity component, a discrimination component, and a feature order component with respect to a feature space.

15. The computer program product of claim 13 , wherein the ordinal quadruplet loss comprises a triplet loss and a log ratio loss.

16. The computer program product of claim 12 , wherein the method further comprises computing, for each of the two candidate labels a label retrieval vector from the each of the two candidate labels to each of present labels of the original data.

17. The computer program product of claim 12 , wherein the method further comprises computing, for each incoming times series to be classified during the inference stage, a test retrieval vector from each of the two candidate labels to a feature center of each of present labels in a feature space.

18. The computer program product of claim 17 , wherein the feature center is computed as an average of features of incoming time series segments having a same label.

19. The computer program product of claim 12 , wherein the hypothesis testing comprises determining outliers for a missing class using statistical methods.

20. A computer processing system for time series classification of missing labels, comprising:

a memory device for storing program code; and

a hardware processor operatively coupled to the memory device for storing the program code to

extract a feature of an incoming time series segment to be classified during an inference stage;

compute, using a neural network model trained on training data, rank-based statistics of the feature to attempt to select two candidate labels from the training data that the incoming time series segment most likely belongs to;

classify the incoming time series segment using k-NN-based classification applied to the training data, responsive to the two candidate labels being present in the training data;

classify the incoming time series segment by hypothesis testing, responsive to only one of the two candidate labels being present in the training data;

classify the incoming time series segment into a class with higher values of the rank-based statistics from among a plurality of classes with different values of the rank-based statistics, responsive to none of the two candidate labels being present in the training data; and

correct a prediction by an applicable one of the classifying steps by majority voting with time windows.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2024
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 069540/0269 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2021
From: LUMEZANU, CRISTIAN; CHEN, YUNCONG; MIZOGUCHI, TAKEHIKO; SONG, DONGJIN; CHEN, HAIFENG; NAZAROVS, JURIJS
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 057255/0937 →
Continuity (2)
Provisional Application 63075859 · Sep 9, 2020
Related Publication 20220075822A1 · Mar 10, 2022
References Cited (14)
US 10489457B1 · Wulf · 2019 [cited by examiner]
US 11675926B2 · Muffat · 2023 [cited by examiner]
US 20190325060A1 · Fenoglio · 2019 [cited by examiner]
US 20190354809A1 · Ralhan · 2019 [cited by examiner]
US 20200317365A1 · Sherry · 2020 [cited by examiner]
US 20200409339A1 · Arashanipalai · 2020 [cited by examiner]
US 20210235293A1 · Chen · 2021 [cited by examiner]
US 20210294818A1 · Savalle · 2021 [cited by examiner]
US 20220070537A1 · Younessian · 2022 [cited by examiner]
US 20220070975A1 · Chen · 2022 [cited by examiner]
US 20220075822A1 · Lumezanu · 2022 [cited by examiner]
Song, Dongjin, et al. “Deep r-th root of rank supervised joint binary embedding for multivariate time series retrieval”, InProceedings of the 24th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining… [cited by applicant]
Diaz, Raul, et al. “Soft labels for ordinal regression”, InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Jun. 2019, pp. 4738-4747. [cited by applicant]
Wu, Peter, et al.. “Ordinal Triplet Loss: Investigating Sleepiness Detection from Speech”, In Interspeech 2019. Sep. 2019, pp. 2403-2407. [cited by applicant]