IP Library › Granted Patent US 12,566,815
Granted Patent B2
US 12,566,815 · App. 17/527,173 · Granted Mar 3, 2026

Computer-implemented method, electronic device, and non-transitory computer-readable storage medium for context-aware classification of physiological signal data

Inventors: Sheng-Chi Huang (Taichung City, TW); Ting Yuan Wang (Taipei City, TW); Jung-Tzu Liu (Hsinchu City, TW); Ya-Wen Lee (Chiayi City, TW)
Assignee: Industrial Technology Research Institute
G06F18/213G06F18/25G06N3/045G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,815
App. No.
17/527,173
Granted
Mar 3, 2026
Kind
B2
Abstract

A computer-implemented method, an electronic device, and a non-transitory computer-readable storage medium for context-aware classification of physiological signal data. The method includes the following. A data string is obtained. The data string is fed into a first deep neural network to generate a first feature map. Multi-dimensional data is generated based on the data string. The multi-dimensional data is fed into a second deep neural network to generate a second feature map. At least the first feature map and the second feature map are fused into a specific feature vector. The specific feature vector is fed into a machine learning model. The machine learning model outputs an identification result corresponding to the data string in response to the specific feature vector.

Claims (91)

1 . A computer-implemented method for context-aware classification of physiological signal data, comprising:

obtaining a first signal stream comprising one-dimensional physiological data;

processing the first signal stream using a first neural network, to generate a first feature map;

transforming the one-dimensional physiological data of the first signal stream into a higher-dimensional representation suitable for spatial processing;

processing the higher-dimensional representation using a second neural network;

projecting the first feature map to generate a reference feature vector aligned with the second feature map;

fusing the reference feature vector and the second feature map; and

classifying to generate a semantic output of a physiological condition.

2 . The method according to claim 1 , wherein the first neural network comprises a first convolutional neural network and a recurrent neural network, and the method comprises:

feeding the first signal stream into the first convolutional neural network, wherein the first convolutional neural network generates a first spatial feature vector in response to the first signal stream; and

feeding the first spatial feature vector into the recurrent neural network, wherein the recurrent neural network generates a first temporal feature vector as the first feature map in response to the first spatial feature vector.

3 . The method according to claim 1 , wherein the second neural network comprises a second convolutional neural network, and the method comprises:

feeding the multi-dimensional data into the second convolutional neural network, wherein the second convolutional neural network generates a second spatial feature vector as the second feature map in response to the multi-dimensional data.

4 . The method according to claim 1 , wherein fusing the reference feature vector and the second feature map comprises:

feeding the first feature map into a third convolutional neural network, wherein the third convolutional neural network generates a first reference feature vector in response to the first feature map;

stacking a plurality of the first reference feature vectors into a second reference feature vector based on a size of the second feature map, wherein the size of the second feature map is equal to a size of the second reference feature vector;

transforming the second reference feature vector into a third reference feature vector; and

generating the specific feature vector based on the second feature map and the third reference feature vector.

5 . The method according to claim 4 , wherein transforming the second reference feature vector into the third reference feature vector comprises:

inputting the second reference feature vector into a Sigmoid function, wherein the Sigmoid function generates the third reference feature vector in response to the second reference feature vector, and each element in the third reference feature vector is within a value range between 0 and 1.

6 . The method according to claim 4 , wherein generating the specific feature vector based on the second feature map and the third reference feature vector comprises:

applying an attention mechanism configured to weight features in the second feature map based on corresponding values in the third reference feature vector, and generating the specific feature vector based on the attention-weighted features.

7 . The method according to claim 1 , further comprising:

processing the first signal stream using a third neural network to generate a third feature map; and wherein fusing the reference feature vector and the second feature map comprises:

fusing the first feature map, the second feature map, and the third feature map into the specific feature vector.

8 . The method according to claim 7 , wherein the first neural network comprises a convolutional neural network, and the method comprises:

processing the first signal stream using the convolutional neural network to generate a first feature map.

9 . The method according to claim 7 , wherein the third neural network comprises a recurrent neural network, and the method comprises:

processing the first signal stream using the recurrent neural network to generate a third feature map.

10 . The method according to claim 7 , wherein fusing the first feature map, the second feature map, and the third feature map to generate the specific feature vector comprises:

processing the first feature map using a fourth convolutional neural network to generate a fourth reference feature vector;

stacking a plurality of the fourth reference feature vectors into a fifth reference feature vector based on a spatial dimension of the second feature map, wherein a size of the second feature map corresponds to a size of the fifth reference feature vector;

transforming the fifth reference feature vector into a sixth reference feature vector;

processing the third feature map using a fifth convolutional neural network to generate a seventh reference feature vector;

stacking a plurality of the seventh reference feature vectors into an eighth reference feature vector based on the spatial dimension of the second feature map, wherein a size of the second feature map corresponds to a size of the eighth reference feature vector;

transforming the eighth reference feature vector into a ninth reference feature vector; and

generating the specific feature vector based on the second feature map, the sixth reference feature vector, and the ninth reference feature vector.

11 . The method according to claim 10 , wherein transforming the fifth reference feature vector into the sixth reference feature vector comprises:

applying a Sigmoid function to the fifth reference feature vector, wherein the Sigmoid function outputs the sixth reference feature vector in response to the fifth reference feature vector, wherein each element in the sixth reference feature vector is between 0 and 1; and

wherein transforming the eighth reference feature vector into the ninth reference feature vector comprises:

applying the Sigmoid function to the eighth reference feature vector, wherein the Sigmoid function outputs the ninth reference feature vector in response to the eighth reference feature vector, wherein each element in the ninth reference feature vector is between 0 and 1.

12 . The method according to claim 10 , wherein generating the specific feature vector based on the second feature map, the sixth reference feature vector, and the ninth reference feature vector comprises:

applying an attention mechanism to compute the specific feature vector from the second feature map, the sixth reference feature vector, and the ninth reference feature vector.

13 . The method according to claim 1 , wherein the higher-dimensional representation comprises a waveform image generated based on the first signal stream.

14 . An electronic device for classifying physiological signal data with contextual awareness, comprising:

a memory storing executable instructions; and

a processor configured to execute the instructions to:

receive a one-dimensional physiological signal;

process the signal using a first neural network to obtain a first feature map;

transform the one-dimensional physiological data of the first signal stream into a higher-dimensional representation suitable for spatial processing;

process the higher-dimensional data using a second neural network to obtain a second feature map;

project the first feature map to generate a reference feature vector aligned with the second feature map;

fuse the reference feature vector and the second feature map; and

classify a result to output a semantic indicator of physiological condition.

15 . The electronic device according to claim 14 , wherein the first neural network comprises a first convolutional neural network and a recurrent neural network, and the processor is configured to:

feed the first signal stream into the first convolutional neural network, wherein the first convolutional neural network generates a first spatial feature vector in response to the first signal stream; and

feed the first spatial feature vector into the recurrent neural network, wherein the recurrent neural network generates a first temporal feature vector as the first feature map in response to the first spatial feature vector.

16 . The electronic device according to claim 14 , wherein the second neural network comprises a second convolutional neural network, and the processor is configured to:

feed the multi-dimensional data into the second convolutional neural network, wherein the second convolutional neural network generates a second spatial feature vector as the second feature map in response to the multi-dimensional data.

17 . The electronic device according to claim 14 , wherein the processor is configured to:

feed the first feature map into a third convolutional neural network, wherein the third convolutional neural network generates a first reference feature vector in response to the first feature map;

stack a plurality of the first reference feature vectors into a second reference feature vector based on a size of the second feature map, wherein the size of the second feature map is equal to a size of the second reference feature vector;

transform the second reference feature vector into a third reference feature vector; and

generate the specific feature vector based on the second feature map and the third reference feature vector.

18 . The electronic device according to claim 17 , wherein the processor is configured to:

input the second reference feature vector into a Sigmoid function, wherein the Sigmoid function generates the third reference feature vector in response to the second reference feature vector, and each element in the third reference feature vector is within a value range between 0 and 1.

19 . The electronic device according to claim 17 , wherein generating the specific feature vector based on the second feature map and the third reference feature vector comprises:

applying an attention mechanism configured to weight features in the second feature map based on corresponding values in the third reference feature vector, and generating the specific feature vector based on the attention-weighted features.

20 . The electronic device according to claim 14 , the processor is further configured to:

process the first signal stream using a third neural network to generate a third feature map; and wherein fusing the reference feature vector and the second feature map comprises:

fusing the first feature map, the second feature map, and the third feature map into the specific feature vector.

21 . The electronic device according to claim 20 , wherein the first neural network comprises a convolutional neural network, and the processor is configured to:

process the first signal stream using the convolutional neural network to generate a first feature map.

22 . The electronic device according to claim 20 , wherein the third neural network comprises a recurrent neural network, and the processor is configured to:

process the first signal stream using the recurrent neural network to generate a third feature map.

23 . The electronic device according to claim 20 , wherein the processor is configured to:

process the first feature map using a fourth convolutional neural network to generate a fourth reference feature vector;

stack a plurality of the fourth reference feature vectors into a fifth reference feature vector based on a spatial dimension of the second feature map, wherein a size of the second feature map corresponds to a size of the fifth reference feature vector;

transform the fifth reference feature vector into a sixth reference feature vector;

process the third feature map using a fifth convolutional neural network to generate a seventh reference feature vector;

stack a plurality of the seventh reference feature vectors into an eighth reference feature vector based on the spatial dimension of the second feature map, wherein a size of the second feature map corresponds to a size of the eighth reference feature vector;

transform the eighth reference feature vector into a ninth reference feature vector; and

generate the specific feature vector based on the second feature map, the sixth reference feature vector, and the ninth reference feature vector.

24 . The electronic device according to claim 23 , wherein the processor is configured to:

apply a Sigmoid function to the fifth reference feature vector, wherein the Sigmoid function outputs the sixth reference feature vector in response to the fifth reference feature vector, wherein each element in the sixth reference feature vector is between 0 and 1; and

wherein transforming the eighth reference feature vector into the ninth reference feature vector comprises:

applying the Sigmoid function to the eighth reference feature vector, wherein the Sigmoid function outputs the ninth reference feature vector in response to the eighth reference feature vector, wherein each element in the ninth reference feature vector is between 0 and 1.

25 . The electronic device according to claim 23 , wherein the processor is configured to:

apply an attention mechanism to compute the specific feature vector from the second feature map, the sixth reference feature vector, and the ninth reference feature vector.

26 . The electronic device according to claim 14 , wherein the higher-dimensional representation comprises a waveform image generated based on the first signal stream.

27 . A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause a computing system to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 26, 2021
From: HUANG, SHENG-CHI; WANG, TING YUAN; LIU, JUNG-TZU; LEE, YA-WEN
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 058211/0718 →
Priority Claims (1)
TW 110138083 · Oct 14, 2021 · national
Continuity (1)
Related Publication 20230122473A1 · Apr 20, 2023
References Cited (75)
US 6735579B1 · Woodall · 2004 [cited by applicant]
US 8311618B2 · Vajdic et al. · 2012 [cited by applicant]
US 8655817B2 · Hasey et al. · 2014 [cited by applicant]
US 10998101B1 · Tran et al. · 2021 [cited by applicant]
US 11568991B1 · Jain · 2023 [cited by examiner]
US 11829413B1 · Hao · 2023 [cited by examiner]
US 20200075167A1 · Srivastava et al. · 2020 [cited by applicant]
US 20200090028A1 · Huang · 2020 [cited by examiner]
US 20200159603A1 · Panda · 2020 [cited by examiner]
US 20200327308A1 · Cheng · 2020 [cited by examiner]
US 20210056413A1 · Cheung · 2021 [cited by examiner]
US 20210100471A1 · Yu · 2021 [cited by examiner]
US 20210118566A1 · Wang · 2021 [cited by examiner]
US 20210142497A1 · Pugh · 2021 [cited by examiner]
US 20220139066A1 · Yang · 2022 [cited by examiner]
US 20220146707A1 · Kim · 2022 [cited by examiner]
US 20220175287A1 · Li · 2022 [cited by examiner]
US 20220318553A1 · Ben Yahia · 2022 [cited by examiner]
US 20220387115A1 · Barbagli · 2022 [cited by examiner]
US 20230018194A1 · Huang · 2023 [cited by examiner]
US 20230046274A1 · Chen · 2023 [cited by examiner]
US 20230089026A1 · Tran · 2023 [cited by examiner]
US 20230134967A1 · Francesca · 2023 [cited by examiner]
US 20230293079A1 · Chen · 2023 [cited by examiner]
US 20240212843A1 · Kim · 2024 [cited by examiner]
CN 105426889 · 2016 [cited by applicant]
CN 109271975 · 2019 [cited by applicant]
CN 110162638 · 2019 [cited by applicant]
CN 110352430 · 2019 [cited by applicant]
CN 110969073 · 2020 [cited by applicant]
CN 109109909 · 2020 [cited by applicant]
CN 111443165 · 2020 [cited by applicant]
CN 111582223 · 2020 [cited by applicant]
CN 112270337 · 2021 [cited by applicant]
CN 112668559 · 2021 [cited by applicant]
TW 201227595 · 2012 [cited by applicant]
TW I718422 · 2021 [cited by applicant]
Cheng Dai et al., “Human action recognition using two-stream attention based LSTM networks”, Applied Soft Computing, vol. 86, Jan. 2020, pp. 1-8. [cited by applicant]
Yuanyao Lu et al., “Automatic Lip-Reading System Based on Deep Convolutional Neural Network and Attention-Based Long Short-Term Memory”, Appl. Sci., vol. 9, Issue 8, Apr. 2019, pp. 1-12. [cited by applicant]
Dandan Liang et al., “Deep convolutional BiLSTM fusion network for facial expression recognition”, The Visual Computer, vol. 36, Feb. 2019 pp. 499-508. [cited by applicant]
Shiyao Chen et al., “A Novel Attention Cooperative Framework for Automatic Modulation Recognition”, IEEE Access, vol. 8, Jan. 2020, pp. 1-14. [cited by applicant]
Erting Pan et al., “Spectral-spatial classification for hyperspectral image based on a single GRU”, Neurocomputing, vol. 387, Apr. 2020, pp. 150-160. [cited by applicant]
Dongli Wang et al., “Human action recognition based on multi-mode spatial-temporal feature fusion”, 2019 22th International Conference on Information Fusion (FUSION), Jul. 2019, pp. 1-7. [cited by applicant]
Ruiying Lu et al., “RAFnet: Recurrent attention fusion network of hyperspectral and multispectral images”, Signal Processing, vol. 177, Dec. 2020, pp. 1-16. [cited by applicant]
Luntian Mou et al., “Driver stress detection via multimodal fusion using attention-based CNN-LSTM”, Expert Systems with Applications, vol. 173, Jul. 2021, pp. 1-10. [cited by applicant]
Rohan Banerjee et al., “A Hybrid CNN-LSTM Architecture for Detection of Coronary Artery Disease from ECG”, Applied Intelligence, Aug. 2021, pp. 1-8. [cited by applicant]
Lidan Fu et al., “Hybrid Network with Attention Mechanism for Detection and Location of Myocardial Infarction Based on 12-Lead Electrocardiogram Signals”, Sensors, vol. 20, Issue 4, Feb. 2020 , pp. 1-24. [cited by applicant]
Dae Ha Kim et al., “Multi-modal emotion recognition using semi-supervised learning and multiple neural networks in the wild”, 19th ACM International Conference on Multimodal Interaction, Nov. 2017, pp. 529-535. [cited by applicant]
Mian Pan et al., “Radar HRRP Target Recognition Model Based on a Stacked CNN-Bi-RNN With Attention Mechanism”, IEEE Transactions on Geoscience and Remote Sensing, Feb. 2021, pp. 1-14. [cited by applicant]
Dian Yu et al., “A Systematic Exploration of Deep Neural Networks for EDA-Based Emotion Recognition”, Information, vol. 11, Issue 4, Apr. 2020, pp. 1-16. [cited by applicant]
Wei-Xun Zhang et al., “Signal-3L 3.0: Improving Signal Peptide Prediction through Combining Attention Deep Learning with Window-Based Scoring”, J. Chem. Inf. Model., vol. 60, Issue 7, Jun. 2020, pp. 3679-3686. [cited by applicant]
Ziping Zhao et al., “Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion Recognition”, Interspeech 2018, Sep. 2018, pp. 272-276. [cited by applicant]
Sen Jia et al., “Cascade Superpixel Regularized Gabor Feature Fusion for Hyperspectral Image Classification”, IEEE Transactions on Neural Networks and Learning Systems , vol. 31, Issue 5, May 2020, pp. 1638-1652. [cited by applicant]
Kia Dashtipour et al., “A hybrid Persian sentiment analysis framework: Integrating dependency grammar based rules and deep neural networks”, Neurocomputing, vol. 380, Mar. 2020, pp. 1-10. [cited by applicant]
Pei Li et al., “Real-time crash risk prediction on arterials based on LSTM-CNN”, Accident Analysis & Prevention, vol. 135, Feb. 2020, pp. 1-9. [cited by applicant]
Aite Zhao et al., “A hybrid spatio-temporal model for detection and severity rating of Parkinson's disease from gait data”, Neurocomputing, vol. 315, Nov. 2018, pp. 1-8. [cited by applicant]
Hao Sun et al., “Spectral-Spatial Attention Network for Hyperspectral Image Classification”, IEEE Transactions on Geoscience and Remote Sensing, vol. 58, Issue 5, May 2020, pp. 3232-3245. [cited by applicant]
Wang Xiaohua et al., “Two-level attention with two-stage multi-task learning for facial emotion recognition”, Journal of Visual Communication and Image Representation, vol. 62, Jul. 2019, pp. 217-225. [cited by applicant]
Ganchao Bao et al., abstract of “Fault diagnosis of reciprocating compressor based on group self-attention network”, Measurement Science and Technology, vol. 31, Issue 6, Apr. 2020, pp. 1-4. [cited by applicant]
Zeeshan Ahmad et al., “CNN-Based Multistage Gated Average Fusion (MGAF) for Human Action Recognition Using Depth and Inertial Sensors”, IEEE Sensors Journal, vol. 62, Jul. 2019, pp. 1-12. [cited by applicant]
Kaijun Zhu et al., “A Cuboid CNN Model With an Attention Mechanism for Skeleton-Based Action Recognition”, IEEE Transactions on Multimedia, vol. 22, Issue 11, Nov. 2020, pp. 2977-2989. [cited by applicant]
Hao Liu et al., “Sequence-based Person Attribute Recognition with Joint CTC-Attention Model”, Computer Vision and Pattern Recognition, arXiv:1811.08115, Nov. 2018, pp. 1-9. [cited by applicant]
Lei Zhang et al., “Deep learning for sentiment analysis: A survey”, Wires, vol. 8, Issue 4, Mar. 2018, pp. 1-34. [cited by applicant]
Ruixi Zhu et al., “Attention-Based Deep Feature Fusion for the Scene Classification of High-Resolution Remote Sensing Images”, Remote Sensing, vol. 11, Aug. 2019, pp. 1-23. [cited by applicant]
Haiman Tian et al., abstract of “Multimodal deep representation learning for video classification”, World Wide Web, vol. 22, May 2018, pp. 1-15. [cited by applicant]
Xuanhan Wang et al., “Two-Stream 3-D convNet Fusion for Action Recognition in Videos With Arbitrary Size and Length”, IEEE Transactions on Multimedia , vol. 20, Issue 3, Mar. 2018, pp. 634-644. [cited by applicant]
Cheng Dai et al., “Human action recognition using two-stream attention based LSTM networks”, Applied Soft Computing, vol. 86, Jan. 2020, pp. 1-2. [cited by applicant]
Luntian Mou et al., “Driver stress detection via multimodal fusion using attention-based CNN-LSTM”, Expert Systems with Applications, vol. 173, Jul. 2021, pp. 502-506. [cited by applicant]
Hari Mohan Rai et al., “A Hybrid CNN-LSTM Architecture for Detection of Coronary Artery Disease from ECG”, Applied Intelligence, Aug. 2021, pp. 1-8. [cited by applicant]
Dian Yu et al., “A Systematic Exploration of Deep Neural Networks for EDA-Based Emotion Recognition”, Information, vol. 11, Issue 4, Apr. 2020, pp. 502-506. [cited by applicant]
Ziping Zhao et al., “Exploring Spatio-Temporal Representations by Integrating Attention-based Bidirectional-LSTM-RNNs and FCNs for Speech Emotion Recognition”, Interspeech 2018, Sep. 2018, pp. 270-276. [cited by applicant]
Ganchao Bao et al., “Fault diagnosis of reciprocating compressor based on group self-attention network”, Measurement Science and Technology, vol. 31, Issue 6, Apr. 2020, pp. 1-4. [cited by applicant]
Kaijun Zhu et al., “A Cuboid CNN Model With an Attention Mechanism for Skeleton-Based Action Recognition”, IEEE Transactions on Multimedia, vol. 22, Issue 11, Nov. 2020, pp. 1-13. [cited by applicant]
Haiman Tian et al., “Multimodal deep representation learning for video classification”, World Wide Web, vol. 22, May 2018, pp. 1325-1341. [cited by applicant]
“Office Action of Taiwan Counterpart Application”, issued on Aug. 1, 2022, p. 1-p. 5. [cited by applicant]