IP Library Granted Patent US 12,307,821
Granted Patent B2
US 12,307,821 · App. 17/886,264 · Granted May 20, 2025

Radar-based gesture classification using a variational auto-encoder neural network

Inventors: Avik Santra (Munich, DE); Souvik Hazra (Munich, DE); Thomas Reinhold Stadelmayer (Wenzenbach, DE)
Assignee: Infineon Technologies AG
G06V40/20G06V10/761G06V10/763G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,821
App. No.
17/886,264
Granted
May 20, 2025
Kind
B2
Abstract

In an embodiment, a method includes: obtaining one or more positional time spectrograms of a radar measurement of a scene comprising an object; and based on the one or more positional time spectrograms and based on a feature embedding of a variational auto-encoder neural network, predicting a gesture class of a gesture performed by the object.

Claims (36)

1. A method comprising:

obtaining one or more positional time spectrograms of a radar measurement of a scene comprising an object; and

based on the one or more positional time spectrograms and based on a feature embedding of a variational auto-encoder neural network, predicting a gesture class of a gesture performed by the object, wherein the gesture class is predicted based on a comparison of a mean of a distribution of the feature embedding of the variational auto-encoder neural network with one or more regions predefined in a feature space of the feature embedding.

2. The method of claim 1 , further comprising:

monitoring a cluster formation of the means of the distribution of the feature embedding of the variational auto-encoder neural network obtained for multiple sets of the one or more positional time spectrograms, the cluster formation being outside of the one or more predefined regions; and

based on the monitoring of the cluster formation, determining a further predefined region in the feature space to enclose a respective cluster.

3. The method of claim 1 , wherein the one or more positional time spectrograms are obtained by time gating measurement data of the radar measurement based on at least one trigger event.

4. The method of claim 3 , wherein the at least one trigger event comprises a comparison between a change rate of a positional observable captured by the measurement data and at least one predefined threshold.

5. The method of claim 3 , wherein the at least one trigger event comprises an output of a gesture detection algorithm.

6. The method of claim 1 , wherein the one or more positional time spectrograms are selected from the group consisting of: a range time spectrogram, a velocity time spectrogram, an azimuthal angle time spectrogram, and an elevation angle time spectrogram.

7. The method of claim 1 , wherein the one or more positional time spectrograms comprise one or more raw positional time spectrograms, and wherein the variational auto-encoder neural network has been trained to reconstruct one or more filtered positional time spectrograms.

8. A method for training a variational auto-encoder neural network for predicting a gesture class of a gesture performed by an object of a scene, the gesture class being selected from a plurality of gesture classes, the method comprising:

obtaining multiple training sets of one or more training positional time spectrograms of a radar measurement of the scene comprising the object, each one of the multiple training sets being associated with a respective ground-truth label indicative of the respective gesture class;

training the variational auto-encoder neural network based on the multiple training sets and the associated ground-truth labels; and

based on class distributions of a feature embedding of the variational auto-encoder neural network obtained for the multiple training sets associated with each one of the plurality of gesture classes, determining predefined regions in a feature space of the feature embedding for gesture class prediction.

9. The method of claim 8 , wherein the training of the variational auto-encoder neural network uses at least one loss that is determined based on at least one statistical distance between a distribution of a feature embedding of the variational auto-encoder neural network obtained for a first training set of the multiple training sets that is associated with a first gesture class of the plurality of gesture classes, and at least one mean of at least one distribution of a feature embedding of the variational auto-encoder neural network obtained for at least one second training set of the multiple training sets that is associated with at least one of the first gesture class or a second gesture class of the plurality of gesture classes.

10. The method of claim 9 , wherein the at least one loss comprises a statistical distance triplet loss determined based on a first statistical distance and a second statistical distance, the first statistical distance being between the distribution of the feature embedding of the variational auto-encoder neural network obtained for an anchor training set of the multiple training sets and the mean of the distribution of the feature embedding of the variational auto-encoder neural network obtained for a positive training set of the multiple training sets, the second statistical distance being between the distribution of the feature embedding of the variational auto-encoder neural network obtained for the anchor training set and the mean of the distribution of the feature embedding of the variational auto-encoder neural network obtained for a negative training set of the multiple training sets.

11. The method of claim 9 , wherein the at least one loss comprises a statistical distance center loss determined based on a statistical distance between a class distribution associated with the first gesture class and means of the distributions of the feature embedding of the variational auto-encoder neural network obtained for all training sets of the multiple training sets associated with the first gesture class.

12. The method of claim 9 , wherein the statistical distance is a Mahalanobis distance.

13. The method of claim 8 , wherein the one or more training positional time spectrograms comprise one or more raw training positional time spectrograms, the method further comprising:

applying an unscented Kalman filter to the one or more raw training positional time spectrograms to obtain one or more filtered training positional time spectrograms, wherein the training of the variational auto-encoder neural network uses at least one reconstruction loss which is based on a difference between a reconstruction of the one or more raw training positional time spectrograms output by the variational auto-encoder neural network and the one or more filtered training positional time spectrograms.

14. A radar system comprising:

a millimeter-wave radar sensor configured to transmit radar signals towards a scene, receive reflected radar signals from the scene, and generate radar measurement data based on the reflected radar signals; and

a processor configured to:

generate one or more positional time spectrograms based on the radar measurement data, and

based on the one or more positional time spectrograms and based on a feature embedding of a variational auto-encoder neural network, predict a gesture class of a gesture performed by an object in the scene, wherein the gesture class is predicted based on a comparison of a mean of a distribution of the feature embedding of the variational auto-encoder neural network with one or more regions predefined in a feature space of the feature embedding.

15. The radar system of claim 14 , wherein the processor is further configured to:

monitor a cluster formation of the mean of the distributions of the feature embedding of the variational auto-encoder neural network obtained for multiple sets of the one or more positional time spectrograms, the cluster formation being outside of the one or more predefined regions; and

based on the monitoring of the cluster formation, determine a further predefined region in the feature space to enclose a respective cluster.

16. The radar system of claim 14 , wherein the one or more positional time spectrograms are obtained by time gating measurement data of a radar measurement based on at least one trigger event, and wherein the at least one trigger event comprises a comparison between a change rate of a positional observable captured by the radar measurement data and at least one predefined threshold.

17. The radar system of claim 14 , wherein the one or more positional time spectrograms comprise at least one of: a range time spectrogram, a velocity time spectrogram, an azimuthal angle time spectrogram, and an elevation angle time spectrogram.

18. The radar system of claim 14 , wherein the processor comprises:

at least one first processor; and

at least one memory with instructions stored thereon, wherein the instructions, when executed by the at least one first processor, enable the processor to generate the one or more positional time spectrograms, and predict the gesture class of the gesture performed by the object in the scene.

19. The radar system of claim 14 , wherein the one or more positional time spectrograms are obtained by time gating measurement data of a radar measurement based on at least one trigger event, and the at least one trigger event comprises an output of a gesture detection algorithm.

20. The radar system of claim 14 , wherein the one or more positional time spectrograms comprise one or more raw positional time spectrograms, and wherein the variational auto-encoder neural network has been trained to reconstruct one or more filtered positional time spectrograms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 19, 2022
From: SANTRA, AVIK; HAZRA, SOUVIK; STADELMAYER, THOMAS REINHOLD
To: INFINEON TECHNOLOGIES AG
Reel/Frame 061284/0423 →
Priority Claims (1)
EP 21190926 · Aug 12, 2021 · regional
Continuity (1)
Related Publication 20230068523A1 · Mar 2, 2023
References Cited (41)
US 10205457B1 · Josefsberg · 2019 [cited by examiner]
US 10466772B2 · Trotta · 2019 [cited by examiner]
US 10699538B2 · Novich · 2020 [cited by examiner]
US 10937438B2 · Narayanan · 2021 [cited by examiner]
US 11036303B2 · Rani · 2021 [cited by examiner]
US 11354823B2 · Lerchner · 2022 [cited by examiner]
US 11432753B2 · Alam · 2022 [cited by examiner]
US 11475910B2 · Jun · 2022 [cited by examiner]
US 11640208B2 · Hazra · 2023 [cited by examiner]
US 11774553B2 · Santra · 2023 [cited by examiner]
US 20210034160A1 · Hof · 2021 [cited by examiner]
US 20230068523A1 · Santra · 2023 [cited by examiner]
Ahmed, S. et al., “Finger-Counting-Based Gesture Recognition within Cars Using Impulse Radar with Convolutional Neural Network,” sensors, MDPI, Mar. 23, 2019, 14 pages. [cited by applicant]
Deng, J. et al., “ArcFace: Additive Angular Margin Loss for Deep Face Recognition,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jan. 9, 2020, 10 pages. [cited by applicant]
Dubey, A. et al., “A Bayesian Framework for Integrated Deep Metric Learning and Tracking of Vulnerable Road Users Using Automotive Radars”, IEEE Acess, May 5, 2021, 20 pages. [cited by applicant]
Erol, B. et al., “Exploitation of motion capture data for improved synthetic micro-doppler signature generation with adversarial learning”, SPIE Proceedings, May 18, 2020, 8 pages. [cited by applicant]
Gurbuz, S. et al., “Radar-Based Human-Motion Recognition With Deep Learning: Promising applications for indoor monitoring”, IEEE Signal Processing Magazine, Jul. 2019, 13 pages. [cited by applicant]
Hazra, S. et al., “Robust Gesture Recognition Using Millimetric-Wave Radar System,” IEEE Sensors Letters, vol. 2, No. 4, Dec. 2018, 4 pages. [cited by applicant]
Hazra, S. et al., “Short-Range Radar-Based Gesture Recognition System Using 3D CNN With Triplet Loss,” IEEE Access, Sep. 17, 2019, 11 pages. [cited by applicant]
He, L. et al., “Softmax Dissection: Towards Understanding Intra- and Inter-Class Objective for Embedding Learning,” The Thirty-Fourth AAAI Conference on Artificial Intelligence (AAAI-20), Feb. 12, 2020, 8 pages. [cited by applicant]
Kim, Y. et al., “Hand Gesture Recognition Using Micro-Doppler Signatures With Convolutional Neural Network,” IEEE Access, Oct. 13, 2016, 6 pages. [cited by applicant]
Kingma, D. et al., “Auto-Encoding Variational Bayes,” arXiv:1312.6114v10, May 1, 2014, 14 pages. [cited by applicant]
Li, G. et al., “Sparsity-Driven Micro-Doppler Feature Extraction for Dynamic Hand Gesture Recognition,” IEEE Transactions on Aerospace and Electronic Systems, Apr. 2018, 21 pages. [cited by applicant]
Lien, J. et al., “Soli: Ubiquitous Gesture Sensing with Millimeter Wave Radar,” ACM Trans. Graph., vol. 35, No. 4, Article 142, Jul. 2016, 19 pages. [cited by applicant]
Liu, W. et al., “SphereFace: Deep Hypersphere Embedding for Face Recognition,” Computer Vision Foundation, Apr. 26, 2017, 9 pages. [cited by applicant]
Molchanov, P. et al., “Online Detection and Classification of Dynamic Hand Gestures with Recurrent 3D Convolutional Neural Networks,” Computer Vision Foundation, Dec. 12, 2016, 9 pages. [cited by applicant]
Molchanov, P. et al., “Short-Range FMCW Monopulse Radar for Hand-Gesture Sensing,” 2015 IEEE Radar Conference (RadarCon), Jun. 25, 2015, 6 pages. [cited by applicant]
Rautaray, S. et al., “Vision based hand gesture recognition for human computer interaction: a survey,” Artificial Intelligence Review, Nov. 6, 2012, 54 pages. [cited by applicant]
Smith, K. et al., “Gesture Recognition Using mm-Wave Sensor for Human-Car Interface,” ResearchGate, Sensors Letters, Feb. 2018, 5 pages. [cited by applicant]
Stadelmayer, T. et al., “Human Activity Classification Using mm-Wave FMCW Radar by Improved Representation Learning,” mmNets'20: Proceedings of the 4th ACM Workshop on Millimeter-Wave Networks and Sensing Systems, Sep. … [cited by applicant]
Sun, Y. et al., “Automatic Radar-based Gesture Detection and Classification via a Region-based Deep Convolutional Neural Network,” ResearchGate, May 2019, 7 pages. [cited by applicant]
Sun, Y. et al., “Real-Time Radar-Based Gesture Detection and Recognition Built in an Edge-Computing Platform,” EEE Sensors Journal, arXiv:2005.10145v1, May 20, 2020, 10 pages. [cited by applicant]
Wan, Q. et al., “Gesture Recognition for Smart Home Applications using Portable Radar Sensors,” ResearchGate, 36th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, Aug. 2014, 5 pa… [cited by applicant]
Wang, F. et al., “Additive Margin Softmax for Face Verification,” IEEE Signal Processing Letters, arXiv:1801.05599v4, May 30, 2018, 7 pages. [cited by applicant]
Wang, H. et al., “CosFace: Large Margin Cosine Loss for Deep Face Recognition,” CVPR, IEEE, Dec. 16, 2018, 10 pages. [cited by applicant]
Wang, S. et al., “Interacting with Soli: Exploring Fine-Grained Dynamic Gesture Recognition in the Radio-Frequency Spectrum,” UIST '16: Proceedings of the 29th Annual Symposium on User Interface Software and Technology,… [cited by applicant]
Wen, Y. et al., “A Discriminative Feature Learning Approach for Deep Face Recognition,” Lecture Notes in Computer Science, Sep. 16, 2016, 17 pages. [cited by applicant]
Wikipedia, “Electromagnetic coil,” https://en.wikipedia.org/w/index.php?title=Electromagnetic_coil&oldid=776415501, Jul. 11, 2018, 6 pages. [cited by applicant]
Wikipedia, “Qi (standard),” https://en.wikipedia.org/w/index.php?title=Qi_(standard)&oldid=803427516, Jul. 11, 2018, 5 pages. [cited by applicant]
Zhang, J. et al., “Doppler-Radar Based Hand Gesture Recognition System Using Convolutional Neural Networks”, arXiv;1711.02254v3, Nov. 22, 2017, 8 pages. [cited by applicant]
Zhang, Z. et al., “Latern: Dynamic Continuous Hand Gesture Recognition Using FMCW Radar Sensor,” IEEE Sensors Journal, vol. 18, No. 8, Apr. 15, 2018, 12 pages. [cited by applicant]