IP Library Granted Patent US 12,412,091
Granted Patent B2
US 12,412,091 · App. 17/781,827 · Granted Sep 9, 2025

Physics-guided deep multimodal embeddings for task-specific data exploitation

Inventors: Han-Pang Chiu (West Windsor, NJ); Zachary Seymour (Pennington, NJ); Niluthpol C. Mithun (Lawrenceville, NJ); Supun Samarasekera (Skillman, NJ); Rakesh Kumar (West Windsor, NJ); Yi Yao (Princeton, NJ)
Assignee: SRI International
G06N3/08G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,412,091
App. No.
17/781,827
Granted
Sep 9, 2025
Kind
B2
Abstract

A method, apparatus and system for object detection in sensor data having at least two modalities using a common embedding space includes creating first modality vector representations of features of sensor data having a first modality and second modality vector representations of features of sensor data having a second modality, projecting the first and second modality vector representations into the common embedding space such that related embedded modality vectors are closer together in the common embedding space than unrelated modality vectors, combining the projected first and second modality vector representations, and determining a similarity between the combined modality vector representations and respective embedded vector representations of features of objects in the common embedding space to identify at least one object depicted by the captured sensor data. In some instances, data manipulation of the method, apparatus and system can be guided by physics properties of a sensor and/or sensor data.

Claims (40)

1. A method for training a common embedding space for combining sensor data captured from a common scene having at least two modalities, the method comprising:

for each of a plurality of the captured sensor data having a first modality of the at least two modalities, creating respective first modality sensor-data vector representations of features of the sensor data having the first modality using a sensor data-specific neural network;

for each of a plurality of the captured sensor data having a second modality of the at least two modalities, creating respective second modality sensor-data vector representations of the features of the sensor data having the second modality using a sensor data-specific neural network;

embedding the first modality sensor-data vector representations and the second modality sensor-data vector representations in a common embedding space such that embedded modality vectors that are related, across modalities, are closer together in the common embedding space than unrelated modality vectors; and

respectively combining the embedded first modality sensor-data vector representations and the second modality vector representations;

wherein at least one of the creating of the first and second modality sensor-data vector representations and the embedding of the first and the second modality sensor-data vector representations are constrained by physics properties of at least one of a respective sensor having captured the first modality sensor data and the second modality sensor data, to prevent captured sensor data having the first modality and sensor data having the second modality captured using physics properties not in compliance with the physics properties of the at least one of the respective sensor from being used to train the common embeddings space, such that the common embedding space is trained using only sensor data having the first modality and sensor data having the second modality captured in compliance with the physics properties of the at least one of the respective sensor.

2. The method of claim 1 , wherein a sensor data-specific neural network is pretrained to recognize features of sensor data having a modality to which the sensor data-specific neural network is to be applied.

3. The method of claim 1 , wherein the first modality sensor-data vector representations and the second modality sensor-data vector representations are combined using late fusion.

4. The method of claim 1 , further comprising:

determining a difference between the plurality of the captured sensor data having the first modality and the captured sensor data having the second modality of the at least two modalities.

5. The method of claim 4 , wherein the determined difference between the captured sensor data having the first modality and the second modality is used to determine missing data of one of the first modality or the second modality from captured data of the other one of the second modality or the first modality.

6. The method of claim 4 , wherein the difference is determined using a generative adversarial network.

7. The method of claim 1 , comprising:

determining a contribution of each of the embedded first modality sensor-data vector representations and the second modality vector representations to the combination.

8. The method of claim 1 , wherein the physics properties comprise at least one of surface reflection, temperature, or humidity.

9. A method for at least one of object detection, object classification, or object segmentation in sensor data having at least two modalities using a common embedding space, comprising:

creating respective first modality sensor-data vector representations of features of sensor data having a first modality of the at least two modalities;

creating respective second modality sensor-data vector representations of features of sensor data having a second modality of the at least two modalities;

projecting the first modality sensor-data vector representations and the second modality sensor-data vector representations into the common embedding space such that embedded modality vectors that are related, across modalities, are closer together in the common embedding space than unrelated modality vectors;

combining the projected first modality sensor-data vector representations and the second modality sensor-data vector representations;

determining a similarity between the combined modality sensor-data vector representations and respective embedded vector representations of features of objects in the common embedding space using a distance function, to identify at least one object depicted by the sensor data having the at least two modalities;

wherein at least one of the creating of the first and second modality sensor-data vector representations and the embedding of the first and the second modality sensor-data vector representations are constrained by physics properties of at least one of a respective sensor having captured the first modality sensor data and the second modality sensor data, to prevent captured sensor data having the first modality and sensor data having the second modality captured using physics properties not in compliance with the physics properties of the at least one of the respective sensor from being used to train the common embeddings space, such that the common embedding space is trained using only sensor data having the first modality and sensor data having the second modality captured in compliance with the physics properties of the at least one of the respective sensor.

10. The method of claim 9 , comprising:

determining a difference between the plurality of the sensor data having the first modality and the sensor data having the second modality of the at least two modalities.

11. The method of claim 10 , wherein at least one of the first modality sensor-data vector representations and the second modality sensor-data vector representations are created using the determined difference between the plurality of the sensor data having the first modality and the sensor data having the second modality.

12. The method of claim 9 , wherein at least one of the first modality sensor-data vector representations and the second modality sensor-data vector representations are created using a sensor data-specific neural network.

13. The method of claim 9 , wherein a contribution of each of the embedded first modality sensor-data vector representations and the second modality vector representations to the combination is predetermined.

14. The method 13 , wherein the first modality sensor-data vector representations and the second modality sensor-data vector representations are combined using attention-based mode fusion.

15. An apparatus for object detection in sensor data having at least two modalities using a common embedding space, comprising:

at least one feature extraction module configured to create respective first modality sensor-data vector representations of features of sensor data having a first modality of the at least two modalities and respective second modality sensor-data vector representations of features of sensor data having a second modality of the at least two modalities;

at least one embedding module configured to project the first modality sensor-data vector representations and the second modality sensor-data vector representations into the common embedding space such that embedded modality vectors that are related, across modalities, are closer together in the common embedding space than unrelated modality vectors;

a fusion module configured to combine the projected first modality sensor-data vector representations and the second modality sensor-data vector representations;

an inference module configured to determine a similarity between the combined modality sensor-data vector representations and respective embedded vector representations of features of objects in the common embedding space using a distance function to identify at least one object depicted by the sensor data having the at least two modalities;

wherein at least one of the creating of the first and second modality sensor-data vector representations and the embedding of the first and the second modality sensor-data vector representations are constrained by physics properties of at least one of a respective sensor having captured the first modality sensor data and the second modality sensor data, to prevent captured sensor data having the first modality and sensor data having the second modality captured using physics properties not in compliance with the physics properties of the at least one of the respective sensor from being used to train the common embeddings space, such that the common embedding space is trained using only sensor data having the first modality and sensor data having the second modality captured in compliance with the physics properties of the at least one of the respective sensor.

16. The apparatus of claim 15 , further comprising:

a generative adversarial network configured to determine a difference between the plurality of the sensor data having the first modality and the sensor data having the second modality of the at least two modalities.

17. The apparatus of claim 16 , wherein the generative adversarial network uses the determined difference between the sensor data having the first modality and the second modality to determine missing data of one of the first modality or the second modality from data of the other one of the second modality or the first modality.

18. The apparatus of claim 15 , wherein the fusion module is configured to determine a contribution of each of the projected first modality sensor-data vector representations and the second modality sensor-data vector representations of the at least two modalities to the combination.

19. The apparatus of claim 18 , wherein the fusion module is configured to apply attention-based mode fusion to combine the first modality sensor-data vector representations and the second modality sensor-data vector representations.

20. The apparatus of claim 15 , wherein the physics properties comprise at least one of surface reflection, temperature, or humidity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2022
From: CHIU, HAN-PANG; SEYMOUR, ZACHARY; MITHUN, NILUTHPOL C.; SAMARASEKERA, SUPUN; KUMAR, RAKESH; YAO, YI
To: SRI INTERNATIONAL
Reel/Frame 061133/0308 →
Continuity (2)
Provisional Application 62987697 · Mar 10, 2020
Related Publication 20230004797A1 · Jan 5, 2023
References Cited (30)
US 20110170781A1 · Bronstein et al. · 2011 [cited by applicant]
US 20180284758A1 · Cella · 2018 [cited by examiner]
US 20180375743A1 · Lee et al. · 2018 [cited by applicant]
US 20190135300A1 · Gonzalez Aguirre et al. · 2019 [cited by applicant]
US 20190197400A1 · Zhang et al. · 2019 [cited by applicant]
US 20190293462A1 · Choi · 2019 [cited by examiner]
US 20190318040A1 · Chaudhury et al. · 2019 [cited by applicant]
US 20190325342A1 · Sikka et al. · 2019 [cited by applicant]
US 20200018852A1 · Walls · 2020 [cited by examiner]
JP 2014512897A · 2014 [cited by applicant]
JP 2019535063A · 2019 [cited by applicant]
KR 1020170098573A · 2017 [cited by applicant]
WO WO2018104563A2 · 2018 [cited by applicant]
WO WO2018124309A1 · 2018 [cited by applicant]
WO WO2019010137A1 · 2019 [cited by applicant]
WO WO2019016968A1 · 2019 [cited by applicant]
WO WO2019057954A1 · 2019 [cited by applicant]
WO WO2019220622A1 · 2019 [cited by examiner]
WO WO2019231624A2 · 2019 [cited by applicant]
WO WO2019049856A1 · 2020 [cited by applicant]
Huang, Z., et al, Multi-Modal Sensor Fusion-Based Deep Neural Network for End-to-End Autonomous Driving With Scene Understanding, [received Jul. 12, 2024]. Retrieved from Internet: <https://ieeexplore.ieee.org/abstract/… [cited by examiner]
Hang, Z. et al, Multi-Modal Sensor Fusion-Based Deep Neural Network for End-to-End Autonomous Driving With Scene Understanding, [received Jul. 12, 2024]. Retrieved from Internet:<https://ieeexplore.ieee.org/abstract/doc… [cited by examiner]
Roheda, S., et al, Robust Multi-Modal Sensor Fusion: An Adversarial Approach, [received [Jul. 12, 2024]. Retrieved from Internet:<https://ieeexplore.ieee.org/abstract/document/9174786> (Year: 2020). [cited by examiner]
Priyasad, D., et al, Attention Driven Fusion for Multi-Modal Emotion Recognition, [received Jul. 12, 2024]. Retrieved from Internet:<https://ieeexplore.ieee.org/abstract/document/9054441> (Year: 2020). [cited by examiner]
Chen, C., et al, Selective Sensor Fusion for Neural Visual-Inertial Odometry, [received Nov. 25, 2024]. Retrieved from Internet:<https://openaccess.thecvf.com/content_CVPR_2019/html/Chen_Selective_Sensor_Fusion_for_Neur… [cited by examiner]
Yufu Qu et al., “Active Multimodal Sensor System for Target Recognition and Tracking”, Sensors, vol. 7, Jun. 28, 2017. [cited by applicant]
Ying Zhang et al., “Deep Cross-Modal Projection Learning for Image-Text Matching”, Computer Vision—ECCV 2018, Oct. 6, 2018. [cited by applicant]
International Search Report for application No. PCT/US2021/017731 dated Jun. 9, 2021. [cited by applicant]
Valentin Vielzeuf et al., “Multilevel Sensor Fusion With Deep Learning” Sensors Letter, Jan. 2019, vol. 3, No. 1, Nov. 7, 2018, pp. 1-12. [cited by applicant]
Yufu Qu et al., “Active Multimodal Sensor System for Target Recognition and Tracking”, Sensors, Jun. 28, 2017, vol. 17, pp. 1-22. [cited by applicant]
Cited By (1)
US 12,531,136