IP Library › Granted Patent US 12,705,871
Granted Patent B2
US 12,705,871 · App. 18/272,849 · Granted Aug 11, 2026

Extracting features from sensor data

Inventors: John Redford (Cambridge, GB); Sina Samangooei (Cambridge, GB); Anuj Sharma (Cambridge, GB); Puneet Dokania (Cambridge, GB)
Assignee: Five AI Limited
G06V10/82G06V10/7715
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,871
App. No.
18/272,849
Filed
Jul 18, 2023
Granted
Aug 11, 2026
Kind
B2
Examiner
LIU, XIAO
Art Unit
2664
USPC
382/155
Abstract

A computer implemented method of training an encoder to extract features from sensor data comprises training a machine learning (ML) system based on a self-supervised loss function applied to a training set, the ML system comprising the encoder. The training set comprises first data representations and corresponding second data representations, wherein the encoder extracts features from each first and second data representation, and wherein the self-supervised loss function encourages the ML system to associate each first data representation with its corresponding second data representation based on their respective features. Each first data representation and its corresponding second data representation represent a common set of sensor data, and at least the second data representation is generated by: applying a 2D object detector to an image other than the first and second data representations, wherein the image contains or is associated with the common set of sensor data, and transforming the common set of sensor data based one or more objects detected in the image, the second data representation representing the transformed sensor data.

Claims (33)

1 . A computer implemented method comprising:

training a machine learning (ML) system based on a self-supervised loss function applied to a training set, the ML system comprising an encoder, wherein the training set comprises first data representations and corresponding second data representations, wherein the encoder extracts features from each first and second data representation, and wherein the self-supervised loss function is configured to cause the ML system to learn feature representations such that each first data representation is more similar to its corresponding second data representation, based on their respective features, extracted by the encoder, than to non-corresponding data representations, wherein each first data representation and its corresponding second data representation represent a common set of sensor data; and

generating at least the second data representation, including

applying a two-dimensional (2D) object detector to an image other than the first and second data representations, wherein the image contains or is associated with the common set of sensor data, and

transforming the common set of sensor data based on one or more objects detected in the image, the common set of sensor data comprising point cloud data or other non-image data, the second data representation representing the transformed common set of sensor data, wherein the common set of sensor data comprises a point cloud encoded in a depth channel of the image and thus represented in a 2D image plane of the image, wherein the first and second data representations represent the point cloud in a 2D plane other than the 2D image plane of the image, and wherein the 2D plane is a bird's-eye view plane lying substantially perpendicular to the 2D image plane.

2 . The method of claim 1 , wherein the first and second data representations are discretised image representation of the point cloud in the 2D plane that optionally include respective height channels.

3 . The method of claim 1 , wherein the common set of sensor data comprises the point cloud encoded in the depth channel of the image and thus represented in the 2D image plane of the image, wherein the first and second data representations represent the point cloud in three-dimensional (3D) space.

4 . The method of claim 3 , wherein the first and second data representations are discretised voxel representations of the point cloud in the 3D space, or non-discretised representations of the point cloud in the 3D space.

5 . The method of claim 1 , wherein the image has been captured substantially simultaneously with the common set of sensor data, the sensor data of a non-image modality; wherein each detected object is matched with a corresponding subset of the common set of sensor data in order to transform the common set of sensor data.

6 . The method of claim 5 , wherein the common set of sensor data comprises a point cloud not encoded in the image.

7 . The method of claim 6 , wherein the point cloud has the non-image modality.

8 . The method of claim 1 , wherein the common set of sensor data is transformed by removing or distorting background sensor data that does not belong to any detected object.

9 . The method of claim 8 , wherein the 2D object detector computes a 2D bounding box for each detected object, wherein the background sensor data is identified as sensor data contained in or associated with a background region of the image outside of any 2D bounding box.

10 . The method of claim 8 , wherein the background sensor data is fully or partially removed and replaced with random noise.

11 . The method of claim 1 , wherein the ML system comprises a trainable projection component which projects the features from a feature space into a projection space, the self-supervised loss defined on the projected features, wherein the trainable projection component is trained simultaneously with the encoder.

12 . The method of claim 1 , wherein each set of sensor data captures a static or dynamic driving scene.

13 . The method of claim 1 , wherein the common set of sensor data comprises:

three-dimensional (3D) spatial data, or

2D spatial data in the 2D plane other than an image plane of the image.

14 . The method of claim 1 , wherein the 2D object detector is a trained machine learning (ML) 2D object detector, whereby knowledge learned in the training of the 2D ML object detector is transferred to the encoder during the training based on the self-supervised loss function.

15 . A computer system comprising:

at least one memory configured to store computer-readable instructions;

at least one hardware processor coupled to the at least one memory and configured to execute the computer-readable instructions, which upon execution cause the at least one hardware processor to perform operations including:

training a machine learning (ML) system based on a self-supervised loss function applied to a training set, the ML system comprising an encoder, wherein the training set comprises first data representations and corresponding second data representations, wherein the encoder is configured to extract features from each first and second data representation, and wherein the self-supervised loss function is configured to cause the ML system to learn feature representations such that each first data representation is more similar to its corresponding second data representation, based on their respective features extracted by the encoder, than to non-corresponding data representations, wherein each first data representation and its corresponding second data representation represent a common set of sensor data; and

generating at least the second data representation, including

applying a two-dimensional (2D) object detector to an image other than the first and second data representations, wherein the image contains or is associated with the common set of sensor data, and

transforming the common set of sensor data based one or more objects detected in the image, the common set of sensor data comprising point cloud data or other non-image data, the second data representation representing the transformed common set of sensor data, wherein the common set of sensor data comprises a point cloud encoded in a depth channel of the image and thus represented in a 2D image plane of the image, wherein the first and second data representations represent the point cloud in a 2D plane other than the image plane of the image, and wherein the 2D plane is a bird's-eye view plane lying substantially perpendicular to the 2D image plane.

16 . The computer system of claim 15 , wherein the at least one hardware processor is configured to implement a perception component, wherein the encoder is configured to receive an input sensor data representation and extract features therefrom, and wherein the perception component is configured to use the extracted features to interpret the input sensor data representation.

17 . A non-transitory medium embodying computer-readable instructions configured, when executed on one or more hardware processors, to cause the one or more hardware processors to perform operations including:

training a machine learning (ML) system based on a self-supervised loss function applied to a training set, the ML system comprising an encoder, wherein the training set comprises first data representations and corresponding second data representations, wherein the encoder extracts features from each first and second data representation, and wherein the self-supervised loss function encourages the ML system to associate each first data representation with its corresponding second data representation based on their respective features, wherein each first data representation and its corresponding second data representation represent a common set of sensor data; and

generating at least the second data representation including

applying a two-dimensional (2D) object detector to an image other than the first and second data representations, wherein the image contains or is associated with the common set of sensor data, and

transforming the common set of sensor data based one or more objects detected in the image, the common set of sensor data comprising point cloud data or other non-image data, the second data representation representing the transformed common set of sensor data, wherein the common set of sensor data comprises a point cloud encoded in a depth channel of the image and thus represented in a 2D image plane of the image, wherein the first and second data representations represent the point cloud in a 2D plane other than the image plane of the image, and wherein the 2D plane is a bird's-eye view plane lying substantially perpendicular to the 2D image plane.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2026
From: REDFORD, JOHN; SAMANGOOEI, SINA; SHARMA, ANUJ; DOKANIA, PUNEET
To: FIVE AI LIMITED
Reel/Frame 074569/0132 →
Priority Claims (1)
GB 2100740 · Jan 20, 2021 · national
Continuity (1)
Related Publication 20240104913A1 · Mar 28, 2024
References Cited (37)
US 11481626B2 · Das et al. · 2022 [cited by applicant]
US 11494597B2 · Nadamuni Raghavan et al. · 2022 [cited by applicant]
US 11663486B2 · Sun et al. · 2023 [cited by applicant]
US 20170304732A1 · Velic · 2017 [cited by examiner]
US 20190383904A1 · Harrison · 2019 [cited by applicant]
US 20210012166A1 · Braley · 2021 [cited by examiner]
US 20210096241A1 · Bongio Karrman · 2021 [cited by examiner]
US 20210326751A1 · Liu et al. · 2021 [cited by applicant]
US 20210327029A1 · Chen · 2021 [cited by examiner]
US 20210406674A1 · Wu · 2021 [cited by examiner]
US 20230214654A1 · Park et al. · 2023 [cited by applicant]
US 20250013912A1 · Jansson Minne et al. · 2025 [cited by applicant]
WO 2019178702A1 · 2019 [cited by applicant]
Beker et al, Monocular Differentiable Rendering for Self-Supervised 3D Object Detection, arXiv:2009.14524v1 (Year: 2020). [cited by examiner]
Cosma et al, Self-Supervised Representation Learning on Document Images, DAS 2020: IAPR International Workshop on Document Analysis Systems (Year: 2020). [cited by examiner]
Chen et al, Multi-View 3D Object Detection Network for Autonomous Driving, CVPR (Year: 2017). [cited by examiner]
Yang et al, PIXOR: Real-time 3D Object Detection from Point Clouds, CVPR (Year: 2018). [cited by examiner]
Zakharov et al, Autolabeling 3D Objects with Differentiable Rendering of SDF Shape Priors. In CVPR (Year: 2020). [cited by examiner]
European Examination Report from related European Application No. 22704296.7, dated Jun. 28, 2024 (7 pages). [cited by applicant]
Comer Joseph F et al: “SAR automatic target recognition with less labels”, SPIE Proceedings [Proceedings of SPIE ISSN 0277-786X], SPIE, US, vol. 11394, Apr. 24, 2020 (Apr. 24, 2020), pp. 113940Q-113940Q, XP060132528, DO… [cited by applicant]
U.S. Appl. No. 18/272,950, filed Jul. 18, 2023, Redford et al. [cited by applicant]
U.S. Appl. No. 18/272,916, filed Jul. 18, 2023, Redford et al. [cited by applicant]
U.S. Appl. No. 18/272,970, filed Jul. 18, 2023, Redford et al. [cited by applicant]
International Search Report; PCT/EP2022/051145; Date: May 16, 2022; By: Authorized Officer: Grigorescu, Simona. [cited by applicant]
Xie Saining et al, “PointContrast: Unsupervised Pre-training for 3D Point Cloud Understanding”, Dec. 3, 2020 (Dec. 3, 2020), Computer Vision—ECCV 2020 : 16th European Conference, Glasgow, UK, Aug. 23-28, 2020 : Proceedi… [cited by applicant]
Bozorgtabar Behzad et al, “SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-Motion”, 2019 IEEE/CVf International Conference on Computer Vision (ICCV), IEEE,Oct. 27, 2019 (Oct. 27, 2019), p… [cited by applicant]
Eric Tzeng et al, “Adapting Deep Visuomotor Representations with Weak Pairwise Constraints”, May 25, 2017 (May 25, 2017), URL:https://arxiv.org/pdf/1511.07111.pdf. [cited by applicant]
Bin Cheng et al, “S^3Net: Semantic-Aware Self-supervised Depth Estimation with Monocular Videos and Synthetic Data”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Jul. 29, 2… [cited by applicant]
Ting Chen et al, “A Simple Framework for Contrastive Learning of Visual Representations”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Jul. 1, 2020 (Jul. 1, 2020). [cited by applicant]
Zhang Bin et al, “Multi-Task Deep Transfer Learning Method for Guided Wave-Based Integrated Health Monitoring Using Piezoelectric Transducers”, Jul. 21, 2020 (Jul. 21, 2020), vol. 20, No. 23, p. 14391-14400. [cited by applicant]
Wenhao Wang et al, “DomainMix: Learning Generalizable Person Re-Identification Without Human Annotations”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Nov. 24, 2020 (Nov. … [cited by applicant]
U.S Office Action dated Jul. 11, 2025, from related U.S. Appl. No. 18/272,970. [cited by applicant]
U.S Office Action dated Oct. 23, 2025, from related U.S. Appl. No. 18/272,916. [cited by applicant]
Yingzi Ma, Self-Supervised Learning of 3D Point Clouds via Feature Transformation and Rotation Prediction, Nov. 2022, 2022 International Conference on Electrical, Computer, Communications and Mechatronics Engineering (Y… [cited by applicant]
Qiangeng Xu et. al., Grid-GCN for Fast and Scalable Point Cloud Learning, 2020, Proceedings ofthe lEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5661-5670 (Year: 2020). [cited by applicant]
Michael Goersele et. al., Ambient Point Clouds for View Interpolation, Jul. 2010, SIGGRAPH'10: ACM SIGGRAPH 2010 papers, article No. 95, pp. 1-6 (Year: 2010). [cited by applicant]
Haoming Lu et. al., Deep Learning for 3D Point Cloud Understanding: A Survey, Sep. 2020, Computer Vision and Pattern Recognition, Machine Learning (Year: 2020). [cited by applicant]