IP Library › Granted Patent US 12,462,532
Granted Patent B2
US 12,462,532 · App. 17/929,405 · Granted Nov 4, 2025

Enriching later-in-time feature maps using earlier-in-time feature maps

Inventors: Balaji Sundareshan (Boston, MA); Akankshya Kar (Santa Monica, CA); Varun Kumar Reddy Bankiti (Bellevue, WA)
Assignee: Motional AD LLC
G06V10/7715G01C21/3822G06V10/761G06V20/58B60W60/001B60W2420/403G06V2201/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,532
App. No.
17/929,405
Granted
Nov 4, 2025
Kind
B2
Abstract

A system may be used to determined object characteristics and/or generate bounding boxes for objects in a vehicle scene by enriching later-in-time feature maps using earlier-in-time feature maps. The system may generate a feature map from a received. Using an earlier-in-time feature map, the system may enrich semantic data of the generated feature map to form an enriched feature map. The system may use the enriched feature map to generate one or more object characteristics of an object in the scene.

Claims (67)

1 . A method, comprising:

receiving a first image at a first time;

generating a first feature map based on the first image;

obtaining a second feature map, the second feature map corresponding to a second image received at a second time, wherein the second time is before the first time;

enriching the first feature map with the second feature map to form a first enriched feature map;

generating a third feature map based on the first enriched feature map;

enriching the third feature map with a fourth feature map to form a second enriched feature map, the fourth feature map based on the second feature map; and

determining a characteristic of an object in the first image based on the second enriched feature map.

2 . The method of claim 1 , wherein generating the first feature map based on the first image comprises generating the first feature map using an image feature extractor that includes a feature pyramid network.

3 . The method of claim 2 , further comprising:

receiving the second image at the second time; and

generating the second feature map based on the second image using the image feature extractor that includes the feature pyramid network.

4 . The method of claim 1 , further comprising:

generating at least one bounding box for the object based on the determined characteristic; and

causing a vehicle to be controlled based on the at least one bounding box.

5 . The method of claim 1 , wherein enriching the first feature map with the second feature map comprises concatenating features of the second feature map with respective features of the first feature map to form the first enriched feature map.

6 . The method of claim 1 , wherein enriching the first feature map with the second feature map comprises:

identifying a particular grid cell in the first feature map;

identifying a set of grid cells in the second feature map associated with the particular grid cell based on a shifting value;

generating a weighting value for each of the set of grid cells relative to the particular grid cell based on a comparison of features of the particular grid cell relative to features of each grid cell of the set of grid cells;

weighting at least one feature of each grid cell of the set of grid cells based on the weighting value to provide at least one weighted feature of the each grid cell of the set of grid cells; and

modifying the at least one feature of the particular grid cell based on the at least one weighted feature of the each grid cell of the set of grid cells.

7 . The method of claim 1 , wherein enriching the first feature map with the second feature map comprises:

identifying a first grid cell in the first feature map;

identifying a set of grid cells in the second feature map associated with the first grid cell based on a shifting value;

generating a set of weighting values for the set of grid cells relative to the first grid cell based on a comparison of features of the first grid cell with features of each grid cell of the set of grid cells, wherein the set of weighting values includes a second weighting value for a second grid cell in the second feature map;

weighting at least one feature of the second grid cell based on the second weighting value to provide at least one weighted feature of the second grid cell; and

modifying at least one feature of the first grid cell based on the at least one weighted feature of the second grid cell.

8 . The method of claim 1 , wherein enriching the third feature map with the fourth feature map comprises concatenating features of the fourth feature map with respective features of the third feature map to form the second enriched feature map.

9 . The method of claim 1 , wherein enriching the third feature map with the fourth feature map comprises:

identifying a third grid cell in the third feature map;

identifying a second set of grid cells in the fourth feature map associated with a fourth grid cell based on a second shifting value;

generating a second set of weighting values for the second set of grid cells relative to the third grid cell based on a comparison of features of the third grid cell with features of each grid cell of the second set of grid cells, wherein the second set of weighting values includes a fourth weighting value for the fourth grid cell in the fourth feature map;

weighting at least one feature of the fourth grid cell based on the fourth weighting value to provide at least one weighted feature of the fourth grid cell; and

modifying at least one feature of the third grid cell based on the at least one weighted feature of the fourth grid cell.

10 . The method of claim 1 , wherein the characteristic of the object in the first image comprises a depth of the object.

11 . The method of claim 1 , wherein the characteristic of the object in the first image comprises a classification of the object.

12 . The method of claim 1 , wherein the characteristic of the object in the first image comprises at least one of a centerness, offset, size, rotation, direction, or velocity of the object.

13 . The method of claim 1 , wherein the characteristic of the object is a first characteristic, the method further comprising:

generating a fifth feature map based on the first enriched feature map;

enriching the fifth feature map with a sixth feature map to form a third enriched feature map, the sixth feature map based on the second feature map; and

determining a second characteristic of the object in the first image based on the third enriched feature map.

14 . The method of claim 13 , further comprising:

generating a seventh feature map based on the first enriched feature map;

enriching the seventh feature map with an eighth feature map to form a fourth enriched feature map, the eighth feature map based on the second feature map; and

determining a third characteristic of the object in the first image based on the fourth enriched feature map.

15 . The method of claim 14 , further comprising:

generating at least one bounding box for the object based on the first characteristic, the second characteristic, and the third characteristic; and

causing a vehicle to be controlled based on the at least one bounding box.

16 . A system, comprising:

a data store storing computer-executable instructions; and

a processor configured to:

receive a first image at a first time;

generate a first feature map based on the first image;

obtain a second feature map, the second feature map corresponding to a second image received at a second time, wherein the second time is before the first time;

enrich the first feature map and the second feature map to form a first enriched feature map; and

generate a third feature map based on the first enriched feature map;

enrich the third feature map with a fourth feature map to form a second enriched feature map, the fourth feature map based on the second feature map; and

determine a characteristic of an object in the first image based on the second enriched feature map.

17 . Non-transitory computer-readable media comprising computer-executable instructions that, when executed by a computing system, causes the computing system to:

receive a first image at a first time;

generate a first feature map based on the first image;

obtain a second feature map, the second feature map corresponding to a second image received at a second time, wherein the second time is before the first time;

enrich the first feature map and the second feature map to form a first enriched feature map;

generate a third feature map based on the first enriched feature map;

enrich the third feature map with a fourth feature map to form a second enriched feature map, the fourth feature map based on the second feature map; and

determine a characteristic of an object in the first image based on the second enriched feature map.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2022
From: SUNDARESHAN, BALAJI; KAR, AKANKSHYA; BANKITI, VARUN KUMAR REDDY
To: MOTIONAL AD LLC
Reel/Frame 061959/0831 →
NUNC PRO TUNC ASSIGNMENT Recorded Sep 6, 2022
From: SUNDARESHAN, BALAJI; KAR, AKANKSHYA; BANKITI, VARUN KUMAR REDDY
To: MOTIONAL AD LLC
Reel/Frame 061001/0635 →
Continuity (1)
Related Publication 20240078790A1 · Mar 7, 2024
References Cited (13)
CN 110916701A · 2020 [cited by examiner]
Fujitake et al., Temporal feature enhancement network with external memory for object detection in surveillance video, 2020 25th International Conference on Pattern Recognition (ICPR), pp. 7684-7691, Jan. 10-15 (Year: 2… [cited by examiner]
Deng et al., Single Shot Video Object Detector, IEEE Transactions on Multimedia, vol. 23, pp. 846-858. (Year: 2021). [cited by examiner]
Sabet et al., Temporal early exits for efficient video object detection arXiv 2106.11208v1 (Year: 2021). [cited by examiner]
Han et al., VISOLO: Grid-Based Space-Time Aggregation for Efficient Online Video Instance Segmentation, IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2886-2895, June (Year: 2022). [cited by examiner]
Perreault et al., RN-VID: A Feature Fusion Architecture for Video Object Detection, arXiv 2003.10898v2 (Year: 2020). [cited by examiner]
SAE On-Road Automated Vehicle Standards Committee, “SAE International's Standard J3016: Taxonomy and Definitions for Terms Related to Driving Automation Systems for On-Road Motor Vehicles”, Jun. 2018, in 35 pages. [cited by applicant]
Tian, Z. et al., “FCOS: Fully Convolutional One-Stage Object Detection”, 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 27-Nov. 2, 2019, in 13 pages. URL: https://doi.org/10.48550/arXiv.1904.0135… [cited by applicant]
Deng, J. et al., “Single Shot Video Object Detector”, IEEE Transactions on Multimedia, 2021, vol. 23, pp. 846-858. [cited by applicant]
Fujitake, M. et al., “Temporal feature enhancement network with external memory for object detection in surveillance video”, 2020 25th International Conference on Pattern Recognition (ICPR), Jan. 10-15, 2021, pp. 7684-7… [cited by applicant]
Han, S. H. et al., “VISOLO: Grid-Based Space-Time Aggregation for Efficient Online Video Instance Segmentation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2022, pp. 2886-2895. [cited by applicant]
International Search Report and Written Opinion received for PCT Patent Application No. PCT/US2023/031755, mailed on Nov. 3, 2023. [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US2023/031755, mailed on Mar. 13, 2025. [cited by applicant]