IP Library › Granted Patent US 12,657,763
Granted Patent B2
US 12,657,763 · App. 17/978,526 · Granted Jun 16, 2026

Aggregating features from multiple images to generate historical data for a camera

Inventors: Paul Lee (Redmond, WA); Michael Bleyer (Seattle, WA); Christian Markus Maekelae (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06T7/74G06T7/251G06T19/006
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,763
App. No.
17/978,526
Granted
Jun 16, 2026
Kind
B2
Abstract

Techniques for generating an aggregated set of features from multiple images generated by a camera are disclosed. A first image generated by the camera is accessed, where the first image was generated at a first time. A first set of features are identified from within the first image. A second image generated by the camera is accessed, where the second image is generated at a subsequent, second time. A second set of features are identified from within the second image. Movement data is obtained. This movement data details a movement of the camera between the first and second times. The movement data is used to reproject a pose embodied in the first image to correspond to a pose embodied in the second image. The embodiments aggregate the two sets of features to generate the aggregated set of features. The aggregated set of features for the camera are then cached.

Claims (54)

1 . A method for generating an aggregated set of features from multiple images generated by a camera, where the aggregated set of features includes multiple permutations of a feature, with each permutation of the feature being tagged with corresponding timing data, said method comprising:

accessing a first image generated by the camera, the first image being generated at a first time;

identifying a first set of features from within the first image by performing feature extraction on the first image;

accessing a second image generated by the camera, the second image being generated at a second time, which is subsequent to the first time;

identifying a second set of features from within the second image by performing feature extraction on the second image;

obtaining movement data detailing a movement of the camera between the first time and the second time;

subsequent to identifying the first set of features and the second set of features, using the movement data to reproject a pose embodied in the first image to correspond to a pose embodied in the second image;

aggregating the first set of features, which underwent said reprojection, with the second set of features to generate the aggregated set of features, wherein:

each feature in the aggregated set of features is tagged with corresponding timing data reflecting when a corresponding image for said each feature was generated, and

the aggregated set of features comprises multiple different permutations for each of at least some of the features, such that each permutation is also associated with corresponding timing data;

caching the aggregated set of features for the camera, resulting in a history of feature extraction results being accessible for an image alignment operation; and

performing the image alignment operation by accessing the history of feature extraction results, wherein the image alignment operation includes performing an alignment using different permutations of a common feature that is represented in different images, where the different permutations have different timing data.

2 . The method of claim 1 , wherein the movement data includes inertial measurement unit (IMU) data.

3 . The method of claim 1 , wherein reprojecting the pose embodied in the first image to correspond to the pose embodied in the second image involves use of a motion model.

4 . The method of claim 1 , wherein the first image and the second image capture a low-contrast scene, and wherein a number of features included in the first set of features is less than a threshold number.

5 . The method of claim 4 , wherein, after aggregating the first set of features with the second set of features, a number of features included in the aggregated set of features at least meets the threshold number.

6 . The method of claim 1 , wherein a number of features included in the first set of features is less than a threshold number, wherein a number of features included in the second set of features is also less than the threshold number, and wherein a number of features included in the aggregated set of features at least meets the threshold number.

7 . The method of claim 1 , wherein a frame rate of the camera is between 30 frames per second (FPS) and 120 FPS.

8 . The method of claim 1 , wherein the first and second images are of a type comprising: a thermal image type, a low light image type, or a visible light image type.

9 . A computer system that generates an aggregated set of features from multiple images generated by a camera, where the aggregated set of features includes multiple permutations of a feature, with each permutation of the feature being tagged with corresponding timing data, said computer system comprising:

at least one processor; and

at least one hardware storage device that stores instructions that are executable by the at least one processor to cause the computer system to:

access a first image generated by the camera, the first image being generated at a first time;

identify a first set of features from within the first image by performing feature extraction on the first image;

access a second image generated by the camera, the second image being generated at a second time, which is subsequent to the first time;

identify a second set of features from within the second image by performing feature extraction on the second image;

obtain movement data detailing a movement of the camera between the first time and the second time;

subsequent to identifying the first set of features and the second set of features, use the movement data to reproject a pose embodied in the first image to correspond to a pose embodied in the second image;

aggregate the first set of features, which underwent said reprojection, with the second set of features to generate the aggregated set of features, wherein:

each feature in the aggregated set of features is tagged with corresponding timing data reflecting when a corresponding image for said each feature was generated, and

the aggregated set of features comprises multiple different permutations for each of at least some of the features, such that each permutation is also associated with corresponding timing data;

cache the aggregated set of features for the camera, resulting in a history of feature extraction results being accessible for an image alignment operation; and

perform the image alignment operation by accessing the history of feature extraction results, wherein the image alignment operation includes performing an alignment using different permutations of a common feature that is represented in different images, where the different permutations have different timing data.

10 . The computer system of claim 9 , wherein additional sets of features identified from additional images generated by the camera are aggregated with the aggregated set of features.

11 . The computer system of claim 9 , wherein the aggregated set of features are compiled into a single unit vector space.

12 . The computer system of claim 9 , wherein the camera is one of a low light camera, a thermal camera, or a visible light camera.

13 . The computer system of claim 9 , wherein a frame rate of the camera is between 30 frames per second (FPS) and 60 FPS.

14 . The computer system of claim 9 , wherein the camera is integrated with the computer system.

15 . The computer system of claim 9 , wherein the camera is peripheral relative to the computer system.

16 . The computer system of claim 9 , wherein the computer system is a wearable mixed-reality (MR) device.

17 . A method for generating an aggregated set of features from multiple images generated by a camera, where the aggregated set of features includes multiple permutations of a feature, with each permutation of the feature being tagged with corresponding timing data, said method comprising:

accessing a first image generated by the camera, the first image being generated at a first time;

identifying a first set of features from within the first image by performing feature extraction on the first image;

accessing a second image generated by the camera, the second image being generated at a second time, which is subsequent to the first time;

identifying a second set of features from within the second image by performing feature extraction on the second image;

obtaining movement data detailing a movement of the camera between the first time and the second time;

subsequent to identifying the first set of features and the second set of features, using the movement data to reproject a pose embodied in the first image to correspond to a pose embodied in the second image;

aggregating the first set of features, which underwent said reprojection, with the second set of features to generate the aggregated set of features, wherein:

each feature in the aggregated set of features is tagged with corresponding timing data reflecting when a corresponding image for said each feature was generated, and

the aggregated set of features comprises multiple different permutations for each of at least some of the features, such that each permutation is also associated with corresponding timing data;

caching the aggregated set of features for the camera, resulting in a history of feature extraction results being accessible for an image alignment operation; and

accessing the history of feature extraction results during the image alignment operation in which an image generated by the camera is aligned with a different image generated by a different camera, wherein the image alignment operation includes performing said aligning using different permutations of a common feature that is represented in both said image and said different image, where the different permutations have different timing data.

18 . The method of claim 17 , wherein a number of features included in the first set of features is less than a threshold number, wherein a number of features included in the second set of features is also less than the threshold number, and wherein a number of features included in the aggregated set of features at least meets the threshold number.

19 . The method of claim 17 , wherein the movement data includes inertial measurement unit (IMU) data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2022
From: LEE, PAUL; BLEYER, MICHAEL; MAEKELAE, CHRISTIAN MARKUS
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 061614/0742 →
Continuity (1)
Related Publication 20240144526A1 · May 2, 2024
References Cited (23)
US 11049277B1 · Price · 2021 [cited by examiner]
US 20190043203A1 · Fleishman · 2019 [cited by examiner]
US 20200198149A1 · Jiang · 2020 [cited by examiner]
US 20220164988A1 · Dotsenko · 2022 [cited by applicant]
US 20230031023A1 · Wang · 2023 [cited by applicant]
US 20230059657A1 · Hu · 2023 [cited by applicant]
US 20230316607A1 · He · 2023 [cited by applicant]
US 20240144496A1 · Lee · 2024 [cited by applicant]
EP 2491532B1 · 2022 [cited by applicant]
Sun et al., “Deep Video Matting via Spatio-Temporal Alignment and Aggregation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6975-6984 (Year: 2021). [cited by examiner]
Badino et al., “Visual Odometry by Multi-frame Feature Integration,” 2013 IEEE International Conference on Computer Vision Workshops, pp. 222-229 (Year: 2013). [cited by examiner]
Yang, B.—“Robust Attentional Aggregation of Deep Feature Sets for Multi-view 3D Reconstruction”—International Journal of Computer Vision 2020—pp. 53-73 (Year: 2020). [cited by examiner]
Badino, et al., “Visual odometry by multi-frame feature integration”, Proceedings of the IEEE International Conference on Computer Vision Workshops, 2013, pp. 222-229. [cited by applicant]
Chen, et al., “Coherent online video style transfer”, In Proceedings of the IEEE International Conference on Computer Vision, 2017, pp. 1105-1114. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2023/033646, (MS #412302-PCT01) Feb. 20, 2024, 11 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US23/033652, (MS#412413-PCT01), Feb. 8, 2024, 15 pages. [cited by applicant]
Jie Zhou et al: “Video Stabilization and Completion Using Two Cameras”, IEEE Transactions on Circuits and Systems for Video Technology, IEEE, USA, vol. 21, issue No. 12, Dec. 1, 2011, pp. 1879-1889. [cited by applicant]
Notice of Allowance mailed on Mar. 10, 2025, in U.S. Appl. No. 17/978,463 (MS# 412302-US01), 12 pages. [cited by applicant]
Notice of Allowance mailed on Jun. 3, 2025, in U.S. Appl. No. 17/978,463, (MS#412302-US01) 13 pages. [cited by applicant]
Supplemental Notice of Allowability mailed on Jun. 26, 2025, in U.S. Appl. No. 17/978,463 (MS# 412302-US01), 09 pages. [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US2023/033646, (MS# 412302-PCT01) May 15, 2025, 07 pages. [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US2023/033652, (MS# 412413-PCT01) May 15, 2025, 9 pages. [cited by applicant]
U.S. Appl. No. 17/978,463, filed Nov. 1, 2022. [cited by applicant]