IP Library Granted Patent US 12,198,358
Granted Patent B2
US 12,198,358 · App. 17/962,624 · Granted Jan 14, 2025

Deep structured scene flow for autonomous devices

Inventors: Raquel Urtasun (Toronto, CA); Wei-Chiu Ma (Toronto, CA); Shenlong Wang (Toronto, CA); Yuwen Xiong (Toronto, CA); Rui Hu (Los Alamitos, CA)
Assignee: AURORA OPERATIONS, INC.
G06T7/285G06T7/246G06T7/593G06T7/97G06V10/40G06V10/82G06V20/56G06T2207/10012G06T2207/10028G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,358
App. No.
17/962,624
Granted
Jan 14, 2025
Kind
B2
Abstract

Systems, methods, tangible non-transitory computer-readable media, and devices associated with motion flow estimation are provided. For example, scene data including representations of an environment over a first set of time intervals can be accessed. Extracted visual cues can be generated based on the representations and machine-learned feature extraction models. At least one of the machine-learned feature extraction models can be configured to generate a portion of the extracted visual cues based on a first set of the representations of the environment from a first perspective and a second set of the representations of the environment from a second perspective. The extracted visual cues can be encoded using energy functions. Three-dimensional motion estimates of object instances at time intervals subsequent to the first set of time intervals can be determined based on the energy functions and machine-learned inference models.

Claims (46)

1. A computer-implemented method comprising:

accessing a first pair of stereo images and a second pair of stereo images of an environment of an autonomous vehicle from a pair of stereo cameras, wherein the first pair of stereo images comprises a first image of the environment from a first perspective at a first time and a second image of the environment from a second perspective at the first time, wherein the second pair of stereo images comprise a first image of the environment from the first perspective at a second time and a second image of the environment from the second perspective at the second time;

generating, in less than 10 seconds, a three-dimensional motion estimate for an object in the environment using a machine-learned motion flow model by:

determining, using the machine-learned motion flow model, a plurality of extracted features from the first pair of stereo images and the second pair of stereo images, wherein at least one feature of the plurality of features describes an object instance associated with an object in the environment; and

processing, using the machine-learned motion flow model, the plurality of extracted features, wherein the machine-learned motion flow model is configured to solve an energy optimization function for characterizing three-dimensional rigid motion of the object instance; and

controlling a motion of the autonomous vehicle based on the three-dimensional motion estimate.

2. The computer-implemented method of claim 1 , wherein the three-dimensional motion estimate is generated in less than 1 second.

3. The computer-implemented method of claim 2 , wherein processing, using the machine-learned motion flow model, the plurality of extracted features comprises:

executing a solver to optimize, over a plurality of steps, one or more energy terms associated with motion of the object instance;

wherein the solver executes in less than 1 second.

4. The computer-implemented method of claim 1 , wherein the machine-learned motion flow model comprises a machine-learned segmentation model.

5. The computer-implemented method of claim 4 , wherein the machine-learned segmentation model is configured to associate the object instance with a portion of at least one of the first pair of stereo images and the second pair of stereo images.

6. The computer-implemented method of claim 1 , wherein the machine-learned motion flow model is trained end-to-end.

7. The computer-implemented method of claim 1 , wherein the plurality of extracted features from the first pair of stereo images and the second pair of stereo images comprise one or more visual cues, the one or more visual cues comprising at least one of an instance segmentation cue, an optical flow cue, or a stereo cue.

8. The computer-implemented method of claim 1 , wherein the machine-learned motion flow model is configured to solve the energy optimization function using a Gaussian-Newton (GN) algorithm implemented as layers in a neural network.

9. The computer-implemented method of claim 1 , the method further comprising:

removing uncertain pixels from the plurality of extracted features before processing the plurality of extracted features with the machine-learned motion flow model.

10. A computing system comprising:

one or more processors; and

one or more tangible non-transitory computer readable media storing computer-readable instructions that are executable by the one or more processors to cause the one or more processors to perform operations, the operations comprising:

accessing a first pair of stereo images and a second pair of stereo images of an environment of an autonomous vehicle from a pair of stereo cameras, wherein the first pair of stereo images comprises a first image of the environment from a first perspective at a first time and a second image of the environment from a second perspective at the first time, wherein the second pair of stereo images comprise a first image of the environment from the first perspective at a second time and a second image of the environment from the second perspective at the second time;

generating, in less than 10 seconds, a three-dimensional motion estimate for an object in the environment using a machine-learned motion flow model by:

determining, using the machine-learned motion flow model, a plurality of extracted features from the first pair of stereo images and the second pair of stereo images, wherein at least one feature of the plurality of features describes an object instance associated with an object in the environment; and

processing, using the machine-learned motion flow model, the plurality of extracted features, wherein the machine-learned motion flow model is configured to solve an energy optimization function for characterizing three-dimensional rigid motion of the object instance; and

controlling a motion of the autonomous vehicle based on the three-dimensional motion estimate.

11. The computing system of claim 10 , wherein the three-dimensional motion estimate is generated in less than 1 second.

12. The computing system of claim 10 , wherein the plurality of extracted features from the first pair of stereo images and the second pair of stereo images comprise one or more visual cues, the one or more visual cues comprising at least one of an instance segmentation cue, an optical flow cue, or a stereo cue.

13. The computing system of claim 10 , wherein processing, using the machine-learned motion flow model, the plurality of extracted features comprises:

executing a solver to optimize, over a plurality of steps, one or more energy terms associated with motion of the object instance;

wherein the solver executes in less than 1 second.

14. The computing system of claim 13 , wherein the plurality of steps are implemented as a plurality of layers of a neural network.

15. The computing system of claim 14 , comprising:

a graphical processing unit (GPU);

wherein the operations comprise:

executing the layers of the neural network on the GPU.

16. The computing system of claim 10 , wherein the machine-learned motion flow model is trained end-to-end.

17. The computing system of claim 10 , wherein the pair of stereo cameras is configured to obtain the first pair of stereo images and the second pair of stereo images of an environment associated with an augmented reality system.

18. One or more tangible non-transitory computer readable media storing computer-readable instructions that are executable by one or more processors to cause the one or more processors to perform operations, the operations comprising:

accessing a first pair of stereo images and a second pair of stereo images of an environment of an autonomous vehicle from a pair of stereo cameras, wherein the first pair of stereo images comprises a first image of the environment from a first perspective at a first time and a second image of the environment from a second perspective at the first time, wherein the second pair of stereo images comprise a first image of the environment from the first perspective at a second time and a second image of the environment from the second perspective at the second time;

generating, in less than 10 seconds, a three-dimensional motion estimate for an object in the environment using a machine-learned motion flow model by:

determining, using the machine-learned motion flow model, a plurality of extracted features from the first pair of stereo images and the second pair of stereo images, wherein at least one feature of the plurality of features describes an object instance associated with an object in the environment; and

processing, using the machine-learned motion flow model, the plurality of extracted features, wherein the machine-learned motion flow model is configured to solve an energy optimization function for characterizing three-dimensional rigid motion of the object instance; and

controlling a motion of the autonomous vehicle based on the three-dimensional motion estimate.

19. The computer-implemented method of claim 3 , wherein the plurality of steps are implemented as a plurality of layers of a neural network.

20. The computer-implemented method of claim 19 , comprising:

executing the layers of the neural network on a graphical processing unit (GPU).

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2024
From: UATC, LLC
To: AURORA OPERATIONS, INC.
Reel/Frame 067733/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2022
From: WANG, SHENLONG; HU, RUI; MA, WEI-CHIU; XIONG, YUWEN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 061424/0051 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2022
From: UBER TECHNOLOGIES, INC.
To: UATC, LLC
Reel/Frame 061426/0489 →
EMPLOYMENT AGREEMENT Recorded Oct 14, 2022
From: SOTIL, RAQUEL URTASUN
To: UBER TECHNOLOGIES, INC.
Reel/Frame 061681/0609 →
Continuity (4)
Continuation 16531720 · Aug 5, 2019
Provisional Application 62851753 · May 23, 2019
Provisional Application 62768774 · Nov 16, 2018
Related Publication 20230038786A1 · Feb 9, 2023
References Cited (60)
US 9679227B2 · Taylor et al. · 2017 [cited by applicant]
US 10037613B1 · Becker · 2018 [cited by examiner]
US 20170295360A1 · Fu · 2017 [cited by examiner]
US 20180157918A1 · Levkova et al. · 2018 [cited by applicant]
US 20180293454A1 · Xu · 2018 [cited by examiner]
US 20180293737A1 · Sun et al. · 2018 [cited by applicant]
US 20190004543A1 · Kennedy · 2019 [cited by examiner]
US 20190050000A1 · Kennedy · 2019 [cited by examiner]
US 20190066326A1 · Tran et al. · 2019 [cited by applicant]
US 20190096032A1 · Li · 2019 [cited by examiner]
US 20190122381A1 · Long · 2019 [cited by examiner]
US 20190145765A1 · Luo et al. · 2019 [cited by applicant]
US 20190162856A1 · Atalla · 2019 [cited by examiner]
US 20210004682A1 · Gong · 2021 [cited by examiner]
M. Menze, “Joint 3D estimation of Vehicles and scene flow” (Year: 2015). [cited by examiner]
Bai et al, Exploiting Semantic Information and Deep Matching for Optical Flow, arXiv:1604v2, Aug. 23, 2016, 16 pages. [cited by applicant]
Basha et al, “Multi-view Scene Flow Estimation: A View Centered Variational Approach”, International Journal of Computer Vision, 2013, 16 pages. [cited by applicant]
Behl et al, “Bounding Boxes, Segmentations and Object Coordinates: How Important is Recognition for 3D Scene Flow Estimation in Autonomous Driving Scenarios?” International Conference on Computer Vision, Oct. 22-29, 201… [cited by applicant]
Black et al, “The Robust Estimation of Multiple Motions: Parametric and Piecewise-Smooth Flow Fields”, Computer Vision and Image Understanding, vol. 63, No. 1, Jan. 1996, pp. 75-104. [cited by applicant]
Boyd et al, “Convex Optimization”, Cambridge University Press, Cambridge, Massachusetts, 2004, 713 pages. [cited by applicant]
Brox et al, “High Accuracy Optical Flow Estimation Based on a Theory for Warping” Conference on Computer Vision, May 11-14, 2004, Prague, Czech Republic, 13 pages. [cited by applicant]
Chang et al, “Pyramid Stereo Matching Network”, arXiv:1803v1, Mar. 23, 2018, 9 pages. [cited by applicant]
Chen et al, “A Deep Visual Correspondence Embedding Model for Stereo Matching Costs”, International Conference on Computer Vision, Dec. 7-13, 2016, Santiago, Chile, 9 pages. [cited by applicant]
Cordts et al, “The Cityscapes Dataset for Semantic Urban Scene Understanding”, arXiv:1604v2, Apr. 7, 2016, 29 pages. [cited by applicant]
Fischer et al, “Flownet: Learning Optical Flow with Convolutional Networks”, arXiv:1504v2, May 4, 2015, 13 pages. [cited by applicant]
Geiger et al, “Are We Ready for Autonomous Driving? The Kitti Vision Benchmark Suite”, Conference on Computer Vision and Pattern Recognition, Jun. 18-20, 2012, Providence, Rhode Island, 8 pages. [cited by applicant]
He et al, “Mask R-CNN”, arXiv:1703v3, Jan. 24, 2018, 12 pages. [cited by applicant]
He et al, “Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition”, arXiv:1406v4, Apr. 23, 2015, 14 pages. [cited by applicant]
Hirschmuller, “Stereo Processing by Semiglobal Matching and Mutual Information”, Transactions on Pattern Analysis and Machine Intelligence, vol. 30, Issue 2, Feb. 2008, 14 pages. [cited by applicant]
Hoff et al, “Surfaces from Stereo: Integrating Feature Matching, Disparity Estimation and Contour Detection”, Transactions on Pattern Analysis and Machine Intelligence, vol. 11, No. 2, Feb. 1989, pp. 121-136. [cited by applicant]
Horn et al, “Determining Optical Flow”, Artificial Intelligence, vol. 17, 1981, pp. 185-203. [cited by applicant]
Huguet et al, “A Variational Method for Scene Flow Estimation from Stereo Sequences”, International Conference on Computer Vision, Oct. 14-20, 20117, Rio de Janeiro, 8 pages. [cited by applicant]
Hui et al, “Liteflownet: A Lightweight Convolutional Neural Network for Optical Flow Estimation”, Jun. 19-21, 2018, Salt Lake City, Utah, pp. 8981-8989. [cited by applicant]
Ilg et al, “Flownet 2.0: Evolution of Optical Flow Estimation with Deep Networks”, arXiv:1612v1, Dec. 6, 2016, 16 pages. [cited by applicant]
Kanade et al, “A Stereo Matching Algorithm with an Adaptive Window: Theory and Experiment”, Technical Report, School of Computer Science, Carnegie Mellon University, Pittsburgh, Pennsylvania, 1990, 30 pages. [cited by applicant]
Kendall et al, “End-to-End Learning of Geometry and Context for Deep Stereo Regression”, arXiv:1703v1, Mar. 13, 2017, 10 pages. [cited by applicant]
Luo et al, “Efficient Deep Learning for Stereo Matching”, Conference on Computer Vision and Pattern Recognition, Jun. 26-Jul. 1, 2016, Las Vegas, Nevada, 9 pages. [cited by applicant]
Lv et al, “A Continuous Optimization Approach for Efficient and Accurate Scene Flow”, arXiv:1607v1, Jul. 27, 2016, 16 pages. [cited by applicant]
Mayer et al, “A Large Dataset to Train Convolutional Networks for Disparity, Optical Flow, and Scene Flow Estimation”, arXiv:1512v1, Dec. 7, 2015, 14 pages. [cited by applicant]
Menze et al, “Object Scene Flow for Autonomous Vehicles”, Conference on Computer Vision and Pattern Recognition, Jun. 8-10, 2015, Boston, Massachusetts, 10 pages. [cited by applicant]
Neoral et al, “Object Scene Flow with Temporal Consistency”, Computer Vision Winter Workshop, Feb. 6-8, 2017, Retz, Austria, 9 pages. [cited by applicant]
Newell et al, “Stacked Hourglass Networks for Human Pose Estimation”, arXiv:1603v2, Jul. 26, 2016, 17 pages. [cited by applicant]
Papenberg et al, “Highly Accurate Optic Flow Computation with Theoretically Justified Warping”, International Journal of Computer Vision, 2006, 18 pages. [cited by applicant]
Pons et al, “Multi-View Stereo Reconstruction and Scene Flow Estimation with a Global Image-Based Matching Score”, International Journal of Computer Vision, 2007, 6 pages. [cited by applicant]
Ranjan et al, “Optical Flow Estimation Using a Spatial Pyramid Network”, arXiv:1611v2, Nov. 21, 2016, 10 pages. [cited by applicant]
Ren et al, “Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks”, arXiv:1506v3, Jan. 6, 2016, 14 pages. [cited by applicant]
Revaud et al, “EpicFlow: Edge-Preserving Interpolation of Correspondences for Optical Flow”, arXiv:1501v2, May 19, 2015, 11 pages. [cited by applicant]
Sun et al, “Models Matter, so Does Training: An Empirical Study of CNNs for Optical Flow Estimation”, arXiv:1809v1, Sep. 14, 2018, 15 pages. [cited by applicant]
Sun et al, “PWC-Net: CNNs for Optical Flow Using Pyramid, Warping, and Cost Volume”, arXiv:1709v3, Jun. 25, 2018, 18 pages. [cited by applicant]
Sun et al, “Secrets of Optical Flow Estimation and Their Principles”, Conference on Computer Vision and Pattern Recognition, Jun. 13-18, 2010, San Francisco, California, 8 pages. [cited by applicant]
Valgaerts et al, “Joint Estimation of Motion, Structure and Geometry from Stereo Sequences”, Conference on Computer Vision, Sep. 5-10, 2010, Crete, Greece, 14 pages. [cited by applicant]
Vedula et al, “Three-Dimensional Scene Flow”, International Conference on Computer Vision, Sep. 20-27, 1999, Corfu, Greece, 8 pages. [cited by applicant]
Vogel et al, “3D Scene Flow Estimation with a Piecewise Rigid Scene Model”, International Journal of Computer Science, vol. 111, No. 3, 2013, 27 pages. [cited by applicant]
Vogel et al, “Piecewise Rigid Scene Flow”, International Conference on Computer Vision, Dec. 3-6, 2013, Sydney, Australia, 8 pages. [cited by applicant]
Wang et al, “Autoscaler: Scale-Attention Networks for Visual Correspondence”, arXiv:1611v1, Nov. 17, 2016, 10 pages. [cited by applicant]
Yamaguchi et al, “Efficient Joint Segmentation, Occlusion Labeling, Stereo and Flow Estimation”, Conference on Computer Vision, Sep. 6-12, 2014, Zurich, Switzerland, 16 pages. [cited by applicant]
Zagoruyko et al, “Learning to Compare Image Patches via Convolutional Neural Networks”, arXiv:1504v1, Apr. 14, 2015, 9 pages. [cited by applicant]
Zbontar et al, “Computing the Stereo Matching Cost with a Convolutional Neural Network”, arXiv:1409v2, Oct. 20, 2015, 8 pages. [cited by applicant]
Zbontar et al, “Stereo Matching by Training a Convolutional Neural Network to Compare Image Patches”, arXiv:1510v2, May 18, 2016, 32 pages. [cited by applicant]
Zhao et al, “Pyramid Scene Parsing Network”, arXiv:1612v2, Apr. 27, 2017, 11 pages. [cited by applicant]