IP Library › Granted Patent US 12,050,461
Granted Patent B2
US 12,050,461 · App. 16/563,314 · Granted Jul 30, 2024

Video system with frame synthesis

Inventor: Alexander Beer Weiss (San Francisco, CA)
Assignee: DOORDASH, INC.
G05D1/0038G05D1/0022G06F16/71G06N3/02H04N5/265H04N7/0127H04N7/0135H04N23/54H04L67/12H04L69/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,050,461
App. No.
16/563,314
Granted
Jul 30, 2024
Kind
B2
Abstract

A remote vehicle control system includes a vehicle mounted sensor system including a video camera system for producing video data. A data handling system is connected to a network to transmit data to and receive data from a remote teleoperation site. A virtual control system is configured to receive the video, provide a user with a live video stream supported by machine intelligence directed frame synthesis, and transmit control instructions to the remote vehicle over the network. The frame synthesis is supported by a convolutional neural network. The frame synthesis may be used to interpolate frames to increase effective frames per second. The frame synthesis may be used to extrapolate frames to replace missing or damaged video frames.

Claims (62)

1. A remote vehicle control system comprising:

a virtual control system configured to:

process video data received from a vehicle mounted sensor system of a vehicle, the vehicle mounted sensor system including a video camera system, wherein the video data includes a plurality of sequential image frames of a video stream, each image frame corresponding to a time position of the video stream, the video stream representing a forward vehicle view, frames of a second vehicle view being prioritized lower than frames of the forward vehicle view;

store the plurality of image frames within a circular buffer;

generate an optical flow map based on at least two consecutive image frames in the circular buffer, the optical flow map predicting pixel positions of a virtual image frame corresponding to a desired time position of the video data;

synthesize the virtual image frame for insertion into the desired time position of the video stream, wherein synthesizing the virtual image frame is supported by a convolutional neural network, the neural network making optical frame predictions at a frequency greater than the frame rate of the video;

create a modified video stream by inserting the virtual image frame into the desired time position of the video stream, the modified video stream representing surroundings of the vehicle;

determining that a collision is occurring based, at least in part, on the predicted pixel positions of the virtual image frame;

generating a warning that the collision is occurring;

provide the modified video stream via a display using distance map data received from the vehicle mounted sensor system;

after the modified video stream has been provided via the display, process input received via an input device, the input indicating control instructions, the control instructions controlling operation of the vehicle; and

transmit the control instructions via a network to the vehicle;

wherein kernels are predicted as a pair of N×1 and 1×N kernels;

wherein convolutions of an encoder-decoder include dilated convolutions.

2. The remote vehicle control system of claim 1 , wherein the network is a cell phone network.

3. The remote vehicle control system of claim 1 , wherein synthesizing the virtual image frame is supported by a convolutional neural network.

4. The remote vehicle control system of claim 1 , wherein synthesizing the virtual image frame is used to interpolate frames to increase effective frames per second.

5. The remote vehicle control system of claim 1 , wherein synthesizing the virtual image frame is used to extrapolate frames to replace missing or damaged video frames.

6. The remote vehicle control system of claim 1 , wherein the video data is received over a cellular network using feed forward correction.

7. The remote vehicle control system of claim 1 , wherein the virtual control system further comprises a virtual reality headset.

8. The remote vehicle control system of claim 1 , the video data including frames ordered at least in part by priority.

9. The remote vehicle control system of claim 1 , the video data including interleaved video frames.

10. The remote vehicle control system of claim 1 , wherein the video data has been compressed using a maximum separable distance erasure code.

11. The remote vehicle control system of claim 1 , wherein the video data has been compressed using Cauchy Reed-Solomon code.

12. The remote vehicle control system of claim 1 , wherein the video data includes packets transmitted by a User Datagram Protocol (UDP).

13. A method comprising:

receiving, at a control system, video data from a vehicle mounted sensor system of a vehicle, the vehicle mounted sensor system including a video camera system, wherein the video data includes a plurality of sequential image frames of a video stream, each image frame corresponding to a time position of the video stream, the video stream representing a forward vehicle view, frames of a second vehicle view being prioritized lower than frames of the forward vehicle view;

storing, by the control system, the plurality of image frames within a circular buffer;

generating, at the control system, an optical flow map based on at least two consecutive image frames in the circular buffer, the optical flow map predicting pixel positions of a virtual image frame corresponding to a desired time position of the video data;

synthesizing, at the control system, the virtual image frame for insertion into the desired time position of the video stream, wherein synthesizing the virtual image frame is supported by a convolutional neural network, the neural network making optical frame predictions at a frequency greater than the frame rate of the video;

creating, at the control system, a modified video stream by inserting the virtual image frame into the desired time position of the video stream, the modified video stream representing surroundings of the vehicle;

determining that a collision is occurring based, at least in part, on the predicted pixel positions of the virtual image frame;

generating a warning that the collision is occurring;

providing, at the control system, the modified video stream via a display using distance map data received from the vehicle mounted sensor system;

after the modified video stream has been provided via the display, processing, at the control system, input received via an input device, the input indicating control instructions, the control instructions controlling operation of the vehicle; and

transmitting, by the control system, the control instructions via a network to the vehicle;

wherein kernels are predicted as a pair of N×1 and 1×N kernels;

wherein convolutions of an encoder-decoder include dilated convolutions.

14. The method of claim 13 , wherein the virtual image frame is an extrapolated image frame and the desired time position is subsequent to the time positions corresponding to the at least two consecutive image frames.

15. The method of claim 14 , further comprising:

determining that an expected image frame corresponding to the desired time position is unavailable; and

inserting the extrapolated image into the video stream subsequent to the at least two consecutive image frames.

16. The method of claim 13 , wherein the virtual image frame is an interpolated image frame and the desired time position is between the time positions corresponding to the at least two consecutive image frames.

17. The method of claim 16 , wherein insertion of the interpolated image increases the frame rate of the video stream.

18. The method of claim 13 , wherein the optical flow map is generated by a convolutional neural network.

19. One or more non-transitory computer readable media having instructions stored thereon for performing a method, the method comprising:

processing, at a control system, video data received from a vehicle mounted sensor system of a vehicle, the vehicle mounted sensor system including a video camera system, wherein the video data includes a plurality of sequential image frames of a video stream, each image frame corresponding to a time position of the video stream, the video stream representing a forward vehicle view, frames of a second vehicle view being prioritized lower than frames of the forward vehicle view;

storing, by the control system, the plurality of image frames within a circular buffer;

generating, at the control system, an optical flow map based on at least two consecutive image frames in the circular buffer, the optical flow map predicting pixel positions of a virtual image frame corresponding to a desired time position of the video data;

synthesizing, at the control system, the virtual image frame for insertion into the desired time position of the video stream, wherein synthesizing the virtual image frame is supported by a convolutional neural network, the neural network making optical frame predictions at a frequency greater than the frame rate of the video;

creating, at the control system, a modified video stream by inserting the virtual image frame into the desired time position of the video stream, the modified video stream representing surroundings of the vehicle;

determining that a collision is occurring based, at least in part, on the predicted pixel positions of the virtual image frame;

generating a warning that the collision is occurring;

providing, at the control system, the modified video stream via a display using distance map data received from the vehicle mounted sensor system;

after the modified video stream has been provided via the display, processing, at the control system, input received via an input device, the input indicating control instructions, the control instructions controlling operation of the vehicle; and

transmitting, by the control system, the control instructions via a network to the vehicle;

wherein kernels are predicted as a pair of N×1 and 1×N kernels;

wherein convolutions of an encoder-decoder include dilated convolutions.

20. The remote vehicle control system of claim 1 , wherein a loss function includes L1 regularization based on summation of absolute weight values.

21. The remote vehicle control system of claim 1 , wherein strided convolutions are used in initial layers of a neural network.

22. The remote vehicle control system of claim 1 , wherein a plurality of convolutions are replaced by bottleneck modules, the bottleneck modules reducing a number of channels via 1×1 convolutions.

23. The remote vehicle control system of claim 1 , wherein an encoder side of a convolutional neural network includes a series of convolutions, followed by application of activation functions, followed by pooling.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2024
From: WEISS, ALEXANDER BEER
To: DOORDASH, INC.
Reel/Frame 067806/0315 →
Continuity (2)
Provisional Application 62728321 · Sep 7, 2018
Related Publication 20200081431A1 · Mar 12, 2020