IP Library Granted Patent US 10,019,652
Granted Patent B2
US 10,019,652 · App. 15/051,005 · Granted Jul 10, 2018

Generating a virtual world to assess real-world video analysis performance

Inventors: Qiao Wang (Phoenix, AZ); Adrien Gaidon (Grenoble, FR); Eleonora Vig (Grenoble, FR)
Assignee: Xerox Corporation
G06K9/6262G06K9/00711G06T7/20G06T19/00G06T2207/10016G06T2207/30236
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,019,652
App. No.
15/051,005
Filed
Feb 23, 2016
Granted
Jul 10, 2018
Kind
B2
Art Unit
2665
USPC
382/159
Abstract

A system and method are suited for assessing video performance analysis. A computer graphics engine clones real-world data in a virtual world by decomposing the real-world data into visual components and objects in one or more object categories and populates the virtual world with virtual visual components and virtual objects. A scripting component controls the virtual visual components and the virtual objects in the virtual world based on the set of real-world data. A synthetic clone of the video sequence is generated based on the script controlling the virtual visual components and the virtual objects. The real-world data is compared with the synthetic clone of the video sequence and a transferability of conclusions from the virtual world to the real-world is assessed based on this comparison.

Claims (54)

1. A method of assessing video performance analysis comprising:

with a camera, acquiring a first set of real-world data including a video sequence, the video sequence including visual components and objects;

with at least one sensor, automatically generating a second set of real-world data including a set of physical measurements for the objects and visual components of the video sequence;

with a computer graphics engine, cloning the first and second sets of real-world data in a virtual world, comprising:

decomposing the visual components and objects in the video sequence of the first set of real-world data into at least one object category,

populating the virtual world with virtual visual components and virtual objects,

generating a script for controlling the virtual visual components and the virtual objects in the virtual world based on the acquired visual components and objects in the first set of real-world data and the automatically generated set of physical measurements for the objects and visual components in the second set of real-world data, and

generating a synthetic clone of the video sequence based on the script for controlling the virtual visual components and the virtual objects;

generating a set of ground truth annotations for the virtual objects in the synthetic clone;

comparing the real-world data with the synthetic clone of the video sequence; and

assessing a transferability of conclusions from the virtual world to the real-world based on the comparison of the real-world data with the synthetic clone of the video sequence.

2. The method of claim 1 , wherein the acquiring the first and second sets of real-world data further comprises annotating the objects in an object category with the set of physical measurements from the at least one sensor, the at least one sensor including at least one of a global positioning system, an inertia measurement unit, an ultraviolet camera, and a 3D laser scanner, and wherein the camera includes at least one of a monochrome camera, a color camera, and an infrared camera.

3. The method of claim 2 , wherein the annotating with the set of physical measurements includes annotating the objects with at least one of a position, size, inertia, and orientation of each of the objects in the at least one object category.

4. The method of claim 1 , further comprising changing a condition of one of the virtual visual components or the virtual objects.

5. The method of claim 4 , further comprising generating a modified synthetic video based on the changed condition and generating a set of ground truth annotations for the virtual objects in the modified synthetic video.

6. The method of claim 4 , wherein the changing a condition includes changing a position, orientation, trajectory, size, color, or shape of at least one of the virtual objects.

7. The method of claim 4 , wherein the changing a condition of one of the virtual visual components includes changing a lighting or a weather condition.

8. The method of claim 4 , wherein the changing a condition includes manually adding, modifying or removing at least one of the virtual objects.

9. The method of claim 1 , further comprising performing a specific task with an algorithm on the first and second sets of real-world data and performing the specific task with the algorithm on the synthetic clone of the video sequence.

10. The method of claim 9 , further comprising evaluating the performance of the algorithm against the set of ground truth annotations for the virtual objects in the synthetic clone.

11. The method of claim 9 , wherein the algorithm is for multi-object tracking.

12. The method of claim 5 , further comprising performing a specific task with an algorithm on the modified synthetic video.

13. The method of claim 12 , further comprising evaluating the performance of the algorithm against the set of ground truth annotations for the virtual objects the modified synthetic video.

14. The method of claim 12 , wherein the algorithm is for multi-object tracking.

15. The method of claim 1 , wherein the at least one object category is selected from a vehicle category, an animal category, a structure category, a signage category, and an environmental category.

16. A computer program product comprising a non-transitory recording medium storing instructions, which when executed on a computer, causes the computer to perform the method of claim 1 .

17. A system comprising memory which stores instructions for performing the method of claim 1 and a processor in communication with the memory which executes the instructions.

18. A system for assessing video performance analysis comprising:

a computer graphics engine component, which:

clones first and second sets of real-world data in a virtual world, the first set of real-world data being acquired by a camera and including a video sequence having visual components and objects, the second set of real-world data being automatically generated by at least one sensor and including sensor data having a set of physical measurements for the objects and visual components of the video sequence,

decomposes the visual components and objects in the video sequence of the first set of real world data into at least one object category, and

populates the virtual world with virtual visual components and virtual objects;

a scripting component which:

generates a script to control the virtual visual components and virtual objects in the virtual world based on the acquired visual components and objects in the first set of real-world data and the automatically generated sensor data for the objects and visual components in the second set of real-world data, and

generates a synthetic video sequence clone;

a modification component for changing a condition of one of the virtual visual components and the virtual objects;

an annotation component for generating a set of ground truth annotations for the virtual objects;

a performance component for performing a specific task with an algorithm and assessing a transferability of conclusions based on a performance of the algorithm; and

a processor which implements the computer graphics engine component, the scripting component, the modification component, the annotation component, and the performance component.

19. The system of claim 18 further comprising a seeding component for annotating the first and second sets of real-world data, enabling the graphics engine component to initialize the virtual world.

20. A method of assessing video performance analysis comprising:

acquiring a first set of real-world data with a camera, the first set of real-world data including a video sequence, the video sequence including visual components and objects;

automatically generating a second set of real-world data with at least one sensor, the second set of real-world data including sensor data, the sensor data including a set of physical measurements for the objects and visual components of the video sequence;

with a computer graphics engine, cloning the first and second sets of real-world data in a virtual world, comprising:

decomposing the visual components and objects in the video sequence of the first set of real world data into at least one object category, and

populating the virtual world with virtual visual components and virtual objects;

generating a script for controlling the virtual visual components and the virtual objects in the virtual world based on the acquired visual components and objects in the first set of real-world data and the automatically generated sensor data for the objects and visual components in the second set of real-world data;

changing a condition of one of the virtual visual components or the virtual objects;

generating a synthetic clone of the video sequence based on the script for controlling the virtual visual components and the virtual objects;

generating a modified synthetic video based on the changed condition;

generating a set of ground truth annotations for the virtual objects in the synthetic clone and the virtual objects in the modified synthetic video;

performing a specific task with an algorithm on the real-world data, on the synthetic clone of the video sequence, and on the modified synthetic video;

evaluating a performance of the algorithm against the set of ground truth annotations; and

assessing a transferability of conclusions from the virtual world to the real-world based on the performance of the algorithm.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2025
From: XEROX CORPORATION
To: GENESEE VALLEY INNOVATIONS, LLC
Reel/Frame 073562/0677 →
SECOND LIEN NOTES PATENT SECURITY AGREEMENT Recorded Jul 2, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 071785/0550 →
FIRST LIEN NOTES PATENT SECURITY AGREEMENT Recorded Apr 11, 2025
From: XEROX CORPORATION
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 070824/0001 →
SECURITY INTEREST Recorded Feb 13, 2024
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 066741/0001 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN PATENTS RECORDED AT RF 064760/0389 Recorded Feb 13, 2024
From: CITIBANK, N.A., AS COLLATERAL AGENT
To: XEROX CORPORATION
Reel/Frame 068261/0001 →
SECURITY INTEREST Recorded Nov 20, 2023
From: XEROX CORPORATION
To: JEFFERIES FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 065628/0019 →
SECURITY INTEREST Recorded Jun 22, 2023
From: XEROX CORPORATION
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 064760/0389 →
RELEASE OF SECURITY INTEREST IN PATENTS AT R/F 062740/0214 Recorded May 18, 2023
From: CITIBANK, N.A., AS AGENT
To: XEROX CORPORATION
Reel/Frame 063694/0122 →
SECURITY INTEREST Recorded Nov 10, 2022
From: XEROX CORPORATION
To: CITIBANK, N.A., AS AGENT
Reel/Frame 062740/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2016
From: WANG, QIAO; GAIDON, ADRIEN; VIG, ELEONORA
To: XEROX CORPORATION
Reel/Frame 037801/0897 →
Continuity (1)
Related Publication 20170243083A1 · Aug 24, 2017
Cited By (2)
US 12,415,418 US 12,561,957