IP Library Granted Patent US 12,675,903
Granted Patent B2
US 12,675,903 · App. 18/529,292 · Granted Jul 7, 2026

System for generating an image dataset for training an artificial intelligence model for object recognition, and method of use thereof

Inventor: Rémi Duquette (Montréal, CA)
G06T7/74G06F30/27G06T15/06G06T15/50G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,675,903
App. No.
18/529,292
Filed
Dec 5, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
2611
USPC
345/426
Abstract

A method of generating a synthetic image for producing a dataset of synthetic images for training an artificial intelligence model for object recognition; it includes generating a virtual space including a synthetic image background comprising one or more synthetic surfaces; spawning one or more virtual objects in the virtual space; simulating freefall and/or simulating environmental conditions of the one or more virtual objects resulting in a simulated impact between the one or more virtual objects and the one or more synthetic surfaces for determining the shape and position of the one or more virtual objects on the synthetic image background; and generating the synthetic image composed of the synthetic image background and the one or more virtual objects with the position of the one or more virtual objects on the synthetic image background determined from the simulated impact and/or simulated environmental conditions.

Claims (55)

1 . A method of generating a dataset of synthetic images for training an artificial intelligence model for object recognition, comprising repeating, for generating each synthetic image of the dataset of synthetic images:

generating a virtual space including a synthetic image background comprising one or more synthetic surfaces;

spawning one or more virtual objects in the virtual space, wherein the spawning comprises:

spawning a plurality of virtual objects in a hidden mode in the virtual space through use of masks that are dimensioned to contain the virtual objects of the plurality of virtual objects, the masks concealing the plurality of virtual objects in the virtual space;

selecting the one or more virtual objects from the plurality of virtual objects; and

revealing the one or more virtual objects for the simulation;

simulating a freefall of the one or more virtual objects resulting in a simulated impact between the one or more virtual objects and the one or more synthetic surfaces for determining a position of the one or more virtual objects on the synthetic image background; and

generating the synthetic image composed of the synthetic image background and the one or more virtual objects with the position of the one or more virtual objects on the synthetic image background determined from the simulated impact.

2 . The method as defined in claim 1 , wherein the one or more virtual objects comprises one or more physical properties, and wherein the outcome of the simulated impact further accounts for the one or more physical properties of the one or more virtual objects.

3 . The method as defined in claim 2 , wherein the one or more physical properties of the one or more virtual objects comprises one or more of:

a type of object;

an object dimension;

an object material;

a center of mass of the object; and

an object weight.

4 . The method as defined in claim 1 , wherein the plurality of virtual objects forms an object map.

5 . The method as defined in claim 1 , further comprising defining an object zone in the virtual space, where the one or more virtual objects are to be positioned in the object zone following the simulated impact.

6 . The method as defined in claim 5 , further comprising removing virtual objects of the one or more virtual objects that are outside of the object zone following the simulated impact or subsequent displacement from the surface of the object zone.

7 . The method as defined in claim 5 , wherein the spawning comprises offsetting the one or more virtual objects from the object zone such that the one or more virtual objects spawn over the object zone while the virtual objects of the one or more virtual objects are not in contact prior to the simulating.

8 . The method as defined in claim 1 , further comprising applying lighting with ray tracing to the one or more virtual objects and the synthetic object background.

9 . The method as defined in claim 8 , wherein the lighting is a first lighting, and further comprising applying a second static lighting above an area of the virtual space where the one or more virtual objects are spawned.

10 . The method as defined in claim 1 , wherein an initial orientation of each of the virtual objects of the one or more virtual objects is determined randomly prior to the simulating.

11 . The method as defined in claim 1 , wherein the generating of the synthetic image is taken at a viewing angle of a virtual camera object, wherein the position of the virtual camera object in the virtual space is determined randomly.

12 . The method as defined in claim 1 , further comprising generating a masked image corresponding to the synthetic image from which the coordinates of a rectangular box containing the object is registered.

13 . The method as defined in claim 1 , wherein the one or more synthetic surfaces each comprise a material texture having one or more physical properties and the simulated impact is generated as a function of at least the one or more physical properties of the one or more synthetic surfaces.

14 . The method as defined in claim 1 , further comprising simulating environmental conditions affecting the position or size of the one or more virtual objects.

15 . A non-transitory storage medium having stored thereon a dataset of synthetic images for training an artificial intelligence model for object recognition generated by performing the method as defined in claim 1 .

16 . A method of training an artificial intelligence model for object recognition using the dataset of synthetic images as defined in claim 15 , comprising:

receiving the dataset of synthetic images;

causing an identification and location of one or more virtual objects in a synthetic image of the dataset of synthetic images; and

comparing the synthetic image with the one or more identified virtual objects to a ground truth to calculate a loss that is back-propagated through the artificial intelligence model to train the artificial intelligence model.

17 . The method as defined in claim 1 , wherein one or more of the following is varied to generate different synthetic images forming the dataset of synthetic images:

lighting with ray tracing to the one or more virtual objects and the synthetic object background;

a viewing angle of a virtual camera related to the perspective taken for the generated synthetic image;

the material texture of the one or more synthetic surfaces; and

one or more types of the one or more virtual objects.

18 . The method according to claim 1 , wherein the plurality of objects forms an object map that occupies a part of the virtual space or a plane of the virtual space, occupying the virtual space in a hidden mode, wherein the object map is a constellation of the plurality of virtual objects, each virtual object of the plurality of virtual objects occupying a position in the virtual space defined in the x, y and z axes, each virtual object of the object map defined at a height (z) above the one or more synthetic surfaces.

19 . A system configured to generate a dataset of synthetic images for training an artificial intelligence model for object recognition, comprising:

a processor;

memory further including program code that, when executed by the processor, causes the processor to repeat, for generating each synthetic image of the dataset of synthetic images:

generating a virtual space including a synthetic image background comprising one or more synthetic surfaces;

spawning one or more virtual objects in the virtual space, wherein the spawning comprises:

spawning a plurality of virtual objects in a hidden mode in the virtual space through use of masks that are dimensioned to contain the virtual objects of the plurality of virtual objects, the masks concealing the plurality of virtual objects in the virtual space;

selecting the one or more virtual objects from the plurality of virtual objects; and

revealing the one or more virtual objects for the simulation;

simulating a freefall of the one or more virtual objects resulting in a simulated impact between the one or more virtual objects and the one or more synthetic surfaces for determining a position the one or more virtual objects on the synthetic image background; and

generating the synthetic image composed of the synthetic image background and the one or more virtual objects with the position of the one or more virtual objects on the synthetic image background determined from the simulated impact.

20 . A non-transitory computer-readable medium having stored thereon program instructions for generating a dataset of synthetic images for training an artificial intelligence model for object recognition training an artificial intelligence model for object recognition, the program instructions executable by a processing unit for repeating, for generating each synthetic image of the dataset of synthetic images:

generating a virtual space including a synthetic image background comprising one or more synthetic surfaces;

spawning one or more virtual objects in the virtual space, wherein the spawning comprises:

spawning a plurality of virtual objects in a hidden mode in the virtual space through use of masks that are dimensioned to contain the virtual objects of the plurality of virtual objects, the masks concealing the plurality of virtual objects in the virtual space;

selecting the one or more virtual objects from the plurality of virtual objects; and

revealing the one or more virtual objects for the simulation;

simulating a freefall of the one or more virtual objects resulting in a simulated impact between the one or more virtual objects and the one or more synthetic surfaces (including environmental simulation effects) for determining a position the one or more virtual objects on the synthetic image background; and

generating the synthetic image composed of the synthetic image background and the one or more virtual objects with the position of the one or more virtual objects on the synthetic image background determined from the simulated impact and environmental condition simulation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2024
From: DUQUETTE, RÉMI
To: MAYA HEAT TRANSFER TECHNOLOGIES LTD.
Reel/Frame 066120/0524 →
Continuity (2)
Provisional Application 63479878 · Jan 13, 2023
Related Publication 20240242380A1 · Jul 18, 2024
References Cited (70)
US 6522312B2 · Ohshima · 2003 [cited by examiner]
US 7948485B1 · Larsen · 2011 [cited by examiner]
US 8422797B2 · Heisele · 2013 [cited by examiner]
US 9665800B1 · Kuffner, Jr. · 2017 [cited by examiner]
US 10277859B2 · Armeni · 2019 [cited by examiner]
US 11182649B2 · Tremblay · 2021 [cited by examiner]
US 11257272B2 · Rowell · 2022 [cited by examiner]
US 11341699B1 · Gottlieb · 2022 [cited by examiner]
US 11944904B2 · Liang · 2024 [cited by examiner]
US 12037013B1 · Linscott · 2024 [cited by examiner]
US 12272019B1 · Iwanowski · 2025 [cited by examiner]
US 20160314613A1 · Nowozin · 2016 [cited by examiner]
US 20160342861A1 · Tuzel · 2016 [cited by examiner]
US 20170355078A1 · Ur · 2017 [cited by examiner]
US 20180232601A1 · Feng · 2018 [cited by examiner]
US 20180250821A1 · Shimodaira · 2018 [cited by examiner]
US 20200125880A1 · Wang · 2020 [cited by examiner]
US 20200167606A1 · Wohlhart · 2020 [cited by examiner]
US 20200388021A1 · Song · 2020 [cited by examiner]
US 20210056247A1 · Audren · 2021 [cited by examiner]
US 20210064122A1 · Fujinawa · 2021 [cited by examiner]
US 20210103776A1 · Jiang · 2021 [cited by examiner]
US 20210182664A1 · Kim · 2021 [cited by examiner]
US 20210197524A1 · Maroli · 2021 [cited by examiner]
US 20210217144A1 · Snellgrove · 2021 [cited by examiner]
US 20210295068A1 · Bergt · 2021 [cited by examiner]
US 20210327127A1 · Hinterstoisser · 2021 [cited by examiner]
US 20220044441A1 · Kalra · 2022 [cited by examiner]
US 20220083807A1 · Zhang · 2022 [cited by examiner]
US 20220215266A1 · Venkataraman · 2022 [cited by examiner]
US 20220237336A1 · Zhao · 2022 [cited by examiner]
US 20220343537A1 · Taamazyan · 2022 [cited by examiner]
US 20230078763A1 · Hata · 2023 [cited by examiner]
US 20230144893A1 · Chan · 2023 [cited by examiner]
US 20230169677A1 · Llic · 2023 [cited by examiner]
US 20230234231A1 · Dong · 2023 [cited by examiner]
US 20230245269A1 · Trott · 2023 [cited by examiner]
US 20230245364A1 · Chen · 2023 [cited by examiner]
US 20230339502A1 · Chi-Johnston · 2023 [cited by examiner]
US 20230339519A1 · Chi-Johnston · 2023 [cited by examiner]
US 20230419532A1 · Strunz · 2023 [cited by examiner]
US 20240037925A1 · Lim · 2024 [cited by examiner]
US 20240054268A1 · Behandish · 2024 [cited by examiner]
US 20240070896A1 · Hayashi · 2024 [cited by examiner]
US 20240104431A1 · Choi · 2024 [cited by examiner]
US 20240135618A1 · Zhang · 2024 [cited by examiner]
US 20240161427A1 · Biehl · 2024 [cited by examiner]
US 20240173855A1 · Thon · 2024 [cited by examiner]
US 20240193845A1 · Hwang · 2024 [cited by examiner]
US 20240221309A1 · Chai · 2024 [cited by examiner]
US 20240303773A1 · Sartor · 2024 [cited by examiner]
US 20240370223A1 · Varekamp · 2024 [cited by examiner]
US 20250054265A1 · Zhu · 2025 [cited by examiner]
US 20250054270A1 · Yamauchi · 2025 [cited by examiner]
US 20250061591A1 · Krupani · 2025 [cited by examiner]
US 20250077765A1 · Xu · 2025 [cited by examiner]
US 20250173891A1 · Manzur · 2025 [cited by examiner]
US 20250181969A1 · Wen · 2025 [cited by examiner]
Sun, Baochen, Xingchao Peng, and Kate Saenko. “Generating large scale image datasets from 3D CAD models.” CVPR 2015 Workshop on the future of datasets in vision. vol. 2, 2015. https://ai.bu.edu/pubs/cvpr2015_workshop_vi… [cited by applicant]
Sooyoung, C., et al. “How to generate image dataset based on 3D model and deep learning method.” Int. J. Eng. Technol 7 (2018): 221-225. https://www.sciencepubco.com/index.php/ijet/article/download/18969/8694. [cited by applicant]
Peng, Xingchao, et al. “Learning deep object detectors from 3d models.” Proceedings of the IEEE international conference on computer vision. 2015. https://arxiv.org/pdf/1412.7122.pdf. [cited by applicant]
Planche, Benjamin, et al. “Depthsynth: Real-time realistic synthetic data generation from cad models for 2.5 d recognition.” 2017 International Conference on 3D Vision (3DV). IEEE, 2017. https://arxiv.org/pdf/1702.08558… [cited by applicant]
Martinez-Gonzalez, Pablo, et al. “Unrealrox: an extremely photorealistic virtual reality environment for robotics simulations and synthetic data generation.” Virtual Reality 24.2 (2020): 271-288. https://arxiv.org/pdf/1… [cited by applicant]
Martinez-Gonzalez, Pablo, et al. “UnrealROX+: An Improved Tool for Acquiring Synthetic Data from Virtual 3D Environments.” 2021 International Joint Conference on Neural Networks (IJCNN). IEEE, 2021. https://arxiv.org/pd… [cited by applicant]
Borkman, Steve, et al. “Unity perception: Generate synthetic data for computer vision.” arXiv preprint arXiv:2107.04259 (2021). https://arxiv.org/pdf/2107.04259.pdf. [cited by applicant]
Wang, Qi, et al. “Learning from synthetic data for crowd counting in the wild.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. https://openaccess.thecvf.com/content_CVPR_2019/pa… [cited by applicant]
Mu, Jiteng, et al. “Learning from synthetic animals.” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2020. https://openaccess.thecvf.com/content_CVPR_2020/papers/Mu_Learning_From_Synt… [cited by applicant]
Ros, German, et al. “The synthia dataset: A large collection of synthetic images for semantic segmentation of urban scenes.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. https://w… [cited by applicant]
YouTube video source: Training AI with 3D Rendered Images https://www.youtube.com/watch?v=WL1yK7RtHT4. Posted Jun. 8, 2020. Retrieved Oct. 27, 2022. [cited by applicant]
YouTube video source: GCC dataset Collector and Labeler (GCC-CL): A generater of crowd scenes. https://www.youtube.com/watch?v=Hvl7xWklueo. Posted Mar. 5, 2019. Retrieved Oct. 27, 2022. [cited by applicant]