IP Library Granted Patent US 11,335,019
Granted Patent B2
US 11,335,019 · App. 15/556,596 · Granted May 17, 2022

Method for the 3D reconstruction of a scene

Inventors: Ieng Sio-Hoi (Montreuil, FR); Benosman Ryad (Pantin, FR); Shi Bertram (Hong Kong, CN)
Assignees: SORBONNE UNIVERSITÉ; CENTRE NATIONAL DE LA RECHERCHE SCIENTIFIQUE—CNRS; INSERM (INSTITUT NATIONAL DE LA SANTE ET DE LA RECHERCHE MEDICALE)
G06T7/593H04N13/239H04N13/257H04N13/296H04N2013/0081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,335,019
App. No.
15/556,596
Granted
May 17, 2022
Kind
B2
Abstract

The present invention concerns a method for the 3D reconstruction of a scene comprising the matching ( 610 ) of a first event from among the first asynchronous successive events of a first sensor with a second event from among the second asynchronous successive events of a second sensor depending on a minimisation ( 609 ) of a cost function (E). The cost function comprises at least one component from: —a luminance component (E i ) that depends at least on a first luminance signal (I u ) convoluted with a convolution core (gσ(t)), the luminance of said pixel depending on a difference between maximums (t e−,u ,t e+,u ) of said first signal; and a second luminance signal (I v ) convoluted with said convolution core, the luminance of said pixel depending on a difference between maximums (t e−,v ,t e+,v ) of said second signal; —a movement component (E M ) depending on at least time values relative to the occurrence of events located at a distance from a pixel of the first sensor and time values relative to the occurrence of events located at a distance from a pixel of the second sensor.

Claims (58)

1. A method of 3D reconstruction of a scene, the method comprising:

receiving a first piece of asynchronous information from a first vision sensor that has a first pixel matrix positioned opposite the scene, the first piece of asynchronous information comprising, for each pixel of the first matrix, the first successive events coming from said pixel of the first matrix;

receiving a second piece of asynchronous information from a second vision sensor that has a second pixel matrix positioned opposite the scene, the second piece of asynchronous information comprising, for each pixel of the second matrix, the second successive events coming from said pixel of the second matrix, the second sensor being separate from the first sensor;

storing the first piece of asynchronous information from the first vision sensor and the second piece of asynchronous information from the second vision sensor at a memory;

matching, by a processor in communication with the memory, a first event from among the first successive events coming from the pixel of the first matrix with a second event from among the second successive events coming from the pixel of the second matrix, the second event being determined to minimize a cost function for the first event, the cost function being a simple sum or a weighted sum of a luminance component and a movement component and a time component and a geometric component for the pixel of the first matrix and the pixel of the second matrix; and

determining, by the processor, a 3D reconstruction of a scene based on the matching,

the luminance component depending on at least:

a first luminance signal coming from the pixel of the first sensor convoluted with a convolution core, the luminance of said first sensor pixel depending on a difference between the maximums of said first signal, and

a second luminance signal coming from the pixel of the second sensor convoluted with said convolution core, the luminance of said second sensor pixel depending on a difference between the maximums of said second signal,

the movement component depending on at least:

time values relating to the occurrence of events spatially located at a predetermined distance from the pixel of the first sensor, and

time values relating to the occurrence of events spatially located at a predetermined distance from the pixel of the second sensor,

the time component depending on a difference between:

a first time value relating to one of the first successive events of the first sensor, and

a second time value relating to one of the second successive events of the second sensor, and

the geometric component depending on:

a spatial distance from the pixel of the second sensor at an epipolar straight line or at an epipolar intersection defined by at least one pixel of the first sensor.

2. The method according to claim 1 , wherein the first luminance signal of the pixel of the first sensor and the second luminance signal of the pixel of the second sensor comprise a maximum, coding an occurrence time of a luminance variation, the convolution core being a predetermined Gaussian variance.

3. The method according to claim 1 , wherein said luminance component additionally depends on:

luminance signals of pixels of the first sensor, spatially located at a predetermined distance from the pixel of the first sensor, convoluted with the convolution core, and

luminance signals of pixels of the second sensor, spatially located at a predetermined distance from the pixel of the second sensor, convoluted with the convolution core.

4. The method according to claim 1 , wherein said movement component depends on:

an average value of the time values relating to the occurrence of pixel events of the first sensor, spatially located at a predetermined distance from the pixel of the first sensor, and

an average value of the time values relating to the occurrence of pixel events of the second sensor, spatially located at a predetermined distance from the pixel of the second sensor.

5. The method according to claim 1 , wherein said movement component depends on, for a given time:

for each current time value relating to the occurrence of events, spatially located at the predetermined distance from the pixel of the first sensor, of a function value decreasing from a distance of said given time to said current time value, and

for each current time value relating to the occurrence of events, spatially located at the predetermined distance from the pixel of the second sensor, of a function value decreasing from a distance of said given time to said current time value.

6. The method according to claim 1 , wherein said movement component depends on:

a first convolution of a decreasing function with a signal comprising a Dirac for each time value relating to the occurrence of events, spatially located at the predetermined distance from the pixel of the first sensor, and

a second convolution of a decreasing function with a signal comprising a Dirac for each time value relating to the occurrence of events, spatially located at the predetermined distance from the pixel of the second sensor.

7. A device for the 3D reconstruction of a scene, the device comprising:

an interface configured to

receive a first piece of asynchronous information from a first sensor that has a first pixel matrix positioned opposite the scene, the first piece of asynchronous information comprising, for each pixel of the first matrix, the first successive events coming from said pixel, and

receive a second piece of asynchronous information from a second sensor that has a second pixel matrix positioned opposite the scene, the second piece of asynchronous information comprising, for each pixel of the second matrix, the second successive events coming from said pixel, the second sensor being separate from the first sensor; and

a processor configured to

match a first event from among the first successive events coming from the pixel of the first matrix with a second event from among the second successive events coming from the pixel of the second matrix, the second event being determined to minimize a cost function for the first event, the cost function being a simple sum or a weighted sum of a luminance component and a movement component and a time component and a geometric component for the pixel of the first matrix and the pixel of the second matrix, and

determine a 3D reconstruction of a scene based on the matching,

the luminance component depending on at least:

a first luminance signal coming from the pixel of the first sensor, convoluted with a convolution core, the luminance of said first sensor pixel depending on a difference between the maximums of said first signal, and

a second luminance signal coming from the pixel of the second sensor, convoluted with said convolution core, the luminance of said second sensor pixel depending on a difference between the maximums of said second signal, and

the movement component depending on at least:

time values relating to the occurrence of events, spatially located at a predetermined distance from the pixel of the first sensor, and

time values relating to the occurrence of event, spatially located at a predetermined distance from the pixel of the second sensor,

the time component depending on a difference between:

a first time value relating to one of the first successive events of the first sensor, and

a second time value relating to one of the second successive events of the second sensor, and

the geometric component depending on:

a spatial distance from the pixel of the second sensor at an epipolar straight line or at an epipolar intersection defined by at least one pixel of the first sensor.

8. A non-transitory computer-readable storage medium comprising instructions for the implementation of the method according to claim 1 , when the method is executed by a processor.

9. The method according to claim 2 , wherein said luminance component additionally depends on:

luminance signals of pixels of the first sensor, spatially located at a predetermined distance from the pixel of the first sensor, convoluted with the convolution core, and

luminance signals of pixels of the second sensor, spatially located at a predetermined distance from the pixel of the second sensor, convoluted with the convolution core.

10. The method according to claim 2 , wherein said movement component depends on:

an average value of the time values relating to the occurrence of pixel events of the first sensor, spatially located at a predetermined distance from the pixel of the first sensor, and

an average value of the time values relating to the occurrence of pixel events of the second sensor, spatially located at a predetermined distance from the pixel of the second sensor.

11. The method according to claim 3 , wherein said movement component depends on:

an average value of the time values relating to the occurrence of pixel events of the first sensor, spatially located at a predetermined distance from the pixel of the first sensor, and

an average value of the time values relating to the occurrence of pixel events of the second sensor, spatially located at a predetermined distance from the pixel of the second sensor.

Assignments (2)
MERGER Recorded Feb 22, 2019
From: UNIVERSITÉ PIERRE ET MARIE CURIE (PARIS 6); UNIVERSITÉ PARIS-SORBONNE
To: SORBONNE UNIVERSITÉ
Reel/Frame 048407/0681 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2018
From: SIO-HOI, IENG; RYAD, BENOSMAN; BERTRAM, SHI
To: UNIVERSITE PIERRE ET MARIE CURIE (PARIS 6); CENTRE NATIONAL DE LA RECHERCHE SCIENTIFIQUE - CNRS; INSERM (INSTITUT NATIONAL DE LA SANTE ET DE LA RECHERCHE MEDICALE)
Reel/Frame 046194/0251 →
Priority Claims (1)
FR 15 52154 · Mar 16, 2015 · national
Continuity (1)
Related Publication 20180063506A1 · Mar 1, 2018