IP Library › Granted Patent US 9,432,625
Granted Patent B2
US 9,432,625 · App. 14/431,101 · Granted Aug 30, 2016

Immersive videoconference method and system

Inventors: Gerard Delegue (Nozay, FR); Nicolas Bouche (Nozay, FR)
Assignee: Alcatel Lucent
H04N7/157H04M3/567H04N13/0239H04N2213/003
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,432,625
App. No.
14/431,101
Granted
Aug 30, 2016
Kind
B2
Abstract

An immersive videoconference method wherein multiple participants ( 21, 22, 23, 24 ) in different locations ( 11, 12, 13 ) remotely interact with each other through a telecommunication network architecture ( 8, 31, 38 ), wherein the method comprises at the location ( 11, 12, 13 ) of a given participant ( 21, 22, 23, 24 ); —capturing video images of the participant by a pair of video cameras ( 4 A, 4 B); —detecting, tracking and determining size and position related parameters of the participant in the video images; —generating a single elementary video stream related to the participant; —associating a room identifier to the elementary video stream, the room identifier being uniquely associated to the given participant; —sending the elementary video stream, the size and position related parameters and the room identifier ( 41 A, 42 A, 43 A) to a centralized entity ( 30 ); —repeating the above steps for each participant ( 21, 22, 23, 24 ) at the different location ( 11, 12, 13 ); wherein the method further comprises at the centralized entity ( 30 ): —creating a virtual room ( 70 ) by combining the elementary video streams ( 41 A, 42 A, 43 A) for all the participants; —staging the elementary video streams of all the participants in said virtual room and computing a scene specification associated to the room identifier of each participant based on the size and position related parameters of all the participants; and —generating, for each participant, a single composite video stream ( 41 B, 42 B, 43 B) of the virtual room ( 70 ) that displays the 2D video of the other participants sized and positioned as if the participants ( 21, 22, 23, 24 ) were in the same virtual room ( 70 ) based on the scene specification and a combination of the elementary video streams of the other participants.

Claims (40)

1. An immersive videoconference method allowing multiple participants in different locations to remotely interact with each other through a telecommunication network architecture, wherein the method comprises at the location of a given participant:

capturing video images of the participant by a pair of video cameras;

detecting, tracking and determining size and position related parameters of the participant in the video images;

generating a single elementary video stream related to the participant;

associating a room identifier to the elementary video stream, the room identifier being uniquely associated to the given participant;

sending the elementary video stream, the size and position related parameters and the room identifier to a centralized entity;

repeating the above for each participant at each different location;

wherein the method further comprises at the centralized entity:

creating a virtual room by combining the elementary video streams for all the participants;

staging the elementary video streams of all the participants in said virtual room and computing a scene specification associated to the room identifier of each participant based on the size and position related parameters of all the participants; and

generating, for each participant, a single composite video stream of the virtual room that displays the 2D video of the other participants sized and positioned as if the participants were in the same virtual room based on the scene specification and a combination of the elementary video streams of the other participants;

wherein detecting and tracking the participant in the video images comprises detecting and tracking a body of the participant without a background from the video images based on a histograms of oriented gradients HOG for the purpose of human detection algorithm and wherein results of said HOG algorithm are further filtered by a depth mapping matrix computed from a pair of video signals of the participant obtained from the pair of video cameras.

2. The immersive videoconference method of claim 1 , wherein the depth mapping matrix is computed based on a pinhole camera model.

3. The immersive videoconference method according to claim 1 , wherein detecting and tracking the participant in the video images comprises determining a 3D position of the participant relatively to a position of one of the video camera based on a binary mask image and the depth mapping matrix.

4. The immersive videoconference method according to claim 1 , wherein generating the elementary video stream comprises encoding images of the elementary video stream with a textured mask, the elementary video stream being a Red Green Blue and Alpha video stream with alpha being the level of transparency.

5. The immersive videoconference method according to claim 1 , wherein generating one composite video stream for the participant comprises translating, zooming and superimposing the elementary video streams received from the other participants based on the scene specification.

6. The immersive videoconference method according to claim 1 , wherein the method further comprises only publishing and displaying said single composite video stream to an appropriate participant based on the corresponding unique room identifier.

7. An immersive videoconference method allowing multiple participants in different locations to remotely interact with each other through a telecommunication network architecture, wherein the method comprises at the location of a given participant:

capturing video images of the participant by a pair of video cameras;

detecting, tracking and determining size and position related parameters of the participant in the video images;

generating a single elementary video stream related to the participant;

associating a room identifier to the elementary video stream, the room identifier being uniquely associated to the given participant;

sending the elementary video stream, the size and position related parameters and the room identifier to a centralized entity;

repeating the above for each participant at each different location;

wherein the method further comprises at the centralized entity:

creating a virtual room by combining the elementary video streams for all the participants;

staging the elementary video streams of all the participants in said virtual room and computing a scene specification associated to the room identifier of each participant based on the size and position related parameters of all the participants; and

generating, for each participant, a single composite video stream of the virtual room that displays the 2D video of the other participants sized and positioned as if the participants were in the same virtual room based on the scene specification and a combination of the elementary video streams of the other participants;

wherein the scene specification comprises z-indexes of the elementary video streams describing whether an elementary video stream related to one participant is in front or behind other elementary video streams related to the other participants in the virtual room, a 2D position of each video describing the positions of each participant relatively to a given point of view in the virtual room, and a zoom scale describing the proximity of one participant relatively to another one.

8. An immersive videoconference system wherein multiple participants in different locations remotely interact with each other through a telecommunication network architecture, the immersive videoconference system comprising:

a pair of video cameras, at the location of each participant, arranged to capture video images of the participant;

a pretreatment module, at the location of each participant, comprising a depth map generator coupled to a tracker arranged to detect and track the participant in the video images, a body position calculator arranged to determine size and position related parameters of the participant in the video images, a video streamer arranged to generate a single elementary video stream related to the participant, and a room identifier requestor arranged to associate a room identifier to the elementary video stream; and

a virtual place building module, at a centralized location, comprising a staging director arranged to create a virtual room by combining the elementary video streams for all the participants, stage the elementary video streams of all the participants in said virtual room and compute a scene specification associated to the room identifier of each participant based on the size and position related parameters of all the participants, and a video mixer arranged to generate, for each participant, a single composite video stream of the virtual room that displays the 2D video of the other participants sized and positioned as if the participants were in the same virtual room based on the scene specification and a combination of the elementary video streams of the other participants;

wherein the tracker is arranged to detect and track a body of the participant without a background from the video images based on a histograms of oriented gradients HOG for the purpose of human detection algorithm and wherein results of said HOG algorithm are further filtered by a depth mapping matrix computed from a pair of video signals of the participant obtained from the pair of video cameras.

9. The immersive videoconference system of claim 8 , wherein the virtual place building module further comprises a video server arranged to publish the composite video streams of the participants, each video stream being associate with a room identifier uniquely associated to the given participant.

10. An immersive videoconference system wherein multiple participants in different locations remotely interact with each other through a telecommunication network architecture, the immersive videoconference system comprising:

a pair of video cameras, at the location of each participant, arranged to capture video images of the participant;

a pretreatment module, at the location of each participant, comprising a depth map generator coupled to a tracker arranged to detect and track the participant in the video images, a body position calculator arranged to determine size and position related parameters of the participant in the video images, a video streamer arranged to generate a single elementary video stream related to the participant, and a room identifier requestor arranged to associate a room identifier to the elementary video stream; and

a virtual place building module, at a centralized location, comprising a staging director arranged to create a virtual room by combining the elementary video streams for all the participants, stage the elementary video streams of all the participants in said virtual room and compute a scene specification associated to the room identifier of each participant based on the size and position related parameters of all the participants, and a video mixer arranged to generate, for each participant, a single composite video stream of the virtual room that displays the 2D video of the other participants sized and positioned as if the participants were in the same virtual room based on the scene specification and a combination of the elementary video streams of the other participants;

wherein the scene specification comprises z-indexes of the elementary video streams describing whether an elementary video stream related to one participant is in front or behind other elementary video streams related to the other participants in the virtual room, a 2D position of each video describing the positions of each participant relatively to a given point of view in the virtual room, and a zoom scale describing the proximity of one participant relatively to another one.

Assignments (11)
PATENT SECURITY AGREEMENT Recorded Aug 6, 2024
From: RPX CORPORATION; RPX CLEARINGHOUSE LLC
To: BARINGS FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 068328/0674 →
RELEASE OF LIEN ON PATENTS Recorded Aug 5, 2024
From: BARINGS FINANCE LLC
To: RPX CORPORATION
Reel/Frame 068328/0278 →
PATENT SECURITY AGREEMENT Recorded Apr 22, 2023
From: RPX CORPORATION
To: BARINGS FINANCE LLC, AS COLLATERAL AGENT
Reel/Frame 063429/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: PROVENANCE ASSET GROUP LLC
To: RPX CORPORATION
Reel/Frame 059352/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: NOKIA US HOLDINGS INC.
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058363/0723 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: CORTLAND CAPITAL MARKETS SERVICES LLC
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058983/0104 →
ASSIGNMENT AND ASSUMPTION AGREEMENT Recorded Feb 14, 2019
From: NOKIA USA INC.
To: NOKIA US HOLDINGS INC.
Reel/Frame 048370/0682 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP, LLC
To: CORTLAND CAPITAL MARKET SERVICES, LLC
Reel/Frame 043967/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2017
From: NOKIA TECHNOLOGIES OY; NOKIA SOLUTIONS AND NETWORKS BV; ALCATEL LUCENT SAS
To: PROVENANCE ASSET GROUP LLC
Reel/Frame 043877/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP LLC
To: NOKIA USA INC.
Reel/Frame 043879/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2015
From: DELEGUE, GERARD; BOUCHE, NICOLAS
To: ALCATEL LUCENT
Reel/Frame 035253/0279 →
Priority Claims (1)
EP 12186744 · Sep 28, 2012 · regional
Continuity (1)
Related Publication 20150244987A1 · Aug 27, 2015