IP Library Granted Patent US 11,700,286
Granted Patent B2
US 11,700,286 · App. 17/577,875 · Granted Jul 11, 2023

Multiuser asymmetric immersive teleconferencing with synthesized audio-visual feed

Inventors: Devon Copley (San Francisco, CA); Prasad Balasubramanian (Fremont, CA)
Assignee: Avatour Technologies, Inc.
H04L65/1066H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,700,286
App. No.
17/577,875
Granted
Jul 11, 2023
Kind
B2
Abstract

Embodiments described herein enable a teleconference among host-side user(s) at a host site and remotely located client-side user(s), wherein a host device is located at the host site, and wherein each of the client-side user(s) uses a respective client device to participate in the teleconference. A host site audio-visual feed is received from the host device. A client data feed is received from the client device of each client-side user. Orientation information for the respective client-side user using a client device is also received. Each client device is provided with the host site audio-visual feed or a modified version thereof. The host device is provided with, for each client device, the client data feed and the orientation information of the client-side user that is using the client device. This enables host-side and client-side users to display visual representations of one another with their respective orientations.

Claims (54)

1. A method for enabling a teleconference among one or more host-side users located at a host site and one or more client-side users that are remotely located relative to the host site, wherein a host device is located at the host site, and wherein each of the one or more client-side users uses a respective client device to participate in the teleconference, the method comprising:

receiving from the host device a host site audio-visual feed that includes audio and video of the host site, captured from the host site by the host device;

at least one of producing or updating a three-dimensional (3D) model of the host site using the host site audio-visual feed;

assigning or otherwise obtaining for each client-side user, of the one or more client-side users, a virtual location within the 3D model of the host site and thereby mapping each client-side user to a corresponding real-world location at the host site;

allocating a respective virtual camera rendering engine to each of the one or more client-side users;

using a said virtual camera rendering engine, allocated to a said client-side user, to render a synthesized audio-visual feed that appears as if it were generated by a virtual camera located at the virtual location of the said client-side user within the 3D model of the host site;

providing the synthesized audio-visual feed, or a modified version thereof, to the client device of the said client-side user to thereby enable consumption thereof by the said client-side user;

receiving, from the client device of the said client-side user, orientation information that includes an indication of a visual portion of the synthesized audio-visual feed, or the modified version thereof, that the said client-side user is viewing at any given time, which corresponds to less than a full field of view (FOV) provided by the synthesized audio-visual feed; and

using the said virtual camera rendering engine, allocated to the said client-side user, to update the said synthesized audio-visual feed for the said client-side user, in response to at least one of the virtual location of the said client-side user changing or the orientation information for the said client-side user changing.

2. The method of claim 1 , wherein the synthesized audio-visual feed rendered for the said client-side user changes in response to the said client side user's virtual location within the 3D model of the host site changing, such that when the said client-side user moves from a first virtual location within the 3D model of the host site to a second virtual location within the 3D model of the host site, the synthesized audio-visual feed changes from appearing as if it were generated by a virtual camera located at the first virtual location to appearing as if it were generated by a virtual camera located at the second virtual location.

3. The method of claim 1 , further comprising:

modifying the synthesized audio-visual feed produced for the said client-side user in near real-time, based on the orientation information, to produce one or more modified versions thereof that include visual data for less than the full FOV visual data for the 3D model of the host site and/or emphasize the visual portion of the synthesized audio-visual feed being viewed on the said client device at any given time, to thereby reduce an amount of data that is provided to the said client device.

4. The method of claim 1 , further comprising causing to display on an augmented reality (AR) device being used by a said host-side user, a visual representation of at least one of the one or more client-side users at the real-world location within the host site that corresponds to their respective virtual location within the 3D model of the host site.

5. The method of claim 4 , wherein when the virtual location of the said client-side user changes, the visual representation of the said client-side user displayed on the AR device changes, such that when the said client-side user moves from a first virtual location within the 3D model of the host site to a second virtual location within the 3D model of the host site, the visual representation of the said client-side user changes from being at a first corresponding real-world location within the host site to being at a second corresponding real-world location within the host site.

6. The method of claim 1 , further comprising causing to display on an augmented reality (AR) device being used by a said host-side user, a visual representation of the said client-side user of the one or more client-side users with their respective orientation, which provides an indication of a visual portion of the synthesized audio-visual feed that the said client-side user is viewing at any given time.

7. The method of claim 1 , wherein the one or more client-side users comprises a plurality of client-side users, and wherein a separate said synthesized audio-visual feed is is rendered using a separate said virtual camera rendering engine for each of at least two of the plurality of client-side users.

8. The method of claim 7 , wherein the synthesized audio-visual feed rendered for a said client side user includes an avatar of another said client side user at their respective virtual location.

9. The method of claim 7 , wherein:

the synthesized audio-visual feed rendered for a said client side user includes a real-world representation of another said client side user at their respective virtual location; and

the real-world representation is captured using a webcam or another type of camera able to capture real-world video.

10. The method of claim 1 , further comprising:

for the said synthesized audio-visual feed rendered for the said client-side user, using the said virtual camera rendering engine allocated to the said client-side user, modifying the said synthesized audio-visual feed in near real-time, based on the orientation information received for the said client-side user, to produce one or more modified versions thereof that include visual data for less than the full FOV generated by the said virtual camera rendering engine allocated to the said client-side user and/or emphasize the visual portion of the synthesized audio-visual feed being viewed at any given time; and

providing to the said client device a said modified version of the host site audio-visual feed resulting from the modifying, to thereby enable consumption thereof by the said client-side user that is using the said client device.

11. A system for enabling a teleconference among one or more host-side users located at a host site and one or more client-side users that are remotely located relative to the host site, wherein a host device is located at the host site, and wherein each of the one or more client-side users uses a respective client device to participate in the teleconference, the system comprising one or more processor configured to:

receive from the host device a host site audio-visual feed that includes audio and video of the host site, captured from the host site by the host device;

at least one of produce or update a three-dimensional (3D) model of the host site using the host site audio-visual feed;

assign or otherwise obtain for each client-side user, of the one or more client-side users, a virtual location within the 3D model of the host site and thereby mapping each client-side user to a corresponding real-world location at the host site;

allocate a respective virtual camera rendering engine to each of the one or more client-side users;

use a said virtual camera rendering engine, allocated to producc for a said client-side user, to render a synthesized audio-visual feed that appears as if it were generated by a virtual camera located at the virtual location of the said client-side user within the 3D model of the host site;

provide the synthesized audio-visual feed, or a modified version thereof, to the client device of the said client-side user to thereby enable consumption thereof by the said client-side user;

receive, from the client device of the said client-side user, orientation information that includes an indication of a visual portion of the synthesized audio-visual feed, or the modified version thereof, that the said client-side user is viewing at any given time, which corresponds to less than a full field of view (FOV) provided by the synthesized audio- visual feed; and

use the said virtual camera rendering engine, allocated to the said client-side user, to update the said synthesized audio-visual feed for the said client-side user, in response to at least one of the virtual location of the said client-side user changing or the orientation information for the said client-side user changing.

12. The system of claim 11 , wherein the one or more processors is/are configured to change the synthesized audio-visual feed rendered for the said client-side user in response to the said client side user's virtual location within the 3D model of the host site changing, such that when the said client-side user moves from a first virtual location within the 3D model of the host site to a second virtual location within the 3D model of the host site, the synthesized audio-visual feed changes from appearing as if it were generated by a virtual camera located at the first virtual location to appearing as if it were generated by a virtual camera located at the second virtual location.

13. The system of claim 11 , wherein the one or more processors is/are configured to:

modify the synthesized audio-visual feed rendered for the said client-side user in near real-time, based on the orientation information, to produce one or more modified versions thereof that include visual data for less than the full FOV visual data for the 3D model of the host site and/or emphasize the visual portion of the synthesized audio-visual feed being viewed on the said client device at any given time, to thereby reduce an amount of data that is provided to the said client device.

14. The system of claim 11 , wherein the one or more processors is/are configured to cause to be displayed on an augmented reality (AR) device being used by a said host-side user, a visual representation of at least one of the one or more client-side users at the real-world location within the host site that corresponds to their respective virtual location within the 3D model of the host site.

15. The system of claim 14 , wherein the one or more processors is/are configured to change the visual representation of the said client-side user displayed on the AR device, when the virtual location of the said client-side user changes, such that when the said client-side user moves from a first virtual location within the 3D model of the host site to a second virtual location within the 3D model of the host site, the visual representation of the said client-side user changes from being at a first corresponding real-world location within the host site to being at a second corresponding real-world location within the host site.

16. The system of claim 11 , wherein the one or more processors is/are configured to cause to display on an augmented reality (AR) device being used by a said host-side user, a visual representation of the said client-side user of the one or more client-side users with their respective orientation, which provides an indication of a visual portion of the synthesized audio-visual feed that the said client-side user is viewing at any given time.

17. The system of claim 11 , wherein the one or more client-side users comprises a plurality of client-side users, and wherein the one or more processors is/are configured to render, using a separate said virtual camera rendering engine, a separate said synthesized audio-visual feed for each of at least two of the plurality of client-side users.

18. The system of claim 17 , wherein:

the synthesized audio-visual feed rendered for a first said client side user includes an avatar of a second said client side user at their respective virtual location; and

the synthesized audio-visual feed rendered for the second said client side user includes a real-world representation of the first said client side user at their respective virtual location.

19. The system of claim 11 , wherein the one or more processors is/are configured to:

for the said synthesized audio-visual feed, modify the said synthesized audio-visual feed in near real-time, based on the orientation information received for the said client-side user, to produce one or more modified versions thereof that include visual data for less than the full FOV generated by the said virtual camera rendering engine allocated to the said-client-side user and/or emphasize the visual portion of the synthesized audio-visual feed being viewed at any given time; and

provide to the said client device, of the one or more client devices, a said modified version of the host site audio-visual feed resulting from the modifying, to thereby enable consumption thereof by the said client-side user that is using the said client device.

20. One or more processor readable storage devices having instructions encoded thereon which when executed cause one or more processors to perform a method for enabling a teleconference among one or more host-side users located at a host site and one or more client-side users that are remotely located relative to the host site, wherein a host device is located at the host site, and wherein each of the one or more client-side users uses a respective client device to participate in the teleconference, the method comprising:

receiving from the host device a host site audio-visual feed that includes audio and video of the host site, captured from the host site by the host device;

at least one of producing or updating a three-dimensional (3D) model of the host site using the host site audio-visual feed;

assigning or otherwise obtaining for each client-side user, of the one or more client-side users, a virtual location within the 3D model of the host site and thereby mapping each client-side user to a corresponding real-world location at the host site;

allocating a respective virtual camera rendering engine to each of the one or more client-side users;

using a said virtual camera rendering engine, allocated to a said client-side user, to render a synthesized audio-visual feed that appears as if it were generated by a virtual camera located at the virtual location of the said client-side user within the 3D model of the host site;

providing the synthesized audio-visual feed, or a modified version thereof, to the client device of the said client-side user to thereby enable consumption thereof by the said client-side user;

receiving, from the client device of the said client-side user, orientation information that includes an indication of a visual portion of the synthesized audio-visual feed, or the modified version thereof, that the said client-side user is viewing at any given time, which corresponds to less than a full field of view (FOV) provided by the synthesized audio-visual feed; and

using the said virtual camera rendering engine, allocated to the said client-side user, to update the said synthesized audio-visual feed for the said client-side user, in response to at least one of the virtual location of the said client-side user changing or the orientation information for the said client-side user changing.

Assignments (4)
SECURITY INTEREST Recorded May 19, 2022
From: AVATOUR TECHNOLOGIES INC.
To: NOMURA STRATEGIC VENTURES FUND 1, LP
Reel/Frame 059954/0524 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: COPLEY, DEVON; BALASUBRAMANIAN, PRASAD
To: IMEVE INC.
Reel/Frame 058686/0110 →
CHANGE OF NAME Recorded Jan 19, 2022
From: IMEVE INC.
To: AVATOUR, INC.
Reel/Frame 058686/0114 →
MERGER Recorded Jan 19, 2022
From: AVATOUR, INC.
To: AVATOUR TECHNOLOGIES, INC.
Reel/Frame 058686/0121 →
Continuity (3)
Continuation 16840780 · Apr 6, 2020
Provisional Application 62830647 · Apr 8, 2019
Related Publication 20220166807A1 · May 26, 2022