IP Library Granted Patent US 12,732,590
Granted Patent B2
US 12,732,590 · App. 18/252,327 · Granted Sep 8, 2026

Methods and apparatus to provide remote telepresence communication

Inventors: Ronald Azuma (San Jose, CA); Joshua Ratcliff (San Jose, CA); Santiago Alfaro (Campbell, CA)
Assignee: Intel Corporation
H04N7/157G06T11/60G06T15/50G06T2215/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,732,590
App. No.
18/252,327
Granted
Sep 8, 2026
Kind
B2
Abstract

Methods, apparatus, systems, and articles of manufacture are disclosed to provide remote telepresence communication. At least one non-transitory machine-readable medium comprises instructions that, when executed, cause a processor to identify features from a plurality of images, the plurality of images representing a first user and a second user, create a first representation of the first user, a second representation of the second user, the representations created using the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer, and construct a first image, the first image including the second model at a specified location within a shared environment, the first image to be presented on a first display.

Claims (56)

1 . An apparatus to provide remote telepresence communication, the apparatus comprising:

interface circuitry;

computer readable instructions; and

processor circuitry to be programmed based on the computer readable instructions to:

identify features from a plurality of images, the plurality of images representing a first user and a second user;

combine a first map set with a second map set, the first map set and the second map set corresponding to a first camera within a camera array, the first camera to capture one or more of the plurality of images;

access a rotated light map of a shared environment, the rotation based on an orientation of the first camera within the camera array;

relight the one or more of the plurality of images based on the features, the rotated light map and the combined map set;

create a first representation of the first user and a second representation of the second user, the representations based on the relit one or more of the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer; and

construct a first image, the first image to include the second representation at a specified location within a shared environment, the first image to be presented on a first display.

2 . The apparatus of claim 1 , wherein the processor circuitry is to construct a second image, the second image including the first representation at a specified location within the shared environment, the second image to be presented on a second display different from the first display.

3 . The apparatus of claim 1 , wherein:

the plurality of images is a first plurality of images; and

the processor circuitry is to:

update the representations based on a second plurality of images; and

construct a second image based on the updated second representation, the first image and second image to represent a first frame and second frame of a video.

4 . The apparatus of claim 1 , wherein:

the plurality of images represents a third user; and

the processor circuitry is to create a third representation of the third user, the third representations based on the plurality of images, the third representations representing the third user at a specified distance and a specified perspectives from a viewer, the first image including the second representation and the third representation at specified locations within the shared environment.

5 . The apparatus of claim 1 , wherein the features include depth maps, foreground extraction, and temporal flow maps.

6 . The apparatus of claim 5 , wherein the first map set includes a first normal map and a first albedo map, the first normal map and a first albedo map generated by a relighting pipeline, wherein the second map set includes a second normal map and a second albedo map, the second normal map and the second albedo map generated based on the temporal flow maps.

7 . The apparatus of claim 1 , wherein the plurality of images represent video feeds from a first camera array associated with the first user and a second camera array associated with the second user.

8 . The apparatus of claim 7 , wherein to create the first representation at specified distances, the processor circuitry is to:

generate a user representation, the user representation to include a three dimensional model or a two dimensional image;

create an image of the first user at a first distance from the first camera array, the first distance larger than a physical distance between the first user and the first camera array; and

reproject the image of the first user at the first distance onto the user representation, the reprojection to remove distortion caused by the physical distance between the first user and the first camera array.

9 . The apparatus of claim 8 , wherein to create the first representation at specified perspectives, the processor circuitry is to map coordinates of a landmark in an image representing the first user to coordinates of the landmark on a user representation, the image in the plurality of images, the mapping based on a determination that the landmark passes a visibility test.

10 . The apparatus of claim 1 , wherein the first display is a laptop screen, one or more computer monitors, or an augmented reality headset.

11 . The apparatus of claim 1 , wherein the processor circuitry includes one or more of:

at least one of a central processor unit, a graphics processor unit, or a digital signal processor, the at least one of the central processor unit, the graphics processor unit, or the digital signal processor having control circuitry to control data movement within the processor circuitry, arithmetic and logic circuitry to perform one or more first operations corresponding to machine-readable data, and one or more registers to store a result of the one or more first operations, the machine-readable data in the apparatus;

a Field Programmable Gate Array (FPGA), the FPGA including logic gate circuitry, a plurality of configurable interconnections, and storage circuitry, the logic gate circuitry and the plurality of the configurable interconnections to perform one or more second operations, the storage circuitry to store a result of the one or more second operations; or

Application Specific Integrated Circuitry (ASIC) including logic gate circuitry to perform one or more third operations.

12 . At least one non-transitory machine-readable medium comprising instructions to cause at least one programmable circuit to at least:

identify features from a plurality of images, the plurality of images representing a first user and a second user;

combine a first texture map with a second texture map, the first texture map and the second texture map corresponding to a first camera within a camera array, the first camera to capture one or more of the plurality of images;

access a rotated light map of a shared environment, the rotation based on an orientation of the first camera within the camera array;

relight the one or more of the plurality of images based on the features, the rotated light map and the combined texture map;

create a first representation of the first user and a second representation of the second user, the representations based on the relit one or more of the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer; and

construct a first image, the first image including the second representation at a specified location within the shared environment, the first image to be presented on a first display.

13 . The at least one non-transitory machine-readable medium of claim 12 , wherein the instructions are to cause one or more of the at least one programmable circuit to construct a second image, the second image including the first representation at a specified location within the shared environment, the second image to be presented on a second display different from the first display.

14 . The at least one non-transitory machine-readable medium of claim 12 , wherein the plurality of images is a first plurality of images, wherein the instructions are to cause one or more of the at least one programmable circuit to:

update the representations based on a second plurality of images; and

construct a second image based on the updated second representation, the first image and second image to represent a first frame and second frame of a video.

15 . The at least one non-transitory machine-readable medium of claim 12 , wherein:

the plurality of images further represents a third user;

the instructions are to cause one or more of the at least one programmable circuit to create a third representation of the third user, the third representations based on the plurality of images, the third representations representing the third user at a specified distance and a specified perspectives from a viewer; and

the first image further includes the second representation and the third representation at specified locations within the shared environment.

16 . The at least one non-transitory machine-readable medium of claim 12 , wherein the features include depth maps, foreground extraction, and temporal flow maps.

17 . A method to provide remote telepresence communication, the method comprising:

identifying features from a plurality of images, the plurality of images representing a first user and a second user;

combining a first texture map with a second texture map, the first texture map and the second texture map corresponding to a first camera within a camera array, the first camera to capture one or more of the plurality of images;

accessing a rotated light map of a shared environment, the rotation based on an orientation of the first camera within the camera array;

relighting the one or more of the plurality of images based on the features, the rotated light map and the combined texture map;

creating a first representation of the first user and a second representation of the second user, the representations based on the relit one or more of the plurality of images, the representations representing the respective users at specified distances and specified perspectives from a viewer; and

constructing a first image, the first image to include the second representation at a specified location within the shared environment, the first image to be presented on a first display.

18 . The method of claim 17 , including creating a second image, the second image including the first representation at a specified location within the shared environment, the second image to be presented on a second display different from the first display.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 11, 2026
From: INTEL CORPORATION
To: INTEL PRODUCTS IP LLC
Reel/Frame 075991/0662 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: RATCLIFF, JOSHUA; AZUMA, RONALD; ALFARO, SANTIAGO
To: INTEL CORPORATION
Reel/Frame 064076/0703 →
Continuity (2)
Provisional Application 63119438 · Nov 30, 2020
Related Publication 20240015263A1 · Jan 11, 2024
References Cited (32)
US 10565734B2 · Bevensee · 2020 [cited by examiner]
US 20140313277A1 · Yarosh · 2014 [cited by examiner]
US 20160205352A1 · Lee · 2016 [cited by applicant]
US 20180374242A1 · Li et al. · 2018 [cited by applicant]
US 20200117270A1 · Gibson · 2020 [cited by examiner]
US 20200349751A1 · Bentovim · 2020 [cited by examiner]
US 20210082185A1 · Ziegler · 2021 [cited by examiner]
US 20220020128A1 · Roy · 2022 [cited by examiner]
US 20220292762A1 · Boubekeur · 2022 [cited by examiner]
US 20220368969A1 · Sheppard · 2022 [cited by examiner]
US 20240015263A1 · Azuma · 2024 [cited by examiner]
KR 20160086226A · 2016 [cited by applicant]
KR 20170059310A · 2017 [cited by applicant]
KR 20170093451A · 2017 [cited by applicant]
KR 20180062045A · 2018 [cited by applicant]
International Searching Authority, “International Search Report,” issued in connection with International Patent Application No. PCT/US2021/060353, dated Mar. 17, 2022, 3 Pages. [cited by applicant]
International Searching Authority, “Written Opinion of the International Searching Authority,” issued in connection with International Patent Application No. PCT/US2021/060353, dated Mar. 17, 2022, 5 Pages. [cited by applicant]
International Searching Authority, “International Preliminary Report on Patentability,” issued in connection with International Patent Application No. PCT/US2021/060353, dated Jun. 15, 2023, 7 pages. [cited by applicant]
Raskar, et al., “The Office of the Future: A Unified Approach to Image-Based Modeling and Spatially Immersive Displays”, SIGGRAPH 98, Computer Graphics Proceedings, Annual Conference Series, 1998, published Jul. 19-24, … [cited by applicant]
Jones et al., “Achieving eye contact in a one-to-many 3D video teleconferencing system”, ACM Trans. Graph. vol. 28, Issue 3, Article No. 64, published Jul. 27, 2009, 7 pages. [cited by applicant]
Fried et al., “Perspective-aware manipulation of portrait photos”, ACM Transactions on Graphics (TOG), vol. 35, Issue 4, Article No. 128, published Jul. 11, 2016, 5 pages. [cited by applicant]
Orts-Escolano et al., “Holoportation: Virtual 3D Teleportation in Real-Time”, Microsoft Research, UIST 2016, published Oct. 16-19, 2016, 14 pages. [cited by applicant]
Wei et al., “VR Facial Animation via Multiview Image Translation”, Facebook Reality Labs, ACM Trans. Graph., vol. 38, No. 4, Article 67, published Jul. 2019, 16 pages. [cited by applicant]
Shih et al., “Distortion-Free Wide-Angle Portraits on Camera Phones”, Google, ACM Trans. Graph., vol. 38, No. 4, Article 61, published Jul. 2019, 12 pages. [cited by applicant]
Zhang et al., “Portrait Shadow Manipulation”, ACM Trans. Graph., vol. 39, No. 4, Article 78, published Jul. 2020, 14 pages. [cited by applicant]
Ray, S., “Video fatigue and a late-night host with no audience inspire a new way to help people feel together, remotely”, Microsoft, Work & Life, published Jul. 8, 2020, 12 pages. [cited by applicant]
Bailenson, Jeremy N., “Nonverbal Overload: A Theoretical Argument for the Causes of Zoom Fatigue”, American Psychological Association, Technology, Mind, and Behavior, published 2021, 6 pages. [cited by applicant]
Bavor, C., “Project Starline: Feel like you're there, together”, Google Research, Retrieved from: https://blog.google/technology/research/project-starline/, published May 18, 2021, 3 pages. [cited by applicant]
Lombardi et al., “Mixture of Volumetric Primitives for Efficient Neural Rendering”, Facebook Reality Labs, ACM Trans. Graph., vol. 40, No. 4, Article 59, published Aug. 2021, 13 pages. [cited by applicant]
Pandey et al., “Total Relighting: Learning to Relight Portraits for Background Replacement”, Google Research, ACM Trans. Graph., vol. 40, No. 4, Article 43, published Aug. 2021, 21 pages. [cited by applicant]
Cutter, C., “Remote work may now last for two years, worrying some bosses,” The Wall Street Journal, published Aug. 22, 2021, 6 pages. [cited by applicant]
Meta Quest, “Introducing Horizon Workrooms: Remote Collaboration Reimagined”, Meta, Retrieved from: https://www.meta.com/en-gb/blog/workrooms/, last updated Dec. 6, 2023, 10 pages. [cited by applicant]