IP Library Granted Patent US 12700115
Granted Patent B2
US 12700115 · App. 18/214,604 · Granted Aug 4, 2026

User representation using depths relative to multiple surface points

Inventors: Brian Amberg (Zollikon, CH); John S. McCarten (Boulder, CO); Nicolas V. Scapel (London, GB); Peter Kaufmann (Zurich, CH); Sebastian Martin (Schwerzenbach, CH)
Assignee: Apple Inc.
G06T7/521G06T13/40G06V40/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700115
App. No.
18/214,604
Granted
Aug 4, 2026
Kind
B2
Abstract

Various implementations disclosed herein include devices, systems, and methods that generates values for a representation of a face of a user. For example, an example process may include obtaining sensor data (e.g., live data) of a user, wherein the sensor data is associated with a point in time, generating a set of values representing the user based on the sensor data, and providing the set of values, where a depiction of the user at the point in time is displayed based on the set of values. In some implementations, the set of values includes depth values that define three-dimensional (3D) positions of portions of the user relative to multiple 3D positions of points of a projected surface and appearance values (e.g., color, texture, opacity, etc.) that define appearances of the portions of the user.

Claims (42)

1 . A method comprising:

at a processor of a device:

obtaining sensor data acquired by sensors with respect to an appearance of portions of a user's face, wherein the sensor data comprises depth data and has been acquired by the sensors at a point in time;

generating a set of values representing the appearance of portions of the user's face based on the sensor data, wherein the set of values comprises:

depth values that define three-dimensional (3D) positions of the appearance of portions of the user's face relative to multiple points of a curved surface, wherein the depth values define a distance between a portion of the user's face and a corresponding point of the curved surface positioned along a ray normal to the curved surface at a position of the corresponding point; and

appearance values that define appearances of the portions of the user's face when the appearance of portions of the user's face is displayed; and

providing the set of values, wherein a depiction of the appearance of portions of the user's face at the point in time is displayed based on the set of values.

2 . The method of claim 1 , wherein the multiple points of the curved surface are spaced at regular intervals along vertical and horizontal lines on the curved surface.

3 . The method of claim 1 , wherein the curved surface is nonplanar.

4 . The method of claim 1 , wherein the curved surface is at least partially-cylindrical.

5 . The method of any of claim 1 , wherein the curved surface is planar.

6 . The method of claim 1 , wherein the set of values is generated based on an alignment such that a subset of the multiple points of the curved surface on a central area of the curved surface correspond to a central portion of the user's face.

7 . The method of claim 1 , wherein generating the set of values is further based on images of the user's face captured while the user is expressing a plurality of different facial expressions.

8 . The method of claim 7 , wherein:

the sensor data corresponds to only a first area of the user; and

the set of image data corresponds to a second area comprising a third area different than the first area.

9 . The method of claim 1 , further comprising:

obtaining additional sensor data of a user associated with a second period of time;

updating the set of values representing the user based on the additional sensor data for the second period of time; and

providing the updated set of values, wherein the depiction of the user is updated at the second period of time based on the updated set of values.

10 . The method of claim 1 , wherein providing the set of values comprises sending a sequence of frames of 3D video data comprising a frame comprising the set of values during a communication session with a second device, wherein the second device renders an animated depiction of the user based on the sequence of frames of 3D video data.

11 . The method of claim 1 , wherein the device comprises a first sensor and a second sensor, wherein the sensor data is obtained from at least one partial image of the user's face from the first sensor from a first viewpoint and from at least one partial image of the user's face from the second sensor from a second viewpoint that is different than the first viewpoint.

12 . The method of claim 1 , wherein the depiction of the user is displayed in real-time.

13 . The method of claim 1 , wherein generating the set of values representing the user's face is based on a machine learning model trained to produce the set of values.

14 . The method of claim 1 , wherein the appearance values comprise color values, texture values, or opacity values.

15 . The method of claim 1 , wherein the device is a head-mounted device (HMD).

16 . The method of claim 15 , wherein the HMD comprises one or more inward facing image sensors and one or more downward facing image sensors, and the sensor data is captured by the one or more inward facing sensors and the one or more downward facing image sensors.

17 . A device comprising:

a non-transitory computer-readable storage medium; and

one or more processors coupled to the non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium comprises program instructions that, when executed on the one or more processors, cause the one or more processors to perform operations comprising:

obtaining sensor data acquired by sensors with respect to an appearance of portions of a user's face, wherein the sensor data comprises depth data and has been acquired by the sensors at a point in time;

generating a set of values representing the appearance of portions of the user's face based on the sensor data, wherein the set of values comprises:

depth values that define three-dimensional (3D) positions of the appearance of portions of the user's face relative to multiple points of a curved surface, wherein the depth values define a distance between a portion of the user's face and a corresponding point of the curved surface positioned along a ray normal to the curved surface at a position of the corresponding point; and

appearance values that define appearances of the portions of the user's face when the appearance of portions of the user's face is displayed; and

providing the set of values, wherein a depiction of the appearance of portions of the user's face at the point in time is displayed based on the set of values.

18 . The device of claim 17 , wherein the multiple points of the curved surface are spaced at regular intervals along vertical and horizontal lines on the curved surface.

19 . A non-transitory computer-readable storage medium, storing program instructions executable on a device to perform operations comprising:

obtaining sensor data acquired by sensors with respect to an appearance of portions of a user's face, wherein the sensor data comprises depth data and has been acquired by the sensors at a point in time;

generating a set of values representing the appearance of portions of the user's face based on the sensor data, wherein the set of values comprises:

depth values that define three-dimensional (3D) positions of the appearance of portions of the user's face relative to multiple points of a curved surface, wherein the depth values define a distance between a portion of the user's face and a corresponding point of the curved surface positioned along a ray normal to the curved surface at a position of the corresponding point; and

appearance values that define appearances of the portions of the user's face when the appearance of portions of the user's face is displayed; and

providing the set of values, wherein a depiction of the appearance of portions of the user's face at the point in time is displayed based on the set of values.