Distributed generation of virtual content
Systems and techniques are described for establishing one or more virtual sessions between users. For instance, a first device can transmit, to a second device, a call establishment request for a virtual representation call for a virtual session and can receive, from the second device, a call acceptance indicating acceptance of the call establishment request. The first device can transmit, to the second device, first mesh information for a first virtual representation of a first user of the first device and first mesh animation parameters for the first virtual representation. The first device can receive, from the second device, second mesh information for a second virtual representation of a second user of the second device and second mesh animation parameters for the second virtual representation. The first device can generate, based on the second mesh information and the second mesh animation parameters, the second virtual representation of the second user.
1 . A method of generating virtual content at a first device in a distributed system, the method comprising:
receiving, by the first device and from a second device associated with a virtual session, input information associated with at least one of the second device or a user of the second device, wherein the input information includes pose information about a body of the user of the second device;
generating, based on the input information and the pose information, a virtual representation of the user of the second device;
generating, by the first device, a virtual scene from a perspective of a user of a third device associated with the virtual session, wherein the virtual scene includes the virtual representation of the user of the second device, and wherein the first device is remote from the second device and the third device; and
transmitting, to the third device, one or more frames depicting the virtual scene from the perspective of the user of the third device.
2 . The method of claim 1 , wherein the input information includes at least one of information representing a face of the user of the second device, information representing a body of the user of the second device, information representing one or more hands of the user of the second device, pose information of the second device, or audio associated with an environment in which the second device is located.
3 . The method of claim 2 , wherein the information representing the body of the user of the second device includes a pose of the body, and wherein the information representing the one or more hands of the user of the second device includes a respective pose of each hand of the one or more hands.
4 . The method of claim 2 , wherein generating the virtual representation comprises:
generating a virtual representation of a face of the user of the second device using the information representing the face of the user, the pose information of the second device, and pose information of the third device;
generating a virtual representation of a body of the user of the second device using the pose information of the second device and pose information of the third device; and
generating a virtual representation of hair of the user of the second device.
5 . The method of claim 4 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the body of the user.
6 . The method of claim 4 , wherein the virtual representation of the body of the user of the second device is generated further using inverse kinematics.
7 . The method of claim 4 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the one or more hands of the user.
8 . The method of claim 4 , further comprising:
combining the virtual representation of the face with the virtual representation of the body to generate a combined virtual representation; and
adding the virtual representation of the hair to the combined virtual representation.
9 . The method of claim 1 , wherein generating the virtual scene comprises:
obtaining a background representation of the virtual scene;
adjusting, based on the background representation of the virtual scene, lighting of the virtual representation of the user of the second device to generate a modified virtual representation of the user; and
combining the background representation of the virtual scene with the modified virtual representation of the user.
10 . The method of claim 1 , further comprising:
generating, based on input information from the third device, a virtual representation of the user of the third device;
generating a virtual scene including the virtual representation of the user of the third device from a perspective of the user of second device; and
transmitting, to the second device, one or more frames depicting the virtual scene from the perspective of the user of the second device.
11 . An apparatus associated with a first device for generating virtual content in a distributed system, comprising:
at least one memory; and
at least one processor coupled to at least one memory and configured to:
receive, by the apparatus and from a second device associated with a virtual session, input information associated with at least one of the second device or a user of the second device, wherein the input information includes a pose of a body of the user of the second device;
generate, based on the input information and the pose information, a virtual representation of the user of the second device;
generate a virtual scene from a perspective of a user of a third device associated with the virtual session, wherein the virtual scene includes the virtual representation of the user of the second device, and wherein the apparatus is remote from the second device and the third device; and
output, for transmission to the third device, one or more frames depicting the virtual scene from the perspective of the user of the third device.
12 . The apparatus of claim 11 , wherein the input information includes at least one of information representing a face of the user of the second device, information representing a body of the user of the second device, information representing one or more hands of the user of the second device, pose information of the second device, or audio associated with an environment in which the second device is located.
13 . The apparatus of claim 12 , wherein the information representing the body of the user of the second device includes a pose of the body, and wherein the information representing the one or more hands of the user of the second device includes a respective pose of each hand of the one or more hands.
14 . The apparatus of claim 12 , wherein the at least one processor is configured to:
generate a virtual representation of a face of the user of the second device using the information representing the face of the user, the pose information of the second device, and pose information of the third device;
generate a virtual representation of a body of the user of the second device using the pose information of the second device and pose information of the third device; and
generate a virtual representation of hair of the user of the second device.
15 . The apparatus of claim 14 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the body of the user.
16 . The apparatus of claim 14 , wherein the virtual representation of the body of the user of the second device is generated further using inverse kinematics.
17 . The apparatus of claim 14 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the one or more hands of the user.
18 . The apparatus of claim 14 , wherein the at least one processor is configured to:
combine the virtual representation of the face with the virtual representation of the body to generate a combined virtual representation; and
add the virtual representation of the hair to the combined virtual representation.
19 . The apparatus of claim 11 , wherein the at least one processor is configured to:
obtain a background representation of the virtual scene;
adjust, based on the background representation of the virtual scene, lighting of the virtual representation of the user of the second device to generate a modified virtual representation of the user; and
combine the background representation of the virtual scene with the modified virtual representation of the user.
20 . The apparatus of claim 11 , wherein the at least one processor is configured to:
generate, based on input information from the third device, a virtual representation of the user of the third device;
generate a virtual scene including the virtual representation of the user of the third device from a perspective of the user of second device; and
output, for transmission to the second device, one or more frames depicting the virtual scene from the perspective of the user of the second device.
21 . A non-transitory computer-readable medium of a first device having stored thereon instructions that, when executed by at least one processor, cause the at least one processor to:
receive, by the first device and from a second device associated with a virtual session, input information associated with at least one of the second device or a user of the second device, wherein the input information includes a pose of a body of the user of the second device;
generate, based on the input information and the pose information, a virtual representation of the user of the second device;
generate, by the first device, a virtual scene from a perspective of a user of a third device associated with the virtual session, wherein the virtual scene includes the virtual representation of the user of the second device, and wherein the first device is remote from the second device and the third device; and
output, for transmission to the third device, one or more frames depicting the virtual scene from the perspective of the user of the third device.
22 . The non-transitory computer-readable medium of claim 21 , wherein the input information includes at least one of information representing a face of the user of the second device, information representing a body of the user of the second device, information representing one or more hands of the user of the second device, pose information of the second device, or audio associated with an environment in which the second device is located.
23 . The non-transitory computer-readable medium of claim 22 , wherein the information representing the body of the user of the second device includes a pose of the body, and wherein the information representing the one or more hands of the user of the second device includes a respective pose of each hand of the one or more hands.
24 . The non-transitory computer-readable medium of claim 22 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to:
generate a virtual representation of a face of the user of the second device using the information representing the face of the user, the pose information of the second device, and pose information of the third device;
generate a virtual representation of a body of the user of the second device using the pose information of the second device and pose information of the third device; and
generate a virtual representation of hair of the user of the second device.
25 . The non-transitory computer-readable medium of claim 24 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the body of the user.
26 . The non-transitory computer-readable medium of claim 24 , wherein the virtual representation of the body of the user of the second device is generated further using inverse kinematics.
27 . The non-transitory computer-readable medium of claim 24 , wherein the virtual representation of the body of the user of the second device is generated further using the information representing the one or more hands of the user.
28 . The non-transitory computer-readable medium of claim 24 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to:
combine the virtual representation of the face with the virtual representation of the body to generate a combined virtual representation; and
add the virtual representation of the hair to the combined virtual representation.
29 . The non-transitory computer-readable medium of claim 21 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to:
obtain a background representation of the virtual scene;
adjust, based on the background representation of the virtual scene, lighting of the virtual representation of the user of the second device to generate a modified virtual representation of the user; and
combine the background representation of the virtual scene with the modified virtual representation of the user.