Synthesizing video from audio using one or more neural networks
Apparatuses, systems, and techniques are presented to reduce an amount of data to be transmitted for media content. In at least one embodiment, one or more neural networks are used to generate video and audio information corresponding to one or more people based, at least in part, on at least one image and voice information corresponding to the one or more people.
1 . One or more processors, comprising:
circuitry to:
receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;
receive a first image of the one or more participants;
receive an indication of a change to one or more characteristics of the one or more participants in the first image;
receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;
generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and
cause the one or more video frames representing the one or more participants of the live transmission to be presented.
2 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.
3 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.
4 . The one or more processors of claim 1 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.
5 . The one or more processors of claim 1 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.
6 . The one or more processors of claim 1 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission.
7 . A system comprising:
one or more processors to:
receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;
receive a first image of the one or more participants;
receive an indication of a change to one or more characteristics of the one or more participants in the first image;
receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;
generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and
cause the one or more video frames representing the one or more participants of the live transmission to be presented.
8 . The system of claim 7 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.
9 . The system of claim 7 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.
10 . The system of claim 7 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.
11 . The system of claim 7 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.
12 . The system of claim 7 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission of speech data.
13 . A method comprising:
receiving one or more transmissions of speech data representing speech of one or more participants of a live transmission;
receiving a first image of the one or more participants;
receiving an indication of a change to one or more characteristics of the one or more participants in the first image;
receiving a second image of the one or more participants, that has been transmitted in response to the indication of the change;
generating, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and
causing the one or more video frames representing the one or more participants of the live transmission to be presented.
14 . The method of claim 13 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.
15 . The method of claim 13 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.
16 . The method of claim 13 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.
17 . The method of claim 13 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of the speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.
18 . The method of claim 13 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission of the speech data.
19 . A non-transitory machine-readable medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:
receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;
receive a first image of the one or more participants;
receive an indication of a change to one or more characteristics of the one or more participants in the first image;
receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;
generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and
cause the one or more video frames representing the one or more participants of the live transmission to be presented.
20 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.
21 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.
22 . The non-transitory machine-readable medium of claim 19 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.
23 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.
24 . The non-transitory machine-readable medium of claim 19 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the live transmission.
25 . A video communication system, comprising:
one or more processors to receive one or more transmissions of speech data representing speech of one or more participants of a live transmission;
receive a first image of the one or more participants;
receive an indication of a change to one or more characteristics of the one or more participants since the first image;
receive a second image of the one or more participants, that has been transmitted in response to the indication of the change;
generate, using one or more neural networks, one or more video frames representing the one or more participants, wherein the one or more video frames are generated based on the first image, the second image, and the speech data; and
cause the one or more video frames representing the one or more participants of the live transmission to be presented; and
store one or more network parameters for the one or more neural networks in one or more memories.
26 . The video communication system of claim 25 , wherein the one or more neural networks are to generate the one or more video frames including one or more representations of the one or more participants uttering speech represented in the live transmission.
27 . The video communication system of claim 25 , wherein the one or more neural networks are to generate one or more representations of the one or more participants based, at least in part, upon image features extracted from at least one image of the one or more participants.
28 . The video communication system of claim 25 , wherein the one or more video frames representing the one or more participants are to be generated based on one or more additional images, the one or more additional images being accessed by the one or more processors in response to the indication of the change to one or more characteristics of the one or more participants.
29 . The video communication system of claim 25 , wherein the one or more neural networks are to generate the one or more video frames based, at least in part, on at least one reference image of the one or more participants captured during the live transmission, or prior to, receiving of the one or more received transmissions of speech data, the at least one reference image being provided with, or separate from, the one or more received transmissions of the speech data.
30 . The video communication system of claim 25 , wherein the one or more neural networks to generate video representative of an amount of emotion or pattern of speech data determined from the one or more received transmissions of speech data.