Multimedia content encoding
Methods and systems configured to embed data representing content output by at least one device for downstream content recognition or other downstream processes. For example, a device operates one or more media data embedding components configured to embed information regarding the output media, such as audio or image content. The embedding component is trained to embed information for a subset of known downstream processes with some known processes deliberately held out from the training. Among other benefits, this approach can help reduce over customization of the embedding component and allow more information to preserved by the component for purposes of downstream operations that may yet be configured.
1 . A computer-implemented method, comprising:
determining a preliminary embedding component configured to process content data to determine embedded content data, the preliminary embedding component corresponding to first parameter data;
determining a plurality of components configured to operate on embedded content data, the plurality of components comprising:
a first component configured to perform a first classification task using the embedded content data,
a second component configured to perform a second classification task using the embedded content data, and
a third component configured to perform a third classification task using the embedded content data;
determining training content data corresponding to training content;
processing a first portion of the training content data using the preliminary embedding component to determine training embedding content data;
processing the training embedding content data using the first component to determine first training classification data;
processing the training embedding content data using the second component to determine second training classification data;
processing the first training classification data and the second training classification data to determine training loss data;
adjusting the preliminary embedding component using the training loss data to determine an adjusted embedding component configured to process the content data to determine the embedded content data, the adjusted embedding component corresponding to second parameter data different from the first parameter data; and
sending the adjusted embedding component to a user device.
2 . The computer-implemented method of claim 1 , further comprising, after sending the adjusted embedding component:
receiving, from the user device, first embedding content data determined by the adjusted embedding component based on first content data processed by the user device;
processing the first embedding content data using the third component to determine output data corresponding to the third classification task; and
performing an action based at least in part on the output data.
3 . The computer-implemented method of claim 1 , further comprising, after sending the adjusted embedding component:
receiving, from the user device, first embedding content data determined by the adjusted embedding component based on first content data processed by the user device;
processing the first embedding content data using a fourth component to determine output data corresponding to a fourth classification task, wherein the plurality of components does not include the fourth component; and
performing an action based at least in part on the output data.
4 . The computer-implemented method of claim 1 , further comprising:
processing a second portion of the training content data using the adjusted embedding component to determine second training embedding image data;
processing the second training embedding image data using the third component to determine third training classification data;
processing the third training classification data with respect to ground truth data to determine output data; and
processing the output data to determine the third training classification data sufficiently corresponded to the ground truth data,
wherein sending the adjusted embedding component is based at least in part on the third training classification data sufficiently corresponding to the ground truth data.
5 . A computer-implemented method, comprising:
receiving, by a user device, first data representing first content to be output;
processing, by the user device, the first data using a first trained embedding component to determine first embedding data, wherein the trained embedding component was trained by adjusting embedding component parameters based on training results relative to operation of a plurality of components different than the trained embedding component and that are configured to process embedded media data;
sending, by the user device to a first processing component to perform a first task of a first type, the first embedding data, wherein the first processing component is different than the plurality of components;
after sending the first embedding data, receiving, by the user device, second data representing second content to be output;
processing, by the user device, the second data using the first trained embedding component to determine second embedding data; and
sending, by the user device to a second processing component to perform a second task of the first type, the second embedding data, wherein the second processing component is configured differently from the first processing component.
6 . The computer-implemented method of claim 5 , further comprising, prior to receiving the first data by the user device:
determining a preliminary embedding component configured to process media data to determine embedded media data, the preliminary embedding component corresponding to first parameter data;
processing a first portion of training media data using the preliminary embedding component to determine training embedding data;
processing the training embedding data using a third processing component of the plurality of components to determine first training classification data;
processing the training embedding data using a third processing component of the plurality of components to determine second training classification data;
processing the first training classification data and the second training classification data to determine training loss data; and
generating an adjusted preliminary embedding component using the training loss data to determine the trained embedding component, the trained embedding component corresponding to second parameter data different from the first parameter data.
7 . The computer-implemented method of claim 6 , further comprising, prior to receiving the first data by the user device:
processing a second portion of the training media data using the trained embedding component to determine second training embedding image data;
processing the second training embedding image data using a fourth processing component to determine third training classification data, wherein the fourth processing component is different from the plurality of components;
processing the third training classification data with respect to ground truth data to determine output data;
processing the output data to determine the third training classification data sufficiently corresponded to the ground truth data; and
sending the trained embedding component to the user device.
8 . The computer-implemented method of claim 7 , wherein the fourth processing component comprises the first processing component.
9 . The computer-implemented method of claim 5 , wherein:
the first data comprises audio data; and
the method further comprises processing the first embedding data by the first processing component to perform a first audio processing task corresponding to the first task.
10 . The computer-implemented method of claim 5 , wherein:
the first data comprises image data; and
the method further comprises processing the first embedding data by the first processing component to perform a first image processing task corresponding to the first task.
11 . The computer-implemented method of claim 5 , further comprising:
processing the first data by an output component of the user device to cause presentation of the content.
12 . A system, comprising:
at least one processor; and
at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:
receive, by a user device, first data representing content to be output;
determine a preliminary embedding component configured to process media data to determine embedded media data;
process a first portion of training media data using the preliminary embedding component to determine training embedding data;
process the training embedding data using a first processing component of a plurality of components to determine first training classification data;
process the training embedding data using a third processing component of the plurality of components to determine second training classification data;
process the first training classification data and the second training classification data to determine training loss data;
generate an adjusted preliminary embedding component using the training loss data to determine a trained embedding component;
process, by the user device, the first data using the trained embedding component to determine first embedding data; and
send, by the user device to a second processing component to perform a first task, the first embedding data, wherein the second processing component is different than the plurality of components.
13 . The system of claim 12 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to, prior to receipt of the first data by the user device:
wherein the preliminary embedding component corresponds to first parameter data; and
wherein the trained embedding component corresponds to second parameter data different from the first parameter data.
14 . The system of claim 12 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to, prior to receipt of the first data by the user device:
process a second portion of the training media data using the trained embedding component to determine second training embedding image data;
process the second training embedding image data using a fourth processing component to determine third training classification data, wherein the fourth processing component is different from the plurality of components;
process the third training classification data with respect to ground truth data to determine output data;
process the output data to determine the third training classification data sufficiently corresponded to the ground truth data; and
send the trained embedding component to the user device.
15 . The system of claim 14 , wherein the fourth processing component comprises the second processing component.
16 . The system of claim 12 , wherein the first embedding data represents additional media information not used by the plurality of components.
17 . The system of claim 12 , wherein:
the first data comprises audio data; and
the plurality of components comprises a first component configured to perform a first audio processing task and a second component configured to perform a second audio processing task.
18 . The system of claim 12 , wherein:
the first data comprises image data; and
the plurality of components comprises a first component configured to perform a first image processing task and a second component configured to perform a second image processing task.
19 . The system of claim 12 , wherein the at least one memory further includes instructions that, when executed by the at least one processor, further cause the system to:
process the first data by an output component of the user device to cause presentation of the content.