IP Library Granted Patent US 12664206
Granted Patent B2
US 12664206 · App. 17/616,866 · Granted Jun 23, 2026

Automatically enhancing streaming media using content transformation

Inventors: Akhilesh Shirbhate (Sunnyvale, CA); Ariyam Das (Sunnyvale, CA)
Assignee: Google LLC
G06F16/433G06F16/483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664206
App. No.
17/616,866
Granted
Jun 23, 2026
Kind
B2
Abstract

A method includes receiving media content comprising audio data for distribution through content distribution platform that requires the media content to include video content, transforming the audio data into textual content, determining, based on a search of a searchable database, that the textual content of the audio data matches characteristics of visual data in the searchable database, integrating the visual data having the matched characteristics with the media content to create an augmented content stream in response to the determination that the textual content of the audio data matches the characteristics of the visual data, and distributing the augmented content stream through the content distribution platform that requires the media content to include video content.

Claims (67)

1 . A method, comprising:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

2 . The method of claim 1 , further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the media content.

3 . The method of claim 2 , wherein the annotation specifies one or more visual data characteristics.

4 . The method of claim 2 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

5 . The method of claim 2 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

6 . The method of claim 1 , further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

7 . The method of claim 6 , further comprising:

identifying a particular temporal location within the audio data based on the first context of the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.

8 . A system comprising:

one or more processors; and

one or more memory elements including instructions that, when executed, cause the one or more processors to perform operations including:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

9 . The system of claim 8 , the operations further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.

10 . The system of claim 9 , wherein the annotation specifies one or more visual data characteristics.

11 . The system of claim 9 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

12 . The system of claim 9 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

13 . The system of claim 8 , the operations further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

14 . The system of claim 13 , the operations further comprising:

identifying a particular temporal location within the audio data based on the first context of the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.

15 . A non-transitory computer-storage medium encoded with instructions that when executed by a distributed computing system cause the distributed computing system to perform operations comprising:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

16 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.

17 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies one or more visual data characteristics.

18 . The non-transitory computer-storage medium of claim 16 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

19 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

20 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

21 . The method of claim 1 , further comprising:

emphasizing portions of the first set of textual content based on audio characteristics other than identification of the first person, wherein selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data comprises selecting the visual data based on the emphasized portions of the first set of textual content.