IP Library Granted Patent US 12,664,206
Granted Patent B2
US 12,664,206 · App. 17/616,866 · Granted Jun 23, 2026

Automatically enhancing streaming media using content transformation

Inventors: Akhilesh Shirbhate (Sunnyvale, CA); Ariyam Das (Sunnyvale, CA)
Assignee: Google LLC
G06F16/433G06F16/483
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,206
App. No.
17/616,866
Filed
Dec 6, 2021
Granted
Jun 23, 2026
Kind
B2
Art Unit
2169
USPC
707/756
Abstract

A method includes receiving media content comprising audio data for distribution through content distribution platform that requires the media content to include video content, transforming the audio data into textual content, determining, based on a search of a searchable database, that the textual content of the audio data matches characteristics of visual data in the searchable database, integrating the visual data having the matched characteristics with the media content to create an augmented content stream in response to the determination that the textual content of the audio data matches the characteristics of the visual data, and distributing the augmented content stream through the content distribution platform that requires the media content to include video content.

Claims (67)

1 . A method, comprising:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

2 . The method of claim 1 , further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the media content.

3 . The method of claim 2 , wherein the annotation specifies one or more visual data characteristics.

4 . The method of claim 2 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

5 . The method of claim 2 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

6 . The method of claim 1 , further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

7 . The method of claim 6 , further comprising:

identifying a particular temporal location within the audio data based on the first context of the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.

8 . A system comprising:

one or more processors; and

one or more memory elements including instructions that, when executed, cause the one or more processors to perform operations including:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

9 . The system of claim 8 , the operations further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.

10 . The system of claim 9 , wherein the annotation specifies one or more visual data characteristics.

11 . The system of claim 9 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

12 . The system of claim 9 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

13 . The system of claim 8 , the operations further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

14 . The system of claim 13 , the operations further comprising:

identifying a particular temporal location within the audio data based on the first context of the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying the visual data with the audio data at the particular temporal location within the audio data.

15 . A non-transitory computer-storage medium encoded with instructions that when executed by a distributed computing system cause the distributed computing system to perform operations comprising:

receiving audio data for distribution through content distribution platform, wherein the audio data includes first speech of a first person and second speech of a second person;

differentiating the first speech of the first person in the audio data from the second speech of the second person in the audio data;

transforming the audio data into textual content;

flagging a first set of the textual content as being the first speech of the first person in the audio data;

flagging a second set of the textual content as being the second speech of the second person in the audio data;

selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data;

integrating the visual data selected based on the first set of textual content with the audio data to create an augmented content stream that includes the visual data, audio of the first speech of the first person, and audio of the second speech of the second person; and

distributing, to a user that has requested the audio data, the augmented content stream through the content distribution platform.

16 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:

detecting an annotation located at a particular temporal location within the audio data; and

wherein integrating the visual data with the audio data to create the augmented content stream comprises overlaying, based on the annotation, the visual data with the audio data at the particular temporal location within the audio data.

17 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies one or more visual data characteristics.

18 . The non-transitory computer-storage medium of claim 16 , wherein integrating the visual data with the audio data to create the augmented content stream further comprises editing the visual data based on one or more visual data characteristics.

19 . The non-transitory computer-storage medium of claim 16 , wherein the annotation specifies that visual data cannot be overlaid with the audio data at the particular temporal location.

20 . The non-transitory computer-storage medium of claim 15 , the operations further comprising:

determining a first context of the audio data based on the textual content of the audio data;

determining a second context of the visual data based on characteristics of the visual data in the searchable database; and

determining that the textual content of the audio data matches the characteristics of the visual data in the searchable database based on the first context matching the second context.

21 . The method of claim 1 , further comprising:

emphasizing portions of the first set of textual content based on audio characteristics other than identification of the first person, wherein selecting, based on a search of a searchable database using the first set of textual content instead of the second set of textual content, visual data in the searchable database based on the first set of the textual content being flagged as being the first speech of the first person in the audio data comprises selecting the visual data based on the emphasized portions of the first set of textual content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2021
From: SHIRBHATE, AKHILESH; DAS, ARIYAM
To: GOOGLE LLC
Reel/Frame 058483/0043 →
Continuity (1)
Related Publication 20220398276A1 · Dec 15, 2022
References Cited (22)
US 10847149B1 · Mok · 2020 [cited by examiner]
US 20090204243A1 · Marwaha et al. · 2009 [cited by applicant]
US 20150339396A1 · Ayers · 2015 [cited by examiner]
US 20170075652A1 · Kikugawa · 2017 [cited by examiner]
US 20200159487A1 · Dawson · 2020 [cited by examiner]
US 20200321005A1 · Iyer et al. · 2020 [cited by applicant]
US 20200357442A1 · Denoue · 2020 [cited by examiner]
CN 105488094 · 2016 [cited by applicant]
CN 109716327 · 2019 [cited by applicant]
EP 3335096 · 2018 [cited by applicant]
GB 2486038 · 2012 [cited by applicant]
WO 2017124116 · 2017 [cited by applicant]
Haijun Xia, Jennifer Jacobs, Manesh Agrawala, “Crosscast: Adding Visuals to Audio Travel Podcasts”, Proceedings of the 33rd annual ACM Symposium on User Interface Software and Technology, Oct. 20-23, Virtual Event, USA … [cited by examiner]
International Search Report and Written Opinion in International Appln. No. PCT/US2020/065559, dated Jul. 16, 2021, 15 pages. [cited by applicant]
Office Action in European Appln. No. 20842427.5, dated Sep. 2, 2022, 6 pages. [cited by applicant]
Xia et al., “Crosscast: Adding Visuals to Audio Travel Podcasts,” Proceedings of the 33rd annual ACM Symposium on User Interface Software and Technology, Oct. 20, 2020, 735-746. [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2020/065559, mailed on Jun. 29, 2023, 9 pages. [cited by applicant]
Summons to Attend Oral Proceedings in European Appln. No. 20842427.5, mailed on Sep. 29, 2023, 10 pages. [cited by applicant]
Office Action in European Appln. No. 20842427.5, mailed on Apr. 3, 2024, 18 pages. [cited by applicant]
Office Action in Indian Appln. No. 202127056820, mailed on Sep. 2, 2024, 8 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 202080046312.X, mailed on Oct. 31, 2024, 12 pages (with English translation). [cited by applicant]
Office Action in Indian Appln. No. 202127056820, mailed on Oct. 31, 2025, 4 pages (with English translation). [cited by applicant]