IP Library › Granted Patent US 12,230,252
Granted Patent B2
US 12,230,252 · App. 17/282,135 · Granted Feb 18, 2025

Generation of interactive audio tracks from visual content

Inventors: Matthew Sharifi (Mountain View, CA); Victor Carbune (Mountain View, CA)
Assignee: GOOGLE LLC
G10L15/083G06F3/167G06V20/64G10L15/063G10L15/1822G10L15/22G10L15/26G10L2015/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,230,252
App. No.
17/282,135
Granted
Feb 18, 2025
Kind
B2
Abstract

Generating audio tracks is provided. The system selects a digital component object having a visual output format. The system determines to convert the digital component object into an audio output format. The system generates text for the digital component object. The system selects, based on context of the digital component object, a digital voice to render the text. The system constructs a baseline audio track of the digital component object with the text rendered by the digital voice. The system generates, based on the digital component object, non-spoken audio cues. The system combines the non-spoken audio cues with the baseline audio form of the digital component object to generate an audio track of the digital component object. The system provides the audio track of the digital component object to the computing device for output via a speaker of the computing device.

Claims (84)

1. A system to transition between different modalities, comprising:

a data processing system comprising one or more processors to:

receive, via a network, data packets comprising an input audio signal detected by a microphone of a computing device remote from the data processing system;

parse the input audio signal to identify a request;

select, based on the request, a digital component object having a visual output format, the digital component object associated with metadata;

determine, based on a type of the computing device, to convert the digital component object into an audio output format;

generate, responsive to the determination to convert the digital component object into the audio output format, text for the digital component object;

select, based on context of the digital component object, a digital voice to render the text;

construct a baseline audio track of the digital component object with the text rendered by the digital voice;

generate, based on the digital component object, non-spoken audio cues;

combine the non-spoken audio cues with the baseline audio form of the digital component object to generate an audio track of the digital component object; and

provide, responsive to the request from the computing device, the audio track of the digital component object to the computing device for output via a speaker of the computing device.

2. The system of claim 1 , comprising:

the data processing system to determine to convert the digital component object into the audio output format based on the type of the computing device comprising a smart speaker.

3. The system of claim 1 , comprising:

the data processing system to determine to convert the digital component object into the audio output format based on the type of the computing device comprising a digital assistant.

4. The system of claim 1 , comprising:

the data processing system to select, responsive to the request, the digital component object based on content selection criteria input into a real-time content selection process, the digital component object selected from a plurality of digital component objects provided by a plurality of third-party content providers.

5. The system of claim 1 , comprising:

the data processing system to select the digital component object based on keywords associated with content rendered by the computing device prior to the request, the digital component object selected from a plurality of digital component objects provided by a plurality of third-party content providers.

6. The system of claim 1 , comprising:

the data processing system to generate, via a natural language generation model, the text for the digital component object based on the metadata of the digital component object.

7. The system of claim 1 , comprising:

the data processing system to select, via a voice model, the digital voice based on context of the digital component object, the voice model trained by a machine learning technique with a historical data set comprising audio and visual media content.

8. The system of claim 1 , comprising the data processing system to:

input the context of the digital component object into a voice model to generate a voice characteristics vector, the voice model trained by a machine learning engine with a historical data set comprising audio and visual media content; and

select the digital voice from a plurality of digital voices based on the voice characteristics vector.

9. The system of claim 1 , comprising:

the data processing system to determine, based on the metadata, to add a trigger word to the audio track, wherein detection of the trigger word in a second input audio signal causes the data processing system or the computing device to perform a digital action corresponding to the trigger word.

10. The system of claim 1 , comprising the data processing system to:

determine a category of the digital component object;

retrieve, from a database, a plurality of trigger words corresponding to a plurality of digital actions associated with the category;

rank, using a digital action model trained based on historical performance of trigger keywords, the plurality of trigger words based on the context of the digital component object and the type of the computing device; and

select a highest ranking trigger keyword to add to the audio track.

11. The system of claim 1 , comprising the data processing system to:

perform image recognition on the digital component object to identify a visual object in the digital component object; and

select, from a plurality of non-spoken audio cues stored in a database, the non-spoken audio cue corresponding to the visual object.

12. The system of claim 1 , comprising the data processing system to:

identify a plurality of visual objects in the digital component object via an image recognition technique;

select, based on the metadata and the plurality of visual objects, a plurality of non-spoken audio cues;

determine a matching score for each of the visual objects that indicates a level of match between each of the visual objects and the metadata;

rank the plurality of non-spoken audio cues based on the matching score;

determine a level of audio interference between each of the plurality of non-spoken audio cues and the digital voice selected based on the context to render the text; and

select, based on a highest rank, the non-spoken audio cue from the plurality of non-spoken audio cues associated with the level of audio interference less than a threshold.

13. The system of claim 1 , comprising:

identify, based on an insertion model trained using historical performance data, an insertion point for the audio track in a digital media stream output by the computing device; and

provide instruction to the computing device to cause the computing device to render the audio track at the insertion point in the digital media stream.

14. A method to transition between different modalities, comprising:

receiving, by one or more processors of a data processing system via a network, data packets comprising an input audio signal detected by a microphone of a computing device remote from the data processing system;

parsing, by the data processing system, the input audio signal to identify a request;

selecting, by the data processing system based on the request, a digital component object having a visual output format, the digital component object associated with metadata;

determining, by the data processing system based on a type of the computing device, to convert the digital component object into an audio output format;

generating, by the data processing system responsive to the determination to convert the digital component object into the audio output format, text for the digital component object;

selecting, by the data processing system based on context of the digital component object, a digital voice to render the text;

constructing, by the data processing system, a baseline audio track of the digital component object with the text rendered by the digital voice;

generating, by the data processing system based on the digital component object, non-spoken audio cues;

combining, by the data processing system, the non-spoken audio cues with the baseline audio form of the digital component object to generate an audio track of the digital component object; and

providing, by the data processing system responsive to the request from the computing device, the audio track of the digital component object to the computing device for output via a speaker of the computing device.

15. The method of claim 14 , comprising:

determining, by the data processing system, to convert the digital component object into the audio output format based on the type of the computing device comprising a smart speaker.

16. The method of claim 14 , comprising:

selecting, by the data processing system responsive to the request, the digital component object based on content selection criteria input into a real-time content selection process, the digital component object selected from a plurality of digital component objects provided by a plurality of third-party content providers.

17. The method of claim 14 , comprising:

selecting, by the data processing system, the digital component object based on keywords associated with content rendered by the computing device prior to the request, the digital component object selected from a plurality of digital component objects provided by a plurality of third-party content providers.

18. A system to transition between different modalities, comprising:

a data processing system comprising one or more processors to:

identify keywords associated with digital streaming content rendered by a computing device;

select, based on the keywords, a digital component object having a visual output format, the digital component object associated with metadata;

determine, based on a type of the computing device, to convert the digital component object into an audio output format;

generate, responsive to the determination to convert the digital component object into the audio output format, text for the digital component object;

select, based on context of the digital component object, a digital voice to render the text;

construct a baseline audio track of the digital component object with the text rendered by the digital voice;

generate, based on the metadata of the digital component object, non-spoken audio cues;

combine the non-spoken audio cues with the baseline audio form of the digital component object to generate an audio track of the digital component object; and

provide the audio track of the digital component object to the computing device for output via a speaker of the computing device.

19. The system of claim 18 , comprising:

the data processing system to determine to convert the digital component object into the audio output format based on the type of the computing device comprising a smart speaker.

20. The system of claim 19 , comprising:

the data processing system to select the digital component object based on the keywords input into a real-time content selection process, the digital component object selected from a plurality of digital component objects provided by a plurality of third-party content providers.

21. The system of claim 1 , wherein the digital voice is selected based on the context of the digital component object, and wherein the context of the digital component object is based on a keyword, a topic, a concept, or a vertical category.

22. The system of claim 1 , wherein the digital voice is further selected based on context of the computing device, the context of the computing device including a mode of transportation, a location, a preference, performance information or other information associated with the computing device.

23. The system of claim 1 , wherein the digital voice is selected from a plurality of digital voice prints, and wherein the digital voice prints are categorized based on a gender, an accent, a phonation, a pitch, a loudness, or a speech rate.

24. The system of claim 1 , comprising the data processing system to:

select the non-spoken audio cue from a plurality of non-spoken audio cues based on a level of audio interference between the non-spoken audio cue and the digital voice selected based on the context.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2021
From: SHARIFI, MATTHEW; CARBUNE, VICTOR
To: GOOGLE LLC
Reel/Frame 055797/0251 →
Continuity (1)
Related Publication 20220157300A1 · May 19, 2022
References Cited (27)
US 10141006B1 · Burciu · 2018 [cited by examiner]
US 11722571B1 · Chenier · 2023 [cited by examiner]
US 11727918B2 · Moreno · 2023 [cited by examiner]
US 11930050B2 · Lewis et al. · 2024 [cited by applicant]
US 20090254345A1 · Fleizach · 2009 [cited by examiner]
US 20110047163A1 · Chechik · 2011 [cited by examiner]
US 20140012586A1 · Rubin · 2014 [cited by examiner]
US 20140380149A1 · Gallo · 2014 [cited by examiner]
US 20160048561A1 · Jones · 2016 [cited by examiner]
US 20170229040A1 · Joshi · 2017 [cited by examiner]
US 20180096675A1 · Nygaard et al. · 2018 [cited by applicant]
US 20180157745A1 · Williams et al. · 2018 [cited by applicant]
US 20200357386A1 · Gao · 2020 [cited by examiner]
US 20200365148A1 · Ji · 2020 [cited by examiner]
US 20210056961A1 · Ding · 2021 [cited by examiner]
US 20210151038A1 · Manjunath · 2021 [cited by examiner]
CN 1549999 · 2004 [cited by applicant]
CN 107464561 · 2017 [cited by applicant]
CN 107516511 · 2017 [cited by applicant]
CN 109392309 · 2019 [cited by applicant]
CN 110753927 · 2020 [cited by applicant]
GB 201714754 · 2018 [cited by applicant]
JP 2015135494 · 2015 [cited by applicant]
WO WO2019216969A1 · 2019 [cited by examiner]
International Search Report and Written Opinion for PCT Appln. Ser. No. PCT/US2020/036749 dated Feb. 23, 2021 (26 pages). [cited by applicant]
International Preliminary Report on Patentability for Application No. PCT/US2020/036749, mailed Dec. 22, 2022, 19 pages. [cited by applicant]
Chinese Search Report for Corresponding Application No. 202080005699.4, dated Aug. 15, 2023. [cited by applicant]