IP Library Granted Patent US 11,734,348
Granted Patent B2
US 11,734,348 · App. 16/136,788 · Granted Aug 22, 2023

Intelligent audio composition guidance

Inventors: Craig M. Trim (Ventura, CA); Gandhi Sivakumar (Bentleigh, AU); Martin G. Keen (Cary, NC); Hernan A. Cunico (Holly Springs, NC)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F16/7834G06F16/24578G06V20/41G10L15/18G10L15/26H04N5/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,734,348
App. No.
16/136,788
Granted
Aug 22, 2023
Kind
B2
Abstract

Embodiments for implementing intelligent audio composition guidance for a video by a processor. One or more acoustic characteristics used in a plurality of video segments may be identified, from a corpus, as having similar acoustic, linguistic, and visual characteristics of a selected video segment.

Claims (49)

1. A method, by a processor, for implementing intelligent audio composition guidance for a video in an Internet of Things (IoT) computing environment, comprising:

receiving, as input from a user through a user interface, a video associated with a current video project;

segmenting the video into video segments and identifying a selected video segment for which an audio composition is to be generated for;

inputting only the selected video segment of the video for which the audio composition is to be generated for into a search query;

analyzing the selected video segment to produce search query metadata;

generating, using a visual composition analysis of visual data of the selected video segment based on the search query metadata, the search query as a search for an audio analysis of audio data of alternative video segments stored in a corpus based on the visual composition analysis of the visual data;

responsive to receiving the selected video segment and the search query metadata as input into the search query, identifying, from the corpus, one or more acoustic characteristics used in the alternative video segments having similar acoustic, linguistic, and visual characteristics of the selected video segment according to a comparison of the selected video segment and the alternative video segments, wherein the similar acoustics, linguistic, and visual characteristics of the selected video segment are determined by performing an artificial intelligence (AI) analysis to derive a topic, tone, and emotional components related to a context of the selected video segment, including indications of characteristics associated with those of the alternative video segments used during a particular historical time period, used to match a similar context of the alternative video segments to identify the one or more acoustic characteristics related thereto notwithstanding whether the similar acoustic, linguistic, and visual characteristics of the selected video segment and the alternative video segments are identical;

assigning a correlation score to each of the alternative video segments according to a degree of similarity to the acoustic, linguistic, and visual characteristics of the selected video segment;

outputting an indication of the identified one or more acoustic characteristics determined from the comparison via the user interface selectively used in adding the audio composition to the selected video segment, wherein the indication is inclusive of a plurality of alternative audio compositions respectively used in the alternative video segments to which the identified one or more acoustic characteristics correspond;

in conjunction with outputting the indication, displaying those of the alternative video segments having the correlation score over a predefined threshold on an interactive graphical user interface (GUI) and listed with the correlation score, wherein the user selects through the plurality of the alternative video segments listed on the GUI to selectively view only those specific portions of the alternative video segments respectively associated with those of the plurality of alternative audio compositions identified by the search query; and

generating the audio composition for the selected video segment using the indicated identified one or more acoustic characteristics, wherein generating the audio composition includes automatically associating the identified one or more acoustic characteristics, output as a result of the search query, with the selected video segment such that at least one of the plurality of alternative audio compositions used in at least one of the alternative video segments selected by the user from the list on the GUI is automatically applied to the selected video segment of the current video project.

2. The method of claim 1 , further including:

analyzing each of the alternative video segments; and

deriving acoustic, linguistic, and visual characteristics from the alternative video segments.

3. The method of claim 1 , further including creating the corpus describing the acoustic, linguistic, and visual characteristics derived from each the alternative video segments.

4. The method of claim 1 , further including aggregating the one or more acoustic characteristics of those of the alternative video segments having a correlation score equal to or greater than a defined threshold.

5. A system for implementing intelligent audio composition guidance for a video, comprising:

one or more computers with executable instructions that when executed cause the system to:

receive, as input from a user through a user interface, a video associated with a current video project;

segment the video into video segments and identifying a selected video segment for which an audio composition is to be generated for;

input only the selected video segment of the video for which the audio composition is to be generated for into a search query;

analyze the selected video segment to produce search query metadata;

generate, using a visual composition analysis of visual data of the selected video segment based on the search query metadata, the search query as a search for an audio analysis of audio data of alternative video segments stored in a corpus based on the visual composition analysis of the visual data;

responsive to receiving the selected video segment and the search query metadata as input into the search query, identify, from the corpus, one or more acoustic characteristics used in the alternative video segments having similar acoustic, linguistic, and visual characteristics of the selected video segment according to a comparison of the selected video segment and the alternative video segments, wherein the similar acoustics, linguistic, and visual characteristics of the selected video segment are determined by performing an artificial intelligence (AI) analysis to derive a topic, tone, and emotional components related to a context of the selected video segment, including indications of characteristics associated with those of the alternative video segments used during a particular historical time period, used to match a similar context of the alternative video segments to identify the one or more acoustic characteristics related thereto notwithstanding whether the similar acoustic, linguistic, and visual characteristics of the selected video segment and the alternative video segments are identical;

assign a correlation score to each of the alternative video segments according to a degree of similarity to the acoustic, linguistic, and visual characteristics of the selected video segment;

output an indication of the identified one or more acoustic characteristics determined from the comparison via the user interface selectively used in adding the audio composition to the selected video segment, wherein the indication is inclusive of a plurality of alternative audio compositions respectively used in the alternative video segments to which the identified one or more acoustic characteristics correspond;

in conjunction with outputting the indication, display those of the alternative video segments having the correlation score over a predefined threshold on an interactive graphical user interface (GUI) and listed with the correlation score, wherein the user selects through the plurality of the alternative video segments listed on the GUI to selectively view only those specific portions of the alternative video segments respectively associated with those of the plurality of alternative audio compositions identified by the search query; and

generate the audio composition for the selected video segment using the indicated identified one or more acoustic characteristics, wherein generating the audio composition includes automatically associating the identified one or more acoustic characteristics, output as a result of the search query, with the selected video segment such that at least one of the plurality of alternative audio compositions used in at least one of the alternative video segments selected by the user from the list on the GUI is automatically applied to the selected video segment of the current video project.

6. The system of claim 5 , wherein the executable instructions:

analyze each of the alternative video segments; and

derive acoustic, linguistic, and visual characteristics from the alternative video segments.

7. The system of claim 5 , wherein the executable instructions create the corpus describing the acoustic, linguistic, and visual characteristics derived from each the alternative video segments.

8. The system of claim 5 , wherein the executable instructions aggregate the one or more acoustic characteristics of those of the alternative video segments having a correlation score equal to or greater than a defined threshold.

9. A computer program product for, by a processor, implementing intelligent audio composition guidance for a video, the computer program product comprising a non-transitory computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

an executable portion that receives, as input from a user through a user interface, a video associated with a current video project;

an executable portion that segments the video into video segments and identifying a selected video segment for which an audio composition is to be generated for;

an executable portion that inputs only the selected video segment of the video for which the audio composition is to be generated for into a search query;

an executable portion that analyzes the selected video segment to produce search query metadata;

an executable portion that generates, using a visual composition analysis of visual data of the selected video segment based on the search query metadata, the search query as a search for an audio analysis of audio data of alternative video segments stored in a corpus based on the visual composition analysis of the visual data;

an executable portion that, responsive to receiving the selected video segment and the search query metadata as input into the search query, identifies, from the corpus, one or more acoustic characteristics used in the alternative video segments having similar acoustic, linguistic, and visual characteristics of the selected video segment according to a comparison of the selected video segment and the alternative video segments, wherein the similar acoustics, linguistic, and visual characteristics of the selected video segment are determined by performing an artificial intelligence (AI) analysis to derive a topic, tone, and emotional components related to a context of the selected video segment, including indications of characteristics associated with those of the alternative video segments used during a particular historical time period, used to match a similar context of the alternative video segments to identify the one or more acoustic characteristics related thereto notwithstanding whether the similar acoustic, linguistic, and visual characteristics of the selected video segment and the alternative video segments are identical;

an executable portion that assigns a correlation score to each of the alternative video segments according to a degree of similarity to the acoustic, linguistic, and visual characteristics of the selected video segment;

an executable portion that outputs an indication of the identified one or more acoustic characteristics determined from the comparison via the user interface selectively used in adding the audio composition to the selected video segment, wherein the indication is inclusive of a plurality of alternative audio compositions respectively used in the alternative video segments to which the identified one or more acoustic characteristics correspond;

an executable portion that, in conjunction with outputting the indication, displays those of the alternative video segments having the correlation score over a predefined threshold on an interactive graphical user interface (GUI) and listed with the correlation score, wherein the user selects through the plurality of the alternative video segments listed on the GUI to selectively view only those specific portions of the alternative video segments respectively associated with those of the plurality of alternative audio compositions identified by the search query; and

an executable portion that generates the audio composition for the selected video segment using the indicated identified one or more acoustic characteristics, wherein generating the audio composition includes automatically associating the identified one or more acoustic characteristics, output as a result of the search query, with the selected video segment such that at least one of the plurality of alternative audio compositions used in at least one of the alternative video segments selected by the user from the list on the GUI is automatically applied to the selected video segment of the current video project.

10. The computer program product of claim 9 , further including an executable portion that:

analyze each of the alternative video segments; and

derive acoustic, linguistic, and visual characteristics from the alternative video segments.

11. The computer program product of claim 9 , further including an executable portion that create the corpus describing the acoustic, linguistic, and visual characteristics derived from each the alternative video segments.

12. The computer program product of claim 9 , further including an executable portion that aggregates the one or more acoustic characteristics of those of the alternative video segments having a correlation score equal to or greater than a defined threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 20, 2018
From: TRIM, CRAIG M.; SIVAKUMAR, GANDHI; KEEN, MARTIN G.; CUNICO, HERNAN A.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 046928/0187 →
Continuity (1)
Related Publication 20200097502A1 · Mar 26, 2020
Cited By (1)
US 12,373,489