IP Library Granted Patent US 10,298,895
Granted Patent B1
US 10,298,895 · App. 15/941,221 · Granted May 21, 2019

Method and system for performing context-based transformation of a video

Inventors: Manjunath Ramachandra (Bangalore, IN); Sethuraman Ulaganathan (Tiruchirapalli, IN)
Assignee: Wipro Limited
H04N9/43G06K9/00523G06K9/00751G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,298,895
App. No.
15/941,221
Granted
May 21, 2019
Kind
B1
Abstract

Disclosed herein is a method and system for performing context-based transformation of a video. In an embodiment, a scene descriptor and a textual descriptor are generated for each scene corresponding to the video. Further, an audio context descriptor is generated based on semantic analysis of the textual descriptor. Subsequently, the audio context descriptor and the scene descriptor are correlated to generate a scene context descriptor for each scene. Finally, the video is translated using the scene context descriptor, thereby transforming the video based on context. In some embodiments, the method of present disclosure is capable of automatically changing one or more attributes, such as color of one or more scenes in the video, in response to change in the context of audio/speech signals corresponding to the video. Thus, the present method helps in effective rendering of a video to users.

Claims (32)

1. A method for performing context-based transformation of a video, the method comprising:

generating, by a video transformation system, a scene descriptor of each of one or more scenes of a video based on executing one or more of a computer vision technique or a deep learning technique to identify and associate one or more scene attributes with one or more scene objects;

generating, by the video transformation system, a textual descriptor based on executing a conversion technique on each of one or more speech segments extracted from audio signals in each of the one or more scenes;

determining, by the video transformation system, an audio context descriptor based on executing a semantic analysis of the textual descriptor to identify and associate one or more audio attributes with one or more of the audio objects;

correlating, by the video transformation system, the one or more audio attributes associated with one or more audio objects in the audio context descriptor with the one or more scene attributes associated with one or more scene objects in the scene descriptor to generate a scene context descriptor for at least one of the scenes; and

translating, by the video transformation system, the at least one of the scenes based on execution of a function based on the scene context descriptor to transform the video.

2. The method as claimed in claim 1 further comprises eliminating one or more redundant scenes corresponding to the video upon detecting a similarity in the scene descriptor between two or more of the scenes.

3. The method as claimed in claim 2 , wherein the detecting is determined by quantifying divergence between the scene descriptor of two or more of the scenes.

4. The method as claimed in claim 1 , wherein the generating the scene descriptor is further based on one or more parameters comprising actions performed by the scene objects, or attributes of background of the scene objects in the one or more scenes.

5. The method as claimed in claim 1 , wherein the generating the scene descriptor further comprises generating labels and description for scene objects present in the one or more scenes.

6. A video transformation system for performing context-based transformation of a video, the video transformation system comprising:

a processor; and

a memory, communicatively coupled to the processor, wherein the memory stores processor-executable instructions, which on execution, cause the processor to:

generate a scene descriptor of each of one or more scenes of a video based on executing one or more of a computer vision technique or a deep learning technique to identify and associate one or more scene attributes with one or more scene objects;

generate a textual descriptor based on executing a conversion technique on each of one or more speech segments extracted from audio signals in each of the one or more scenes;

determine an audio context descriptor based on executing a semantic analysis of the textual descriptor to identify and associate one or more audio attributes with one or more of the audio objects;

correlate the one or more audio attributes associated with one or more audio objects in the audio context descriptor with the one or more scene attributes associated with one or more scene objects in the scene descriptor to generate a scene context descriptor for at least one of the scenes; and

translate the at least one of the scenes based on execution of a function based on the scene context descriptor to transform the video.

7. The video transformation system as claimed in claim 6 , wherein the processor eliminates one or more redundant scenes corresponding to the video upon determining similarity in the scene descriptor between two or more of the scenes.

8. The video transformation system as claimed in claim 7 , wherein the processor quantifies divergence between the scene descriptor of two or more of the scenes to determine the similarity among the scene descriptor of the two or more scenes.

9. The video transformation system as claimed in claim 6 , wherein the processor for the generate the scene descriptor is further based on one or more parameters comprising actions performed by the scene objects or attributes of background of the scene objects in the one or more scenes.

10. The video transformation system as claimed in claim 6 , wherein the generate the scene descriptor further comprises generating labels and description for scene objects present in the one or more scenes.

11. A non-transitory computer readable medium having stored thereon instructions for performing context-based transformation of a video comprising executable code which when executed by one or more processors, causes the one or more processors to:

generate a scene descriptor of each of one or more scenes of a video based on executing one or more of a computer vision technique or a deep learning technique to identify and associate one or more scene attributes with one or more scene objects;

generate a textual descriptor based on executing a conversion technique on each of one or more speech segments extracted from audio signals in each of the one or more scenes;

determine an audio context descriptor based on executing a semantic analysis of the textual descriptor to identify and associate one or more audio attributes with one or more of the audio objects;

correlate the one or more audio attributes associated with one or more audio objects in the audio context descriptor with the one or more scene attributes associated with one or more scene objects in the scene descriptor to generate a scene context descriptor for at least one of the scenes; and

translate the at least one of the scenes based on execution of a function based on the scene context descriptor to transform the video.

12. The medium as claimed in claim 11 , wherein the executable code when executed by the one or more processors further causes the one or more processors to eliminate one or more redundant scenes corresponding to the video upon determining a similarity in the scene descriptor between two or more of the scenes.

13. The medium as claimed in claim 12 , wherein the executable code when executed by the one or more processors further causes the one or more processors to quantify divergence between the scene descriptor of two or more of the scenes to determine the similarity among the scene descriptor of the two or more scenes.

14. The medium as claimed in claim 11 , wherein the executable code when executed by the one or more processors further causes the one or more processors for the generate the scene descriptor is further based on one or more parameters comprising actions performed by the scene objects or attributes of background of the scene objects in the one or more scenes.

15. The medium as claimed in claim 11 , wherein the generate the scene descriptor further comprises generating labels and description for scene objects present in the one or more scenes.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2026
From: IP3 2023, SERIES 923 OF ALLIED SECURITY TRUST I
To: RP INTELLECTUAL PARTNERS LLC
Reel/Frame 075782/0624 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: WIPRO LIMITED
To: IP3 2023, SERIES 923 OF ALLIED SECURITY TRUST I
Reel/Frame 066195/0967 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 3, 2018
From: RAMACHANDRA, MANJUNATH; ULAGANATHAN, SETHURAMAN
To: WIPRO LIMITED
Reel/Frame 045830/0623 →
Priority Claims (1)
IN 201841005827 · Feb 15, 2018 · national
Cited By (1)
US 12,243,438