IP Library Granted Patent US 10,460,732
Granted Patent B2
US 10,460,732 · App. 15/473,029 · Granted Oct 29, 2019

System and method to insert visual subtitles in videos

Inventors: Chitralekha Bhat (Thane, IN); Sunil Kumar Kopparapu (Thane, IN); Ashish Panda (Thane, IN)
Assignee: Tata Consultancy Services Limited
G10L15/265G10L15/04G10L15/183G10L15/25G10L15/26G10L21/06G10L21/10G10L25/57G10L25/63G10L25/81G11B27/036G10L2015/025G10L2021/065G10L2021/105
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,732
App. No.
15/473,029
Granted
Oct 29, 2019
Kind
B2
Abstract

A system and method to insert visual subtitles in videos is described. The method comprises segmenting an input video signal to extract the speech segments and music segments. Next, a speaker representation is associated for each speech segment corresponding to a speaker visible in the frame. Further, speech segments are analyzed to compute the phones and the duration of each phone. The phones are mapped to a corresponding viseme and a viseme based language model is created with a corresponding score. Most relevant viseme is selected for the speech segments by computing a total viseme score. Further, a speaker representation sequence is created such that phones and emotions in the speech segments are represented as reconstructed lip movements and eyebrow movements. The speaker representation sequence is then integrated with the music segments and super imposed on the input video signal to create subtitles.

Claims (44)

1. A computer implemented system for creation of visually subtitled videos, wherein said system comprises:

a memory storing instructions and coupled to the processor, wherein the processor is configured to execute the instructions to perform operations to:

segment an audio signal with at least one speaker in a video frame from an input video signal and a music segment from the input video signal;

map a plurality of recognized phones from the segmented audio signal to at least one viseme using a language model;

determine a likelihood score of the recognized phones;

determine a viseme based language model score (VLM) for the at least one mapped viseme;

compute a total viseme score for the at least one mapped viseme by determining a sum of the likelihood score and the VLM score when more than one phone maps to the same viseme or determining a product of the likelihood score and VLM score when one phone maps to a single viseme;

determine a most relevant viseme from the segmented audio signal by comparing the total viseme score and selecting a viseme with the highest total score;

generate a speaker representation sequence of the segmented audio signal; and

integrate the segmented audio signal and viseme sequence to generate visual subtitles.

2. The system as claimed in claim 1 , wherein the processor further executes the instructions to create a speaker representation by separating speech segments for at least one speaker visible in the video frame and separating music segments from the video frame.

3. The system as claimed in claim 2 , wherein the processor further executes the instructions to extract a speaker face of the at least one speaker visible in the video frame and associating the speaker face with the speech segments.

4. The system as claimed in claim 3 , wherein the processor further executes the instructions to create a speaker representation for at least one speaker not visible in the video frame by creating a generic speaker face, wherein a gender of the at least one speaker not visible in the video frame is identified from the speech segments.

5. The system as claimed in claim 2 , wherein the processor further executes the instructions to:

remove a lip region and a mouth region of the speaker representation; and

reconstruct a lip movement and an eyebrow movement sequence using a series of facial animation points executed through an interpolation simulation corresponding to the most relevant viseme for a computed time duration of the most relevant viseme.

6. The system as claimed in claim 1 , wherein the processor further executes the instructions to integrate the music segment and a speaker representation sequence in time synchrony with the input video signal.

7. A computer implemented method for creation of visually subtitled videos, wherein the method comprises:

segmenting at least one audio signal from an input video frame of a video signal;

mapping a plurality of recognized phones from the segmented audio signal to at least one viseme using a language model;

determining a likelihood score of the recognized phones;

determining a viseme based language model (VLM) score for the at least one mapped viseme;

computing a total viseme score for the at least one mapped viseme by determining a sum of the likelihood score and the VLM score when more than one phone maps to the same viseme or determining a product of the likelihood score and VLM score when one phone maps to a single viseme;

analysing the segmented audio signal for the input video frame to select the most relevant viseme by comparing the total viseme score and selecting a viseme with the highest total score;

generating a speaker representation sequence subsequent to audio segmentation; and

integrating the video signal with the speaker representation sequence to generate visual subtitles.

8. The method as claimed in claim 7 , wherein the segmenting of the at least one audio signal from an input video frame further comprises:

separating speech segments for at least one speaker visible in the video frame and separating music segments from the video frame; and

associating at least one acoustic model to a speaker representation.

9. The method as claimed in claim 8 , further comprising extracting a speaker face of the at least one speaker visible in the video frame and associating the speaker face with the speech segments.

10. The method as claimed in claim 8 , further comprising creating a generic speaker face for at least one speaker not visible in the video frame, wherein a gender of the at least one speaker not visible in the video frame is identified from the speech segments.

11. The method as claimed in claim 7 , wherein generating the speaker representation sequence further comprises removing a lip region and a mouth region of at least one speaker representation of the speaker representation sequence and reconstructing a lip movement and an eyebrow movement sequence.

12. The method as claimed in claim 11 , wherein generating the speaker representation sequence further comprises using a series of facial animation points to form a lip shape or an eyebrow shape, executed through an interpolation simulation corresponding to the most relevant viseme for a computed time duration of the most relevant viseme.

13. The method as claimed in 12 , wherein the eyebrow movement and the lip movement correspond to emotion detected in the speech segments of the segmented audio.

14. The method as claimed in claim 7 , wherein the integrating the video signal further comprises integrating a viseme sequence and a notational representation of a music segment in time synchrony with the video signal.

15. One or more non-transitory machine readable information storage mediums comprising one or more instructions which when executed by one or more hardware processors causes:

segmenting at least one audio signal from an input video frame of a video signal;

mapping a plurality of recognized phones from the segmented audio signal to at least one viseme using a language model;

determining a likelihood score of the recognized phones;

determining a viseme based language model score (VLM) for the at least one mapped viseme;

computing a total viseme score for the at least one mapped viseme by determining a sum of the likelihood score and the VLM score when more than one phone maps to the same viseme or determining a product of the likelihood score and VLM score when one phone maps to a single viseme;

analysing the segmented audio signal for the input video frame to select the most relevant viseme by comparing the total viseme score and selecting a viseme with the highest total score;

generating a speaker representation sequence subsequent to audio segmentation; and

integrating the video signal with the speaker representation sequence to generate visual subtitles.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2017
From: BHAT, CHITRALEKHA; KOPPARAPU, SUNIL KUMAR; PANDA, ASHISH
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 041797/0458 →
Priority Claims (1)
IN 201621011523 · Mar 31, 2016 · national
Continuity (1)
Related Publication 20170287481A1 · Oct 5, 2017
Cited By (1)
US 12,615,025