IP Library Granted Patent US 10,917,607
Granted Patent B1
US 10,917,607 · App. 16/601,102 · Granted Feb 9, 2021

Editing text in video captions

Inventors: Vincent Charles Cheung (San Carlos, CA); Marc Layne Hemeon (Haleiwa, HI); Nipun Mathur (Belmont, CA)
Assignee: Facebook Technologies, LLC
H04N5/9305G10L15/26G11B27/036G11B27/34
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,917,607
App. No.
16/601,102
Granted
Feb 9, 2021
Kind
B1
Abstract

This disclosure describes techniques that include modifying text associated with a sequence of images or a video sequence to thereby generate new text and overlaying the new text as captions in the video sequence. In one example, this disclosure describes a method that includes receiving a sequence of images associated with a scene occurring over a time period; receiving audio data of speech uttered during the time period; transcribing into text the audio data of the speech, wherein the text includes a sequence of original words; associating a timestamp with each of the original words during the time period; generating, responsive to input, a sequence of new words; and generating a new sequence of images by overlaying each of the new words on one or more of the images.

Claims (58)

1. A system comprising:

a language processing engine configured to:

receive a sequence of images and audio data associated with a scene occurring over a time period, wherein the audio data includes data representing speech uttered during the time period,

transcribe the audio data of the speech into text, wherein the text includes a sequence of original words,

associate a timestamp with each of the original words during the time period, and

generate, responsive to input, a sequence of new words from the sequence of original words; and

a video processing engine configured to generate a new sequence of images by overlaying each of the new words on one or more of the images, wherein each of the new words is overlaid on one or more of the images based on the timestamps associated with the original words.

2. The system of claim 1 , wherein the language processing engine is further configured to:

associate each of the new words with one or more corresponding original words.

3. The system of claim 1 , further comprising:

a data capture system configured to capture the sequence of images and the audio data.

4. The system of claim 2 , wherein to generate the new sequence of images, the video processing engine is further configured to:

overlay each new word of the new words on one or more of the images based on the timestamp associated with the one or more corresponding original words for the new word.

5. The system of claim 2 ,

wherein to generate the sequence of new words, the language processing engine is further configured to replace an original word with a new word, and

wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to associate the new word with the original word as a corresponding original word.

6. The system of claim 2 ,

wherein to generate the sequence of new words, the language processing engine is further configured to add a new word before an original word, and

wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to associate the new word with the original word as a corresponding original word.

7. The system of claim 2 ,

wherein to generate the sequence of new words, the language processing engine is further configured to add a new word after an original word, and

wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to associate the new word with the original word as a corresponding original word.

8. The system of claim 2 ,

wherein to generate the sequence of new words, the language processing engine is further configured to remove a deleted original word from the sequence of original words, and

wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to associate each of the new words with one or more original words without associating any of the new words with deleted original word.

9. The system of claim 1 , wherein to associate a timestamp with each of the original words, the language processing engine is further configured to:

associate, for each of the original words, a starting timestamp corresponding to the start of the original word during the time period and an ending timestamp corresponding to the end of the original word during the time period.

10. The system of claim 1 , wherein to associate a timestamp with each of the original words, the language processing engine is further configured to:

associate each of the original words with a time period corresponding to a pause between each of the original words.

11. The system of claim 1 , wherein the sequence of original words includes a plurality of unchanged original words representing original words not changed in the sequence of new words, and wherein to generate the new sequence of images, the video processing engine is further configured to:

overlay each of the unchanged original words as a caption on one or more of the images in the sequence of images based on the respective timestamps associated with the unchanged original words.

12. The system of claim 2 , wherein to generate the sequence of new words, the language processing engine is further configured to:

generate a sequence of foreign language words by translating the sequence of original words.

13. The system of claim 2 ,

wherein to generate a sequence of new words, the language processing engine is further configured to generate a sequence of new words modifying each of the original words in the sequence of original words; and

wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to associate each of the new words with one or more corresponding original words based on pacing of the original words in the audio data.

14. The system of claim 2 , wherein to associate each of the new words with one or more corresponding original words, the language processing engine is further configured to:

associate each of the new words with one or more corresponding original words further based on pacing of the original words in the audio data.

15. A method comprising:

receiving, by a computing system, a sequence of images and audio data associated with a scene occurring over a time period, wherein the audio data includes data representing speech uttered during the time period;

transcribing, by the computing system, the audio data of the speech into text, wherein the text includes a sequence of original words;

associating, by the computing system, a timestamp with each of the original words during the time period; and

generating, by the computing system and responsive to input, a sequence of new words from the sequence of original words; and

generating, by the computing system, a new sequence of images by overlaying each of the new words on one or more of the images, wherein each of the new words is overlaid on one or more of the images based on the timestamps associated with the original words.

16. The method of claim 15 , further comprising:

associating, by the computing system, each of the new words with one or more corresponding original words.

17. The method of claim 16 , wherein generating the new sequence of images includes:

overlaying each of new word of the new words on one or more of the images based on the timestamp associated with the one or more corresponding original words for the new word.

18. The method of claim 16 , wherein generating a sequence of new words includes replacing an original word with a new word, and wherein associating each of the new words with one or more corresponding original words includes:

associating the new word with the original word as a corresponding original word.

19. The method of claim 16 , wherein generating a sequence of new words includes adding a new word before an original word, and wherein associating each of the new words with one or more corresponding original words includes:

associating the new word with the original word as a corresponding original word.

20. A non-transitory computer-readable storage medium comprising instructions that, when executed, configure processing circuitry of a computing system to perform operations comprising:

receiving a sequence of images and audio data associated with a scene occurring over a time period, wherein the audio data includes data representing speech uttered during the time period;

parsing the audio data of the speech into text, wherein the text includes a sequence of original words;

associating a timestamp with each of the original words during the time period; and

generating, responsive to input, a sequence of new words from the sequence of new words; and

generating a new sequence of images by overlaying each of the new words on one or more of the images, wherein each of the new words is overlaid on one or more of the images based on the timestamps associated with the original words.

Assignments (2)
CHANGE OF NAME Recorded Jul 21, 2022
From: FACEBOOK TECHNOLOGIES, LLC
To: META PLATFORMS TECHNOLOGIES, LLC
Reel/Frame 060802/0799 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2019
From: CHEUNG, VINCENT CHARLES; HEMEON, MARC LAYNE; MATHUR, NIPUN
To: FACEBOOK TECHNOLOGIES, LLC
Reel/Frame 050826/0039 →
Cited By (1)
US 12,621,518