IP Library › Granted Patent US 10,021,276
Granted Patent B1
US 10,021,276 · App. 15/905,934 · Granted Jul 10, 2018

Method and device for processing video, electronic device and storage medium

Inventor: Hanwen Chang (Beijing, CN)
Assignee: BEIJING KINGSOFT INTERNET SECURITY SOFTWARE CO., LTD.
H04N5/04G06F17/28G06K9/00275G06K9/00288G10L15/265H04N5/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,021,276
App. No.
15/905,934
Granted
Jul 10, 2018
Kind
B1
Abstract

Embodiments of the present disclosure provide a method and a device for processing a video, an electronic device and a storage medium. The method includes the followings. A target recognition is performed to A first video segments and B first speech segments to obtain M second video segments and N second speech segments. A speech processing is performed to the N second speech segments to obtain N target speech text files. First representation information is extracted from the M second video segments to obtain Q pieces of first representation information. A second sound matched with the target character is determined according to the Q pieces of first representation information. The second sound is merged with the N target speech text files to obtain N target speech segments.

Claims (69)

1. A method for processing a video, comprising:

performing a target recognition to A first video segments and B first speech segments to obtain M second video segments and N second speech segments, wherein the A first video segments and the B first speech segments are obtained by dividing an inputted video, the M second video segments comprise a first face image of a target character, the N second speech segments comprise a first sound of the target character, A is an integer greater than 1, B is a positive integer, M is a positive integer less than or equal to A, and N is a positive integer less than or equal to B;

performing a speech processing to the N second speech segments to obtain N target speech text files, wherein the N second speech segments correspond to the N target speech text files one by one;

extracting first representation information of the first face image from the M second video segments to obtain Q pieces of first representation information, wherein Q is an integer greater than or equal to M;

determining a second sound matched with the target character according to the Q pieces of first representation information; and

merging the second sound with the N target speech text files to obtain N target speech segments, wherein the N target speech text files correspond to the N target speech segments one by one.

2. The method according to claim 1 , wherein performing a speech processing to the N second speech segments to obtain N target speech text files comprises:

performing a speech recognition to the N second speech segments to obtain N text files, wherein the N second speech segments correspond to the N text files one by one; and

translating the N text files according to a specified language to obtain the N target speech text files, wherein the N text files correspond to the N target speech text files one by one.

3. The method according to claim 1 , wherein, extracting first representation information of the first face image from the M second video segments to obtain Q pieces of first representation information comprises:

performing a first representation information extraction to the first face image of each of the M second video segments or the first face image of each of L frames comprising the first face image in the M second video segments, so as to obtain the Q pieces of first representation information, wherein L is a positive integer.

4. The method according to claim 1 , wherein determining a second sound matched with the target character according to the Q pieces of first representation information comprises:

classifying the Q pieces of first representation information to obtain P classes of first representation information, wherein P is a positive integer less than or equal to Q; and

determining the second sound according to one of the P classes with a longest playing period among the inputted video.

5. The method according to claim 1 , wherein, after obtaining the Q pieces of first representation information, the method further comprises:

determining a second face image matched with the target character according to the Q pieces of first representation information;

replacing the first face image in the M second video segments with the second face image to obtain M target video segments, wherein the M video segments correspond to the M target video segments one by one; and

merging the N target speech segments with the M target video segments to obtain an outputted video.

6. The method according to claim 5 , further comprising:

pre-processing the second face image to obtain a third face image;

replacing facial features of the third face image with those of the first image face to obtain a fourth face image;

rectifying the fourth face image with a loss function to obtain a fifth face image;

merging the fifth face image with rest of the target face image except the second face image, to obtain an outputted image.

7. The method according to claim 6 , wherein the pre-processing comprises at least one of a face alignment, an image enhancement and a normalization.

8. The method according to claim 1 , wherein, before performing a target recognition to A first video segments and B first speech segments, the method further comprises:

dividing the inputted video into the A first video segments according to a preset period or a playing period of the inputted video; and

dividing the inputted video into the B first speech segments according to a preset volume threshold.

9. The method according to claim 1 , further comprising:

extracting second representation information of the first sound from the N second speech segments to obtain R pieces of second representation information, wherein R is an integer greater than or equal to N; and

determining the second sound according to the Q pieces of first representation information and the R pieces of second representation information.

10. An electronic device, comprising: a housing, a processor, a memory, a circuit board and a power supply circuit; wherein the circuit board is enclosed by the housing; the processor and the memory are positioned on the circuit board; the power supply circuit is configured to provide power for respective circuits or components of the electronic device; the memory is configured to store executable program codes; and the processor is configured to run a program corresponding to the executable program codes by reading the executable program codes stored in the memory, to perform a method for processing a video, the method comprising:

performing a target recognition to A first video segments and B first speech segments to obtain M second video segments and N second speech segments, wherein the A first video segments and the B first speech segments are obtained by dividing an inputted video, the M second video segments comprise a first face image of a target character, the N second speech segments comprise a first sound of the target character, A is an integer greater than 1, B is a positive integer, M is a positive integer less than or equal to A, and N is a positive integer less than or equal to B;

performing a speech processing to the N second speech segments to obtain N target speech text files, wherein the N second speech segments correspond to the N target speech text files one by one;

extracting first representation information of the first face image from the M second video segments to obtain Q pieces of first representation information, wherein Q is an integer greater than or equal to M;

determining a second sound matched with the target character according to the Q pieces of first representation information; and

merging the second sound with the N target speech text files to obtain N target speech segments, wherein the N target speech text files correspond to the N target speech segments one by one.

11. The electronic device according to claim 10 , wherein the processor is configured to perform a speech processing to the N second speech segments to obtain N target speech text files by acts of:

performing a speech recognition to the N second speech segments to obtain N text files, wherein the N second speech segments correspond to the N text files one by one; and

translating the N text files according to a specified language to obtain the N target speech text files, wherein the N text files correspond to the N target speech text files one by one.

12. The electronic device according to claim 10 , wherein the processor is configured to extract first representation information of the first face image from the M second video segments to obtain Q pieces of first representation information by acts of:

performing a first representation information extraction to the first face image of each of the M second video segments or the first face image of each of L frames comprising the first face image in the M second video segments, so as to obtain the Q pieces of first representation information, wherein L is a positive integer.

13. The electronic device according to claim 10 , wherein the processor is configured to determine a second sound matched with the target character according to the Q pieces of first representation information by acts of:

classifying the Q pieces of first representation information to obtain P classes of first representation information, wherein P is a positive integer less than or equal to Q; and

determining the second sound according to one of the P classes with a longest playing period among the inputted video.

14. The electronic device according to claim 10 , wherein the processor is configured to, after obtaining the Q pieces of first representation information, perform acts of:

determining a second face image matched with the target character according to the Q pieces of first representation information;

replacing the first face image in the M second video segments with the second face image to obtain M target video segments, wherein the M video segments correspond to the M target video segments one by one; and

merging the N target speech segments with the M target video segments to obtain an outputted video.

15. The electronic device according to claim 14 , wherein processor is configured to perform acts of:

pre-processing the second face image to obtain a third face image;

replacing facial features of the third face image with those of the first image face to obtain a fourth face image;

rectifying the fourth face image with a loss function to obtain a fifth face image;

merging the fifth face image with rest of the target face image except the second face image, to obtain an outputted image.

16. The electronic device according to claim 15 , wherein the pre-processing comprises at least one of a face alignment, an image enhancement and a normalization.

17. The electronic device according to claim 10 , wherein the processor is configured to, before performing a target recognition to A first video segments and B first speech segments, perform acts of:

dividing the inputted video into the A first video segments according to a preset period or a playing period of the inputted video; and

dividing the inputted video into the B first speech segments according to a preset volume threshold.

18. The electronic device according to claim 10 , wherein the processor is configured to perform acts of:

extracting second representation information of the first sound from the N second speech segments to obtain R pieces of second representation information, wherein R is an integer greater than or equal to N; and

determining the second sound according to the Q pieces of first representation information and the R pieces of second representation information.

19. A non-transitory computer readable storage medium, having computer programs stored therein, when the computer programs are executed by a processor, a method for processing a video is realized, the method comprising:

performing a target recognition to A first video segments and B first speech segments to obtain M second video segments and N second speech segments, wherein the A first video segments and the B first speech segments are obtained by dividing an inputted video, the M second video segments comprise a first face image of a target character, the N second speech segments comprise a first sound of the target character A is an integer greater than 1, B is a positive integer, M is a positive integer less than or equal to A, and N is a positive integer less than or equal to B;

performing a speech processing to the N second speech segments to obtain N target speech text files, wherein the N second speech segments correspond to the N target speech text files one by one;

extracting first representation information of the first face image from the M second video segments to obtain Q pieces of first representation information, wherein Q is an integer greater than or equal to M;

determining a second sound matched with the target character according to the Q pieces of first representation information; and

merging the second sound with the N target speech text files to obtain N target speech segments, wherein the N target speech text files correspond to the N target speech segments one by one.

20. The non-transitory computer readable storage medium according to claim 19 , wherein performing a speech processing to the N second speech segments to obtain N target speech text files comprises:

performing a speech recognition to the N second speech segments to obtain N text files, wherein the N second speech segments correspond to the N text files one by one; and

translating the N text files according to a specified language to obtain the N target speech text files, wherein the N text files correspond to the N target speech text files one by one.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2023
From: JOYINME PTE. LTD.
To: JUPITER PALACE PTE. LTD.
Reel/Frame 064190/0756 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2021
From: BEIJING KINGSOFT INTERNET SECURITY SOFTWARE CO., LTD.
To: JOYINME PTE. LTD.
Reel/Frame 055658/0541 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 13, 2018
From: CHANG, HANWEN
To: BEIJING KINGSOFT INTERNET SECURITY SOFTWARE CO., LTD.
Reel/Frame 045189/0693 →
Priority Claims (1)
CN 2017 1 0531697 · Jun 30, 2017 · national
Cited By (1)
US 12,400,438