IP Library › Granted Patent US 11,227,638
Granted Patent B2
US 11,227,638 · App. 16/935,222 · Granted Jan 18, 2022

Method, system, medium, and smart device for cutting video using video content

Inventors: Qianya Lin (Guangdong, CN); Tian Xia (Guangdong, CN); RemyYiYang Ho (Guangdong, CN); Zhenli Xie (Guangdong, CN); Pinlin Chen (Guangdong, CN); Rongchan Liu (Guangdong, CN)
Assignee: Sunday Morning Technology (Guangzhou) Co., Ltd.
G11B27/06G06K9/00281G06K9/00335G06K9/00744G10L15/26G10L21/0208G10L25/57G11B27/036G11B27/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,227,638
App. No.
16/935,222
Granted
Jan 18, 2022
Kind
B2
Abstract

The present invention discloses a method and system for cutting video using video content. The method comprises: acquiring recorded video produced by user's recording operation; extracting features of recorded audio in the recorded video and judging whether the recorded audio is damaged; and if not, extracting human voice data from the recorded audio which has been filtered out background sound, intercepting video segment corresponding to effective human voice, and displaying the video segment as clip video; and if yes, extracting image feature data of person's mouth shape and human movements in the recorded video after image processing, fitting the image feature data and the human voice data which has been filtered out background sound, and displaying the video segment with the highest fitting degree as clip video.

Claims (29)

1. A system for cutting video using video content, wherein the system comprises:

an acquisition module, the acquisition module is configured to acquire recorded video produced by user's recording operation;

a judgement module, the judgement module is configured to extract features of recorded audio in the recorded video and judge whether the recorded audio is damaged;

an intercept module, the intercept module is configured to extract human voice data from the recorded audio which has been filtered out background sound, intercept video segment corresponding to effective human voice, and display the video segment as clip video; and

a fitting module, the fitting module is configured to extract image feature data of person's mouth shape and human movements in the recorded video after image processing, fit the image feature data and the human voice data which has been filtered out background sound, and display the video segment with the highest fitting degree as clip video;

wherein the intercept module comprises:

a recording unit, the recording unit is configured to identify human voice video segments in the recorded video through AI model, extract effective human voice data in the human voice video segments, filter out background sound, and record first time range corresponding to the effective human voice data;

an adjustment unit, the adjustment unit is configured to convert the effective human voice data which has been filtered out background sound into a text, record second time range corresponding to the text, and adjust the first time range according to the second time range; and

a synthesizing unit, the synthesizing unit is configured to clip and synthesize the video segments including the effective human voice data according to the effective human voice data, the text corresponding to the effective human voice data, the adjusted time range and video picture in the recorded video, and display the effect of the obtained clip video.

2. A non-transitory medium on which a computer program is stored, wherein when the program is executed by a processor, a method for cutting video using video content is implemented;

wherein the method comprises the following steps:

acquiring recorded video produced by user's recording operation;

extracting features of recorded audio in the recorded video and judging whether the recorded audio is damaged; and

if not, extracting human voice data from the recorded audio which has been filtered out background sound, intercepting video segment corresponding to effective human voice, and displaying the video segment as clip video; and

if yes, extracting image feature data of person's mouth shape and human movements in the recorded video after image processing, fitting the image feature data and the human voice data which has been filtered out background sound, and displaying the video segment with the highest fitting degree as clip video;

wherein the method for extracting human voice data from the recorded audio which has been filtered out background sound, intercepting video segment corresponding to effective human voice, and displaying the video segment as clip video comprises:

identifying human voice video segments in the recorded video through AI model, extracting effective human voice data in the human voice video segments, filtering out background sound, and recording first time range corresponding to the effective human voice data;

converting the effective human voice data which has been filtered out background sound into a text, recording second time range corresponding to the text, and adjusting the first time range according to the second time range; and

clipping and synthesizing the video segments including the effective human voice data according to the effective human voice data, the text corresponding to the effective human voice data, the adjusted time range and video picture in the recorded video, and displaying the effect of the obtained clip video.

3. A smart device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, a method for cutting video using video content is implemented;

wherein the method comprises the following steps:

acquiring recorded video produced by user's recording operation;

extracting features of recorded audio in the recorded video and judging whether the recorded audio is damaged; and

if not, extracting human voice data from the recorded audio which has been filtered out background sound, intercepting video segment corresponding to effective human voice, and displaying the video segment as clip video; and

if yes, extracting image feature data of person's mouth shape and human movements in the recorded video after image processing, fitting the image feature data and the human voice data which has been filtered out background sound, and displaying the video segment with the highest fitting degree as clip video;

wherein the method for extracting human voice data from the recorded audio which has been filtered out background sound, intercepting video segment corresponding to effective human voice, and displaying the video segment as clip video comprises:

identifying human voice video segments in the recorded video through AI model, extracting effective human voice data in the human voice video segments, filtering out background sound, and recording first time range corresponding to the effective human voice data;

converting the effective human voice data which has been filtered out background sound into a text, recording second time range corresponding to the text, and adjusting the first time range according to the second time range; and

clipping and synthesizing the video segments including the effective human voice data according to the effective human voice data, the text corresponding to the effective human voice data, the adjusted time range and video picture in the recorded video, and displaying the effect of the obtained clip video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 23, 2020
From: LIN, QIANYA; XIA, TIAN; HO, REMYYIYANG; XIE, ZHENLI; CHEN, PINLIN; LIU, RONGCHAN
To: SUNDAY MORNING TECHNOLOGY (GUANGZHOU) CO., LTD.
Reel/Frame 053291/0458 →
Priority Claims (1)
CN 202010281326.1 · Apr 10, 2020 · national
Continuity (1)
Related Publication 20210319809A1 · Oct 14, 2021