IP Library Granted Patent US 11,189,320
Granted Patent B2
US 11,189,320 · App. 16/934,562 · Granted Nov 30, 2021

System and methods for concatenating video sequences using face detection

Inventors: Michal Shafir Nir (Tel Aviv, IL); Liat Sade-Sternberg (Pacific Palisades, CA); Rhona Horiner Rosen (Haifa, IL); Tamar Raviv (Burgata, IL)
Assignee: FUSIT, INC.
G11B27/036G06K9/00228G06K9/00765G11B27/031G11B27/28H04N21/2365H04N21/23418H04N21/23424H04N21/234363H04N21/234381H04N21/2665H04N21/44008H04N21/440263H04N21/440281H04N21/4788H04N21/47205H04N21/854
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,189,320
App. No.
16/934,562
Granted
Nov 30, 2021
Kind
B2
Abstract

There are provided methods and devices for media processing, comprising: providing at least one media asset source selected from a media asset sources library, the at least one media asset source comprising at least one source video, via a network or client device; receiving via the network or the client device a media recording comprising a client video recorded by a user of the client device; parsing the client video and the source video, respectively, to a plurality of client video frames and a plurality of source video frames; identifying at least one face in at least one frame of the plurality of source video frames and at least another face in at least one frame of the plurality of client video frames by face detection; superposing one or more markers on the identified at least one face of the plurality of source video frames; processing said client video frames to fit the size or shape of said source video frames by using said one or more markers; concatenating said processed client video frames with said source video frames, wherein said concatenation comprises matching the frame rate and resolution of the processed client video frames to the frame rate and resolution of the plurality of client video frames to yield a mixed media asset.

Claims (55)

1. A method for media processing, comprising:

providing at least one media asset source selected from a media asset sources library, the at least one media asset source comprising at least one source video, via a network or client device, wherein the at least one source video comprises a plurality of source video frames;

receiving via the network or the client device a media recording comprising a client video recorded by a user of the client device, wherein the client video comprises a plurality of client video frames and wherein the frame rate and bit rate of the at least one source video is different from the frame rate and bit rate of the client video;

parsing the client video and the source video, respectively, to the plurality of client video frames and the plurality of source video frames;

identifying at least one face image in at least one frame of the plurality of source video frames and at least other face image in at least one frame of the plurality of client video frames using a face detection method;

superposing one or more markers on the identified at least one face image in the at least one frame of the plurality of source video frames;

processing said client video frames and said source video frames using an editing module and an analyzer module to fit the size or shape of elements in said client video frames to the size or shape of elements in said source video frames;

mixing said processed client video frames with said source video frames, using a mixing module, to yield a mixed media asset;

filtering the new mixed media asset using a filter generator module, wherein said filtering comprises:

matching the frame rate, the bit rate and resolution of the processed client video frames to the frame rate, bit rate and resolution of the plurality of client video frames; and grouping the processed client video frames of the plurality of client video frames to yield a mixed and coherent media asset.

2. The method of claim 1 , wherein said processing comprises:

extracting said one or more markers from said at least identified one face image and superposing the extracted one or more markers on the other at least one identified face; and

resizing the other at least one identified face to match the size or shape of the at least one identified face in the at least one frame of the plurality of source video frames.

3. The method of claim 2 , wherein said processing further comprises cropping or scaling the at least one other identified face image in the at least one frame of the plurality of client video frames to match the size or shape of said identified at least one face image in the source video frames.

4. The method of claim 1 , wherein said filtering comprises one or more of the following filtering procedures:

fast forward filtering; slow motion filtering; color filtering; black and white filtering.

5. The method of claim 1 , wherein said face detection method comprise identifying the position of one or more face elements in at least one frame of the plurality of video source frames and other one or more face elements in the at least one client video frames.

6. The method of claim 1 , wherein said media asset source is in a form selected from the group consisting of:

graphics interchange format (GIF); or MP4; VP8; m4v; mov; avi; fav; mpg; wmv; h265, and the client video form is in a different form.

7. The method of claim 1 , wherein said face detection are selected from a face detection algorithms group consisting of:

SMQT Features and SNOW Classifier Method (SFSC),

Efficient and Rank Deficient Face Detection Method (ERDFD),

Gabor-Feature Extraction and Neural Network Method (GFENN),

an efficient face candidates selector Features Method (EFCSF).

8. An apparatus for media processing, comprising:

a memory which is configured to hold one or more source media videos, wherein each of the one or more source media videos comprises a plurality of source video frames; and

a processor which is configured to:

transmit the one or more source media videos to a client device;

receive, via the network or the client device, a media recording comprising a client video recorded by a user of the client device, wherein the client video comprises a plurality of client video frames and wherein the frame rate and bit rate of the one or more source media videos is different from the frame rate and bit rate of the client video;

parse the client video and the source video, respectively, to the plurality of client video frames and the plurality of source video frames;

identify at least one face image in at least one frame of the plurality of source video frames and at least other face in at least one frame of the plurality of client video frames using face detection method;

superpose one or more markers on the identified at least one face image in the at least one frame of the plurality of source video frames;

process said client video frames and said source video frames using an editing module and an analyzer module to fit the size or shape of elements in said client video frames to the size or shape of elements in said source video frames;

mix said processed client video frames with said source video frames, using a mixing module, to yield a new mixed media asset; and

filter the new mixed media asset, wherein said filtering comprises:

matching the frame rate, the bit rate and the resolution of the processed client video frames to the frame rate, bit rate and the resolution of the plurality of client video frames; and

grouping the processed client video frames of the plurality of client video frames to generate a mixed and coherent media asset.

9. The apparatus of claim 8 , wherein said processing comprises:

extracting said one or more markers from said identified at least one face image and superposing the extracted one or more markers on the other at least one identified face; and

resizing the other at least one identified face to match the size and shape of the at least one identified face in the at least one frame of the plurality of source video frames.

10. The apparatus of claim 8 , wherein said processing further comprises cropping or scaling the at least one other identified face image in the at least one frame of the plurality of client video frames to match the size or shape of said identified at least one face image in the source video frames.

11. A computer software product, comprising a non-transitory computer-readable medium in which program instructions are stored, which instructions, when read by a computer, cause the computer to provide at least one media asset source selected from a media asset sources library, the at least one media asset source comprising at least one source video, via a network to a client device, wherein the at least one source video comprises a plurality of source video frames;

receive via the network from the client device a media recording comprising a client video recorded by a user of the client device, wherein the client video comprises a plurality of client video frames and wherein the frame rate and bit rate of the one or more source media videos is different from the frame rate and bit rate of the client video;

parse the client video and the source video, respectively, to the plurality of client video frames and the plurality of source video frames;

identify at least one face image in at least one frame of the plurality of source video frames and at least other face in at least one frame of the plurality of client video frames using face detection method;

superpose one or more markers on the identified at least one face image in the at least one frame of the plurality of source video frames;

simultaneously process and edit said client video frames and said source video frames using an editing module and an analyzer module to fit the size or shape of elements in said client video frames to the size or shape of elements in said source video frames, said processing and editing comprises:

extracting said one or more markers from said identified at least one face image in the at least one frame of the plurality of source video frames; superposing the extracted one or more markers on the at least other face image; and

resizing the at least other face image to match the size or shape of the identified at least one face image, using the superposed one or more markers;

mix said processed client video frames with said source video frames, using a mixing module, to yield a new mixed media asset;

filter the new mixed media asset, wherein said filtering comprises:

matching the frame rate, the bit rate and the resolution of the processed client video frames to the frame rate, the bit rate and the resolution of the plurality of client video frames; and

grouping the processed client video frames of the plurality of client video frames to generate a mixed and coherent media asset.

12. The method of claim 1 , comprising analyzing and editing simultaneously said client video frames and said source video frames using the editing module and the analyzer module to fit the size or shape of elements in said client video frames to the size or shape of elements in said source video frames, by analyzing successively each of the frames of said client video frames and said source video frames to identify one of more faces at said client video frames and said source video frames, and editing the identified face at each frame of said client video frames according to the location of the identified face at each of the related preceding frame.

13. The apparatus of claim 8 , wherein the editing module and the analyzer module are configured to simultaneously fit the size or shape of elements in said client video frames to the size or shape of elements in said source video frames by analyzing successively each of the frames of said client video frames and said source video frames to identify one of more faces at said client video frames and said source video frames, and editing the identified face at each frame of said client video frames according to the location of the identified face at each of the related preceding frame.

Assignments (3)
NUNC PRO TUNC ASSIGNMENT Recorded Jul 21, 2020
From: FUSIC LTD.
To: NEW DE FUSIT, INC.
Reel/Frame 053268/0389 →
MERGER Recorded Jul 21, 2020
From: NEW DE FUSIT, INC.
To: FUSIT, INC.
Reel/Frame 053268/0414 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 21, 2020
From: NIR, MICHAL SHAFIR; SADE-STERNBERG, LIAT; ROSEN, RHONA HORINER; RAVIV, TAMAR
To: FUSIC LTD.
Reel/Frame 053269/0977 →
Continuity (3)
Continuation 15897270 · Feb 15, 2018
Provisional Application 62459620 · Feb 16, 2017
Related Publication 20200349978A1 · Nov 5, 2020