IP Library › Granted Patent US 12,272,384
Granted Patent B1
US 12,272,384 · App. 18/805,819 · Granted Apr 8, 2025

Synchronization of lip movement images to audio voice signal

Inventors: Kyrylo Sydorchuk (Kharkov, UA); Volodymyr Cherniavskyi (Kharkov, UA); Stanislav Mihailevschii (Ribnita, MD); Oleh Vallas (Kharkov, UA); Ivan Shuhaienko (Kharkov, UA); Daniil Krasylnikov (Kharkov, UA); Yurii Astafiev (Kharkov, UA)
Assignee: Pheon, Inc.
G11B27/031G06V10/82G06V20/46G06V20/49G06V40/169G06V40/175G10L15/02G10L15/04G10L15/16G10L25/57
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,384
App. No.
18/805,819
Granted
Apr 8, 2025
Kind
B1
Abstract

Systems and methods for synchronization of lip movement images to an audio voice signal are provided. A method includes acquiring a source video; dividing the source video into a set of image frames and a set of audio frames; generating, by the computing device, a vector database based on the set of image frames and the set of audio frames, wherein a vector of the vector database includes a face vector and an audio vector; receiving a target image frame and a target audio frame; determining a target image vector based on the target image frame and a target audio vector based on the target audio frame; searching the vector database to select a pre-determined number of vectors corresponding to the target image vector and the target audio frame; and generating, based on the pre-determined number of vectors, an output image frame of an output video.

Claims (59)

1. A method comprising:

acquiring, by a computing device, a source video;

dividing, by the computing device, the source video into a set of image frames and a set of audio frames;

generating, by the computing device, a vector database based on the set of image frames and the set of audio frames, wherein a vector of the vector database includes a face vector and an audio vector, the face vector being determined based on an image frame of the set of image frames and the audio vector being determined based an audio frame of the set of audio frames, the audio frame corresponding to the image frame;

receiving, by the computing device, a target image frame and a target audio frame, the target audio frame being selected from a target audio record, wherein the source video is a first record and the target audio record is a second record, the second record being different from the first record;

determining, by the computing device, a target image vector based on the target image frame and a target audio vector based on the target audio frame;

searching, by the computing device, the vector database to select a pre-determined number of vectors corresponding to the target image vector and the target audio frame; and

generating, by the computing device and based on the pre-determined number of vectors, an output image frame of an output video, the output video being a third record, the third record being different from the first record.

2. The method of claim 1 , wherein the acquiring the source video includes capturing, by the computing device, a video featuring a user.

3. The method of claim 1 , wherein the face vector and the target image vector are generated by a vocabulary encoder including a pre-trained neural network.

4. The method of claim 1 , wherein the face vector includes an angle of a rotation of a face in the image frame around an axis.

5. The method of claim 1 , wherein the audio vector and the target audio vector are generated by a speech encoder including a pre-trained neural network.

6. The method of claim 1 , wherein the selection of the pre-determined number of vectors includes:

determining a first metric based on the target image vector and the face vector;

determining a second metric based on the target audio vector and the audio vector; and

combining the first metric and the second metric into a third metric; and

determining that the third metric is below a predetermined threshold.

7. The method of claim 6 , wherein:

the first metric includes a distance between the target image vector and the face vector; and

the second metric includes a scaled dot product of the target audio vector and the audio vector.

8. The method of claim 1 , further comprising, prior to generating the output image frame, extracting style information from the set of image frames, wherein:

the style information indicates a presence or an absence of an emotional expression in a face in the image frames of the set of image frames; and

the output image frame is generated based on the style information.

9. The method of claim 8 , wherein the output image frame is generated by a decoder including a pre-trained neural network.

10. The method of claim 8 , wherein the target image frame is selected from the set of image frames.

11. A computing device comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the computing device to:

acquire, by the computing device, a source video;

divide, by the computing device, the source video into a set of image frames and a set of audio frames;

generate, by the computing device, a vector database based on the set of image frames and the set of audio frames, wherein a vector of the vector database includes a face vector and an audio vector, the face vector being determined based on an image frame of the set of image frames and the audio vector being determined based an audio frame of the set of audio frames, the audio frame corresponding to the image frame;

receive, by the computing device, a target image frame and a target audio frame, the target audio frame being selected from a target audio record, wherein the source video is a first record and the target audio record is a second record, the second record being different from the first record;

determine, by the computing device, a target image vector based on the target image frame and a target audio vector based on the target audio frame;

search, by the computing device, the vector database to select a pre-determined number of vectors corresponding to the target image vector and the target audio frame; and

generate, by the computing device and based on the pre-determined number of vectors, an output image frame of an output video, the output video being a third record, the third record being different from the first record.

12. The computing device of claim 11 , wherein the acquiring the source video includes capturing, by the computing device, a video featuring a user.

13. The computing device of claim 11 , wherein the face vector and the target image vector are generated by a vocabulary encoder including a pre-trained neural network.

14. The computing device of claim 11 , wherein the face vector includes an angle of a rotation of a face in the image frame around an axis.

15. The computing device of claim 11 , wherein the audio vector and the target audio vector are generated by a speech encoder including a pre-trained neural network.

16. The computing device of claim 11 , wherein the selection of the pre-determined number of vectors includes:

determining a first metric based on the target image vector and the face vector;

determining a second metric based on the target audio vector and the audio vector; and

combining the first metric and the second metric into a third metric; and

determining that the third metric is below a predetermined threshold.

17. The computing device of claim 16 , wherein:

the first metric includes a distance between the target image vector and the face vector; and

the second metric includes a scaled dot product of the target audio vector and the audio vector.

18. The computing device of claim 11 , wherein the instructions further configure the computing device to, prior to generating the output image frame, extract style information from the set of image frames, wherein:

the style information indicates a presence or an absence of an emotional expression in a face in the image frames of the set of image frames; and

the output image frame is generated based on the style information.

19. The computing device of claim 18 , wherein the output image frame is generated by a decoder including a pre-trained neural network.

20. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that, when executed by a computing device, cause the computing device to:

acquire a source video;

divide the source video into a set of image frames and a set of audio frames;

generate a vector database based on the set of image frames and the set of audio frames, wherein a vector of the vector database includes a face vector and an audio vector, the face vector being determined based on an image frame of the set of image frames and the audio vector being determined based an audio frame of the set of audio frames, the audio frame corresponding to the image frame;

receive a target image frame and a target audio frame, the target audio frame being selected from a target audio record, wherein the source video is a first record and the target audio record is a second record, the second record being different from the first record;

determine a target image vector based on the target image frame and a target audio vector based on the target audio frame;

search the vector database to select a pre-determined number of vectors corresponding to the target image vector and the target audio frame; and

generate, based on the pre-determined number of vectors, an output image frame of an output video, the output video being a third record, the third record being different from the first record.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2024
From: SYDORCHUK, KYRYLO; CHERNIAVSKYI, VOLODYMYR; MIHAILEVSCHII, STANISLAV; VALLAS, OLEH; SHUHAIENKO, IVAN; KRASYLNIKOV, DANIIL; ASTAFIEV, YURII
To: PHEON, INC.
Reel/Frame 068654/0618 →
References Cited (15)
US 7246314B2 · Foote · 2007 [cited by examiner]
US 8780756B2 · Ogikubo · 2014 [cited by examiner]
US 9338467B1 · Gadepalli · 2016 [cited by examiner]
US 11756291B2 · Turkelson · 2023 [cited by examiner]
US 12153627B1 · Lee · 2024 [cited by examiner]
US 20050232497A1 · Yogeshwar · 2005 [cited by examiner]
US 20190073520A1 · Ayyar · 2019 [cited by examiner]
US 20220245655A1 · Faith · 2022 [cited by examiner]
US 20230238002A1 · Hirano · 2023 [cited by examiner]
US 20230291909A1 · Liang · 2023 [cited by examiner]
US 20240428776A1 · Kang · 2024 [cited by examiner]
Chen et al. “Hierarchical cross-modal talking face generation with dynamic pixel-wise loss”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (2019), pp. 7832-7841. [cited by applicant]
K R Prajwal et al. “A Lip Sync Expert Is All You Need for Speech to Lip Generation in the Wild”, MM '20: Proceedings of the 28th ACM International Conference on Multimedia, Oct. 2020, pp. 484-493. [cited by applicant]
Siarohin et al. “First Order Motion Model for Image Animation”, Conference on Neural Information Processing Systems (NeurIPS), Dec. 2019. [cited by applicant]
Zhou, Hang et al. “Pose-Controllable Talking Face Generation by Implicitly Modularized Audio-Visual Representation.” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2021): pp. 4176-4186. [cited by applicant]