IP Library › Granted Patent US 12,456,180
Granted Patent B2
US 12,456,180 · App. 18/059,395 · Granted Oct 28, 2025

System and method for an audio-visual avatar evaluation

Inventors: Ilya Baimetov (Redmond, WA); Denis Parkhomenko (Moscow, RU); Marcel de Korte (Stockholm, SE); Ivan Kirillov (Moscow, RU); Dmitriy Obukhov (Istanbul, TR); Alexey Rybak (Istanbul, TR); Laurent Dedenis (Singapore, SG); Serg Bell (Singapore, SG); Stanislav Protasov (Singapore, SG)
Assignee: Constructor Technology AG
G06T7/0002G10L15/02G10L15/22G10L25/57G10L25/60G10L25/84G10L25/90G06T2207/10016G06T2207/20081G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,180
App. No.
18/059,395
Granted
Oct 28, 2025
Kind
B2
Abstract

The present disclosure relates to a system to evaluate an avatar generated by an avatar generator. The system comprises an evaluation module including an audio evaluation module for evaluating audio features and a video evaluation module for evaluating video features. Evaluation of the avatar includes extracting audio and video features from the avatar and applying a set of evaluation metrics for generating audio and video evaluation scores. The scores are combined to generate a final score. For avatar generator evaluation, audio clip and video clip are provided to the audio evaluation module and video evaluation module, respectively. A set of evaluation metrics is applied for evaluation. Each metric can generate a score. All scores are combined to generate a final evaluation score.

Claims (48)

1. A method for automated evaluation of an avatar generated by an avatar generator comprising the steps of:

obtaining, by an audio evaluator, a speech generated by a text-to-speech module;

obtaining, by a video evaluator, a video clip generated by a video generator;

obtaining, by the audio evaluator, the audio features of the target person;

obtaining, by the video evaluator, the video features of the target person;

comparing the speech with the audio features of the target person using a set of audio metrics, and generating an audio evaluation score for the speech;

wherein generating the audio evaluation score comprises evaluating speech intelligibility using automatic-speech-recognition (ASR) based evaluation metrics, evaluating audio noise level using voice-activity-detection (VAD) based evaluation metrics, evaluating naturalness of speech intonation using pitch-based metrics, evaluating voice similarities using equal-error-rate (EER) and cosine (COS) metrics, and evaluating speech pronunciation statistics;

comparing the video clip with the video features of the target person using a set of video metrics, and generating a video evaluation score for the video clip;

combining the audio evaluation score and the video evaluation score; and

generating a combined naturalness score for the avatar generator based on the combined score of the audio evaluation score and the video evaluation score.

2. The method of claim 1 , wherein evaluation scores generated by each of the set of audio metrics are combined to generate the audio evaluation score.

3. The method of claim 1 , wherein the step of evaluating the video clip further comprises:

evaluating a video quality with a reference image using a peak signal-to-noise ratio (PSNR), a multi-scale structured similarity indexing method (MS-SSIM), a feature similarity indexing method (FSIM), a learned perceptual image patch similarity (LPIPS), a video multimethod assessment fusion (VMAF), visual information fidelity, (VIF), or natural language processing (NLP) metrics;

evaluating a video quality with no reference images using a deep neural network-based IQA model (WaDIQaM), a deep bilinear convolutional neural network (DUBCNN), a transformer, relative ranking, and self consistency (TReS) model, or a chip quality assurance (ChipQA) model;

evaluating distribution using distribution-based metrics;

evaluating lip synchronization using lip synchronization metrics; and

evaluating an identity of the target using identity metrics.

4. The method of claim 3 , wherein the evaluation scores generated by each of the set of video metrics are combined to generate a video evaluation score.

5. The method of claim 1 , wherein the step of generating a combined naturalness score includes generating one or more human-interpretable scores of an avatar.

6. The method of claim 1 , wherein the step of combining the audio evaluation score and the video evaluation score comprises combining using at least one of a weighted average with fixed weights method and a trainable combination method.

7. The method of claim 6 , wherein the weighted average with fixed weights method comprises scaling all evaluation scores to a predefined range of weights and determining an average of the weights.

8. The method of claim 6 , wherein the trainable combination method comprises: using a dataset containing pairs of a video of the target person and corresponding mean opinion scores to train a regression module to predict the final score.

9. An evaluation system for automated evaluation of an avatar generated by an avatar generator comprising:

a processor configured to:

obtain a speech generated by a text-to-speech module of the avatar generator;

obtain the audio features of the target person;

compare the speech and the audio features using a set of audio metrics; wherein the audio metrics comprise automatic-speech-recognition (ASR) based evaluation metrics voice-activity-detection (VAD) based evaluation metrics, gross-pitch-error (GPE), or F0-frame-error (FFE) metrics, and equal-error-rate (EER) and cosine (COS) metrics, and generate an audio evaluation score based on the comparison;

the processor further configured to:

obtain a video clip generated by a video generator of the avatar generator;

obtain the video features of the target person;

compare the video clip with the video features using a set of video metrics; and

generate a video evaluation score based on the comparison; and

a score combination module configured to:

combine the audio evaluation score and the video evaluation score;

generate a combined naturalness score for the avatar generator based on the combination; and

generate an overall naturalness score based on the naturalness score and the combined naturalness score.

10. The system of claim 9 , wherein evaluation scores generated by each of the set of audio metrics are combined to generate the audio evaluation score.

11. The system of claim 9 , wherein the set of video metric includes:

a peak signal-to-noise ratio PSNR, a multi-scale structured similarity indexing method MS-SSIM, a feature similarity indexing method FSIM, a learned perceptual image patch similarity LPIPS, a video multimethod assessment fusion VMAF, visual information fidelity VIF, or natural language processing NLP metrics;

a deep neural network-based IQA model (WaDIQaM), a deep bilinear convolutional neural network (DUBCNN), a transformer, relative ranking, and self consistency (TReS), or chip quality assurance (ChipQA) model;

distribution-based metrics;

lip synchronization metrics; or

identity metrics.

12. The system of claim 9 , wherein the evaluation scores generated by each of the set of video metrics are combined to generate a video evaluation score.

13. The system of claim 9 , wherein the score combination module generates one or more human-interpretable scores of an avatar.

14. The system of claim 9 , wherein the score combination module is configured to combine the audio evaluation score and the video evaluation score using at least one of a weighted average with fixed weights method and a trainable combination method.

15. The system of claim 14 , wherein the score combination module using the weighted average with fixed weights method is configured to scale all evaluation scores to a predefined range of weights and determine an average of the weights.

16. The system of claim 14 , wherein the score combination module using the trainable combination method is configured to use a dataset containing pairs of a video of the target person and corresponding mean opinion scores to train a regression module to predict the final score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2025
From: BAIMETOV, ILYA; PARKHOMENKO, DENIS; DE KORTE, MARCEL; KIRILLOV, IVAN; OBUKHOV, DMITRIY; RYBAK, ALEXEY; DEDENIS, LAURENT; BELL, SERG; PROTASOV, STANISLAV
To: CONSTRUCTOR TECHNOLOGY AG
Reel/Frame 072261/0685 →
Continuity (1)
Related Publication 20240177283A1 · May 30, 2024
References Cited (35)
US 8937620B1 · Teller · 2015 [cited by examiner]
US 10592733B1 · Ramanarayanan et al. · 2020 [cited by applicant]
US 12167169B1 · Ganju · 2024 [cited by examiner]
US 20030028378A1 · August · 2003 [cited by examiner]
US 20060247046A1 · Choi · 2006 [cited by examiner]
US 20140358541A1 · Colibro · 2014 [cited by examiner]
US 20150036883A1 · Deri · 2015 [cited by examiner]
US 20180350144A1 · Rathod · 2018 [cited by examiner]
US 20220036617A1 · Biswas · 2022 [cited by examiner]
US 20230071994A1 · Day · 2023 [cited by examiner]
US 20230237722A1 · Stewart · 2023 [cited by examiner]
US 20230260183A1 · Anderson · 2023 [cited by examiner]
US 20230334904A1 · Monti · 2023 [cited by examiner]
US 20240104816A1 · Shaw · 2024 [cited by examiner]
US 20240202245A1 · Mahadevan · 2024 [cited by examiner]
US 20240233229A1 · Tumanov · 2024 [cited by examiner]
US 20240303947A1 · Domae · 2024 [cited by examiner]
US 20240420225A1 · Johri · 2024 [cited by examiner]
CN 112116589A · 2020 [cited by applicant]
CN 106250400B · 2021 [cited by applicant]
WO WO2021257868A1 · 2021 [cited by applicant]
WO WO2022087147A1 · 2022 [cited by applicant]
Minna Pakanen, Paula Alavesa, Niels Van Berkel, Timo Koskela, and Timo Ojala: “nice to see you virtually”: Thoughtful design and evaluation of virtual avatar of the other user in ar and vr based telexistence systems. En… [cited by applicant]
Wu, Y., Wang, Y., Jung, S., Hoermann, S., and Lindeman, R. W. (2021): Using a Fully Expressive Avatar to Collaborate in Virtual Reality: Evaluation of Task Performance, Presence, and Attraction. Front. Virtual Real. 2, … [cited by applicant]
Prajwal, K. R., et al.:“A lip sync expert is all you need for speech to lip generation in the wild.” Proceedings of the 28th ACM International Conference on Multimedia. 2020. [cited by applicant]
Video Multi-Method Assessment Fusion https://github.com/Netflix/vmaf. [cited by applicant]
Heusel, Martin, et al.:“Gans trained by a two time-scale update rule converge to a local nash equilibrium.” Advances in neural information processing systems 30, 2017. [cited by applicant]
Salimans, Tim, et al.:“Improved techniques for training gans.” Advances in neural information processing systems 29, 2016. [cited by applicant]
Zhang, Richard, et al.:“The unreasonable effectiveness of deep features as a perceptual metric.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2018. [cited by applicant]
IEEE Trans: “IEEE Recommended Practice for Speech Quality Measurements,”. Audio Electroacoust., vol. 17, No. 3, pp. 225-246, 1969, doi: 10.1109/TAU.1969.1162058. [cited by applicant]
M. Viswanathan and M. Viswanathan: “Measuring speech quality for text-to-speech systems: development and assessment of a modified mean opinion score (MOS) scale,” Comput. Speech Lang., vol. 19, No. 1, pp. 55-83, Jan. 20… [cited by applicant]
A. W. Rix, J. G. Beerends, M. P. Hollier, and A. P. Hekstra: “Perceptual evaluation of speech quality (PESQ)—a new method for speech quality assessment of telephone networks and codecs,” in 2001 IEEE International Confe… [cited by applicant]
C.-C. Lo et al.: “MOSNet: Deep Learning-based Objective Assessment for Voice Conversion.”. [cited by applicant]
Adrian Łancucki. Fastpitch: Parallel text-to-speech with pitch prediction. ' arXiv preprint arXiv:2006.06873, 2020. [cited by applicant]
J. Kim, J. Kong, and J. Son: “Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech,” arXiv preprint arXiv:2106.06103, 2021. [cited by applicant]