IP Library Granted Patent US 12,450,809
Granted Patent B2
US 12,450,809 · App. 18/098,428 · Granted Oct 21, 2025

Method and apparatus for providing interactive avatar services

Inventors: Jaeeun Yang (Suwon-si, KR); Jaehong Kim (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T13/40G06T13/205G06T17/20G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,809
App. No.
18/098,428
Granted
Oct 21, 2025
Kind
B2
Abstract

A method of providing an avatar service includes obtaining a user-uttered voice and a spatial information of a user-utterance space, transmitting the user-uttered voice and the spatial information to a server, receiving, from the server, a first avatar voice answer and an avatar facial expression sequence corresponding to the first avatar voice, which are determined based on the user-uttered voice and the spatial information, determining first avatar facial expression data, based on the first avatar voice answer and the avatar facial expression sequence, identifying a certain event during reproduction of a first avatar animation created based on the first avatar voice answer and the first avatar facial expression data, determining second avatar facial expression data or a second avatar voice answer, based on the certain event, and reproducing a second avatar animation created based on the second avatar facial expression data or the second avatar voice answer.

Claims (55)

1. A method, performed by an electronic device, of providing an avatar service, comprising:

obtaining a user-uttered voice and a spatial information of a user-utterance space where a user utters the user-uttered voice, the spatial information comprising spatial characteristics of the user-utterance space based on at least one of images captured by a camera and a sound obtained through a microphone, wherein the spatial information comprises at least one of: whether the user-utterance space is public or private; or whether the user-utterance space is quiet or noisy;

transmitting the user-uttered voice and the spatial information to a server;

receiving, from the server, a first avatar voice answer and an avatar facial expression sequence corresponding to the first avatar voice answer, which are determined based on the user-uttered voice and an avatar response mode, wherein the avatar response mode is determined based on the spatial information;

determining first avatar facial expression data, based on the first avatar voice answer and the avatar facial expression sequence;

identifying a certain event during reproduction of a first avatar animation created based on the first avatar voice answer and the first avatar facial expression data;

determining second avatar facial expression data or a second avatar voice answer, based on the certain event; and

stopping reproduction of the first avatar animation, and reproducing a second avatar animation created based on the second avatar facial expression data or the second avatar voice answer.

2. The method of claim 1 , wherein the spatial information comprises information about whether the user-utterance space is a public place and a level of noise in the user-utterance space.

3. The method of claim 1 , wherein the first avatar facial expression data and the second avatar facial expression data each comprise a set of coefficients for each of a plurality of reference three-dimensional (3D) meshes for modeling a facial expression of the first avatar animation and the second avatar animation, respectively.

4. The method of claim 1 , wherein:

the second avatar facial expression data comprises lip sync data, and

the lip sync data is obtained using an artificial intelligence (AI) model.

5. The method of claim 4 , wherein the AI model is trained using data normalized based on an available range according to the lip sync data.

6. The method of claim 1 , wherein the certain event comprises at least one of an utterance mode change event, an observation mode event, or a refresh mode event.

7. The method of claim 6 , wherein, based on the certain event being the refresh mode event, the stopping reproduction of the first avatar animation, and the reproducing of the second avatar animation comprises:

stopping the reproduction of the first avatar animation at a point in time;

reproducing a preset refresh animation; and

reproducing the first avatar animation from the point in time at which the first avatar animation is stopped.

8. The method of claim 6 , wherein, based on the certain event being the utterance mode change event, the determining of the second avatar facial expression data or the second avatar voice answer comprises:

determining the second avatar facial expression data by modifying the first avatar facial expression data, based on an utterance mode obtained as a result of the certain event; and

modifying the first avatar voice answer, based on the utterance mode.

9. The method of claim 6 , wherein, based on the certain event being the observation mode event, the determining of the second avatar facial expression data or the second avatar voice answer comprises determining the second avatar facial expression data by changing a face direction or eye direction of the first avatar animation.

10. A method, performed by a server, of providing an avatar service through an electronic device, the method comprising:

receiving, from the electronic device, a user-uttered voice and spatial information of a user-utterance space where a user utters the user-uttered voice, the spatial information comprising spatial characteristics of the user-utterance space based on at least one of images captured by a camera and a sound obtained through a microphone, wherein the spatial information comprises at least one of: whether the user-utterance space is public or private; or whether the user-utterance space is quiet or noisy;

determining an avatar response mode for the user-uttered voice, based on the spatial information;

generating a first avatar voice answer for an avatar to respond to the user-uttered voice and an avatar facial expression sequence corresponding to the first avatar voice answer, based on the user-uttered voice and the avatar response mode; and

transmitting the first avatar voice answer and the avatar facial expression sequence for generating a first avatar animation to the electronic device.

11. An electronic device for providing an avatar service, comprising:

a communication interface;

a storage storing at least one instruction; and

at least one processor configured to execute the at least one instruction stored in the storage, wherein the at least one processor is configured to execute the at least one instruction to:

obtain a user-uttered voice and a spatial information of a user-utterance space where a user utters the user-uttered voice, the spatial information comprising spatial characteristics of the user-utterance space based on at least one of images captured by a camera and a sound obtained through a microphone, wherein the spatial information comprises at least one of: whether the user-utterance space is public or private; or whether the user-utterance space is quiet or noisy;

transmit the user-uttered voice and the spatial information to a server;

receive, from the server through the communication interface, a first avatar voice answer and an avatar facial expression sequence corresponding to the first avatar voice answer, which are determined based on the user-uttered voice and an avatar response mode, wherein the avatar response mode is determined based on the spatial information;

determine first avatar facial expression data, based on the first avatar voice answer and the avatar facial expression sequence;

identify a certain event during reproduction of a first avatar animation created based on the first avatar voice answer and the first avatar facial expression data;

determine second avatar facial expression data or a second avatar voice answer, based on the certain event; and

reproduce a second avatar animation created based on the second avatar facial expression data or the second avatar voice answer.

12. The electronic device of claim 11 , wherein the spatial information comprises information about whether the user-utterance space is a public place and a level of noise in the user-utterance space.

13. The electronic device of claim 11 , wherein the first avatar facial expression data and the second avatar facial expression data each comprise a set of coefficients for each of a plurality of reference three-dimensional (3D) meshes for modeling a facial expression of the first avatar animation and the second avatar animation, respectively.

14. The electronic device of claim 11 , wherein:

the second avatar facial expression data comprises lip sync data, and

the lip sync data is obtained using an artificial intelligence (AI) model.

15. The electronic device of claim 14 , wherein the AI model is trained using data normalized based on an available range according to the lip sync data.

16. The electronic device of claim 11 , wherein the certain event comprises at least one of an utterance mode change event, an observation mode event, or a refresh mode event.

17. The electronic device of claim 16 , wherein, when based on the certain event being the refresh mode event, the second avatar animation is a preset refresh animation.

18. A non-transitory computer-readable recording medium for storing computer readable program code or instructions which are executable by a processor to perform a method of providing an avatar service, the method comprising:

obtaining a user-uttered voice and a spatial information of a user-utterance space where a user utters the user-uttered voice, the spatial information comprising spatial characteristics of the user-utterance space based on at least one of images captured by a camera and a sound obtained through a microphone, wherein the spatial information comprises at least one of: whether the user-utterance space is public or private; or whether the user-utterance space is quiet or noisy;

transmitting the user-uttered voice and the spatial information to a server;

receiving, from the server, a first avatar voice answer and an avatar facial expression sequence corresponding to the first avatar voice answer, which are determined based on the user-uttered voice and an avatar response mode, wherein the avatar response mode is determined based on the spatial information;

determining first avatar facial expression data, based on the first avatar voice answer and the avatar facial expression sequence;

identifying a certain event during reproduction of a first avatar animation created based on the first avatar voice answer and the first avatar facial expression data;

determining second avatar facial expression data or a second avatar voice answer, based on the certain event; and

stopping reproduction of the first avatar animation, and reproducing a second avatar animation created based on the second avatar facial expression data or the second avatar voice answer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2023
From: YANG, JAEEUN; KIM, JAEHONG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062413/0474 →
Priority Claims (1)
KR 10-2022-0007400 · Jan 18, 2022 · national
Continuity (2)
Continuation PCTKR2023000721 · Jan 16, 2023
Related Publication 20230230303A1 · Jul 20, 2023
References Cited (47)
US 6753857B1 · Matsuura · 2004 [cited by examiner]
US 8941642B2 · Tadaishi · 2015 [cited by examiner]
US 9330483B2 · Du · 2016 [cited by examiner]
US 9799133B2 · Tong · 2017 [cited by examiner]
US 10293260B1 · Evans et al. · 2019 [cited by applicant]
US 10521946B1 · Roche et al. · 2019 [cited by applicant]
US 10586369B1 · Roche · 2020 [cited by examiner]
US 10726836B2 · Baik et al. · 2020 [cited by applicant]
US 11562520B2 · Lee · 2023 [cited by examiner]
US 11593984B2 · Hussen Abdelaziz · 2023 [cited by examiner]
US 11651541B2 · Baszucki · 2023 [cited by examiner]
US 20050253850A1 · Kang · 2005 [cited by examiner]
US 20120130717A1 · Xu et al. · 2012 [cited by applicant]
US 20120310717A1 · Kankainen et al. · 2012 [cited by applicant]
US 20150213604A1 · Li · 2015 [cited by examiner]
US 20160005206A1 · Li · 2016 [cited by examiner]
US 20170206694A1 · Jiao · 2017 [cited by examiner]
US 20170206797A1 · Solomon · 2017 [cited by examiner]
US 20170256086A1 · Park · 2017 [cited by examiner]
US 20180253897A1 · Satake · 2018 [cited by examiner]
US 20180316734A1 · Nakabo · 2018 [cited by examiner]
US 20180335930A1 · Scapel · 2018 [cited by examiner]
US 20180336713A1 · Avendano · 2018 [cited by examiner]
US 20180373413A1 · Sawaki · 2018 [cited by examiner]
US 20190138266A1 · Takechi et al. · 2019 [cited by applicant]
US 20190143527A1 · Favis et al. · 2019 [cited by applicant]
US 20190250934A1 · Kim · 2019 [cited by examiner]
US 20190340419A1 · Milman · 2019 [cited by examiner]
US 20200410739A1 · Shin et al. · 2020 [cited by applicant]
US 20210027511A1 · Shang · 2021 [cited by examiner]
US 20210056747A1 · Hefny · 2021 [cited by examiner]
US 20220241692A1 · Fukushige · 2022 [cited by examiner]
US 20220301250A1 · Ko · 2022 [cited by examiner]
CN 116762103A · 2023 [cited by examiner]
KR 1020110059178A · 2011 [cited by applicant]
KR 1020110081364A · 2011 [cited by applicant]
KR 101089184A · 2011 [cited by applicant]
KR 1020180084582A · 2018 [cited by applicant]
KR 1020190046371A · 2019 [cited by applicant]
KR 101992424B1 · 2019 [cited by applicant]
KR 102108422A · 2020 [cited by applicant]
KR 1020210060196A · 2021 [cited by applicant]
WO 2020159621A1 · 2020 [cited by applicant]
International Search Report (PCT/ISA/210) issued May 1, 2023 from the International Searching Authority in International Application No. PCT/KR2023/000721. [cited by applicant]
“DeepSpeech Model”, Mozilla DeepSpeech, https://deepspeech.readthedocs.io/en/master/DeepSpeech.html, 2020, (3 pages total). [cited by applicant]
D. Cudeiro et al., “VOCA Coice Operated Character Animation”, Computer Vision and Pattern Recognition (CVPR), http://voca.is.tue.mpg.de/, 2019, (4 pages total). [cited by applicant]
Communication dated Feb. 7, 2025, issued by European Patent Office in European Patent Application No. 23743425.3. [cited by applicant]