IP Library › Granted Patent US 12,277,637
Granted Patent B2
US 12,277,637 · App. 17/669,666 · Granted Apr 15, 2025

AI avatar-based interaction service method and apparatus

Inventors: Han Seok Ko (Seoul, KR); Jeong Min Bae (Seoul, KR); Miguel Alba (Seoul, KR)
Assignee: DATUM POINT LABS, INC.
G06T13/40G06N3/006G06V40/161G06V40/172G10L17/22H04R1/406H04R3/005H04R2201/401
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,277,637
App. No.
17/669,666
Granted
Apr 15, 2025
Kind
B2
Abstract

Artificial intelligence avatar-based interaction service is performed in a system including an unmanned information terminal and an interaction service device. A sound signal is collected from a microphone array mounted in the unmanned information terminal and an image signal collected from a vision sensor to the interaction service device. A sensing area is set based on the received sound signal and image signal by the interaction service device; recognizing an active speaker based on a voice signal of a user and an image signal of the user collected in the sensing area, by the interaction service device. A response for the recognized active speaker is generated to provide a 3D rendering an artificial intelligence avatar to which the response is reflected. The rendered artificial intelligence avatar is provided to the unmanned information terminal by the interaction service device.

Claims (46)

1. An artificial intelligence (AI) avatar-based interaction service method performed in a system including an unmanned information terminal and an interaction service device, the method comprising:

transmitting a sound signal collected from a microphone array mounted in the unmanned information terminal and an image signal collected from a vision sensor to the interaction service device;

setting a sensing area based on a received sound signal and image signal by the interaction service device;

recognizing an active speaker based on a voice signal of a user and an image signal of the user collected in the sensing area, by the interaction service device;

determining a voice information based on the voice signal in response to recognizing the active speaker;

determining a non-verbal information based on the image signal in response to recognizing the active speaker;

generating a response for the recognized active speaker, 3D rendering an artificial intelligence avatar, said artificial intelligence avatar reflecting a desired response, wherein generating the response includes applying a first weight to the voice information and a second weight to the non-verbal information in response to the voice information and the non-verbal information having consistent results, and applying a third weight different than the first weight to the voice information and a fourth weight different than the second weight to the non-verbal information in response to the voice information and the non-verbal information having inconsistent results; and

using the interaction service device to provide the rendered artificial intelligence avatar to the unmanned information terminal.

2. The artificial intelligence avatar-based interaction service method according to claim 1 , further comprising:

using the interaction service device to estimate a sound source direction based on the received sound signal by sound source direction estimation;

limiting an input of a sound from a side by sidelobe signal cancellation; and

limiting image input after an object recognized by applying background separation to the received image signal.

3. The artificial intelligence avatar-based interaction service method according to claim 1 , further comprising:

when recognizing the active speaker, using the interaction service device to determine a number of people from the image signal of the user in the sensing area by facial recognition; and

in the case of recognizing a plurality of people in the sensing area, selecting a person as a speaker as an active speaker using any one or more of sound source position estimation, voice recognition, and mouth-shape recognition.

4. The artificial intelligence avatar-based interaction service method according to claim 1 , further comprising:

determining the non-verbal information by analyzing information including at least one of the group consisting of a facial expression, a pose, a gesture, or a voice tone of a speaker from the received image or audio signal of the user.

5. The artificial intelligence avatar-based interaction service method according to claim 4 , further comprising:

determining the voice information by any one or more of voice recognition, natural language understanding, or speech-to-text.

6. The artificial intelligence avatar-based interaction service method according to claim 4 , further comprising, in the providing of the artificial intelligence avatar to the unmanned information terminal, analyzing information including at least one of the group consisting of a facial expression, a gesture, and a voice tone from the audio or image of the user to recognize an emotional state of the user to change an expression, a gesture, or a voice tone of the AI avatar in response to the recognized emotional state or add an effect.

7. The artificial intelligence avatar-based interaction service method according to claim 1 , wherein the third weight is lower than the first weight, and the fourth weight is lower than the second weight.

8. The artificial intelligence avatar-based interaction service method according to claim 1 , wherein the voice information and the non-verbal information have consistent results based on the voice information being positive, and the non-verbal information being positive.

9. The artificial intelligence avatar-based interaction service method according to claim 1 , wherein the voice information and the non-verbal information have consistent results based on the voice information being negative, and the non-verbal information being negative.

10. The artificial intelligence avatar-based interaction service method according to claim 1 , wherein the voice information and the non-verbal information have inconsistent results based on the voice information being positive, and the non-verbal information being negative.

11. An artificial intelligence (AI) avatar-based interaction service apparatus, comprising:

an unmanned information terminal which includes a microphone array and a vision sensor, configured to:

collect a sound signal from the microphone array, and

collect an image signal from the vision sensor; and

an interaction service device, configured to:

receive the sound signal and the image signal,

set a sensing area based on the sound signal and the image signal,

recognize an active speaker based on a voice signal of a user and an image signal of the user collected in the sensing area,

determine a voice information based on the voice signal in response to recognizing the active speaker,

determine a non-verbal information based on the image signal in response to recognizing the active speaker,

generate a response for the recognized active speaker, including applying a first weight to the voice information and a second weight to the non-verbal information in response to the voice information and the non-verbal information having consistent results, and applying a third weight different than the first weight to the voice information and a fourth weight different than the second weight to the non-verbal information in response to the voice information and the non-verbal information having inconsistent results,

3D render the artificial intelligence avatar that reflects the response, and

provide the rendered artificial intelligence avatar to the unmanned information terminal.

12. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the interaction service device estimates a sound source direction based on the received sound signal by sound source direction estimation, limits an input of a sound from a side by sidelobe signal cancellation, and limits image input after an object recognized by applying a background separation to the received image signal.

13. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the interaction service device checks a number of people from the image signal of the user in the sensing area by face recognition and in the case of recognizing a plurality of people in the sensing area, the interaction service device selects a person as an active speaker using any one or more of sound source position estimation, voice recognition, and mouth-shape recognition.

14. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the interaction service device is configured to determine the non-verbal information based on any one or more of a facial expression, a pose, a gesture, or a voice tone of a speaker from the received image signal of the user.

15. The artificial intelligence avatar-based interaction service apparatus according to claim 14 , wherein the interaction service device is configured to determine the voice information based on any one or more of voice recognition, natural language understanding, and speech-to-text.

16. The artificial intelligence avatar-based interaction service apparatus according to claim 14 , wherein the interaction service device analyzes a facial expression, a gesture, and a voice tone from the image of the user to recognize an emotional state of the user to change an expression, a gesture, or a voice tone of the AI avatar in response to the recognized emotional state or add an effect.

17. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the third weight is lower than the first weight, and the fourth weight is lower than the second weight.

18. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the voice information and the non-verbal information have consistent results based on the voice information being positive, and the non-verbal information being positive.

19. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the voice information and the non-verbal information have consistent results based on the voice information being negative, and the non-verbal information being negative.

20. The artificial intelligence avatar-based interaction service apparatus according to claim 11 , wherein the voice information and the non-verbal information have inconsistent results based on the voice information being positive, and the non-verbal information being negative.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2025
From: DM LAB CO., LTD.
To: BRAND ENGAGEMENT NETWORK, INC.
Reel/Frame 070104/0731 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 4, 2025
From: BRAND ENGAGEMENT NETWORK, INC.
To: DATUM POINT LABS, INC.
Reel/Frame 070106/0681 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 17, 2022
From: KO, HAN SEOK; BAE, JEONG MIN; ALBA, MIGUEL
To: DMLAB. CO., LTD.
Reel/Frame 059029/0841 →
Priority Claims (2)
KR 10-2021-0034756 · Mar 17, 2021 · national
KR 10-2022-0002347 · Jan 6, 2022 · national
Continuity (1)
Related Publication 20220301251A1 · Sep 22, 2022
References Cited (8)
US 10206036B1 · Feng · 2019 [cited by examiner]
US 20170068551A1 · Vadodaria · 2017 [cited by examiner]
US 20180124466A1 · Park · 2018 [cited by examiner]
US 20190246075A1 · Khadloya · 2019 [cited by examiner]
US 20200410980A1 · Yamada · 2020 [cited by examiner]
US 20210327449A1 · Shin · 2021 [cited by examiner]
Nunamaker, Jay F., et al. “Embodied conversational agent-based kiosk for automated interviewing.” Journal of Management Information Systems 28.1 (2011): 17-48. (Year: 2011). [cited by examiner]
Kwon, Dong-Soo, et al. “Emotion interaction system for a service robot.” RO-MAN 2007—The 16th IEEE International Symposium on Robot and Human Interactive Communication. IEEE, 2007. (Year: 2007). [cited by examiner]