IP Library Granted Patent US 11,600,290
Granted Patent B2
US 11,600,290 · App. 17/015,902 · Granted Mar 7, 2023

System and method for talking avatar

Inventor: Carl Adrian Woffenden (Bartenheim la Chaussée, FR)
Assignee: LEXIA LEARNING SYSTEMS LLC
G10L21/10G06T13/00G09B5/065G09B19/06G10L19/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,600,290
App. No.
17/015,902
Granted
Mar 7, 2023
Kind
B2
Abstract

Aspects of this disclosure provide techniques for generating a viseme and corresponding intensity pair. In some embodiments, the method includes generating, by a server, a viseme and corresponding intensity pair based at least on one of a clean vocal track or corresponding transcription. The method may include generating, by the server, a compressed audio file based at least on one of the viseme, the corresponding intensity, music, or visual offset. The method may further include generating, by the server or a client end application, a buffer of raw pulse-code modulated (PCM) data based on decoding at least a part of the compressed audio file, where the viseme is scheduled to align with a corresponding phoneme.

Claims (43)

1. A method for generating a viseme and corresponding intensity pair, comprising:

generating, by a server, a viseme and corresponding intensity pair based at least on one of a clean vocal track or corresponding transcription;

generating, by the server, a compressed audio file based at least on one of the viseme, the corresponding intensity, music, or visual offset; and

generating, by the server or a client end application, a buffer of raw pulse-code modulated (PCM) data based on decoding at least a part of the compressed audio file,

wherein the viseme is scheduled to align with a corresponding phoneme, and scheduled to coincide when a user hears a sound to allow two or more mouth shapes or facial expressions to be in an expected position using at least one visual offset, the at least one visual offset being used to compensate for longer blending between the two or more mouth shapes, and wherein the scheduling is based at least on one of:

a size of a decoder audio buffer for the compressed audio file,

a size of a processing buffer, or

a latency between transferring the decoder audio buffer to the client end application and the sound being heard.

2. The method of claim 1 further comprising storing the viseme and corresponding intensity pair in an intermediary file.

3. The method of claim 2 , wherein the viseme and corresponding intensity pair is stored in the intermediary file with specific timings for the viseme and corresponding intensity.

4. The method of claim 1 , wherein the frequency of decoding the compressed audio file is based on the size of the audio buffer.

5. The method of claim 1 further comprising feeding, by the server, the PCM data to a user equipment.

6. The method of claim 1 further comprising transmitting, by the server, the PCM data to a user equipment upon request.

7. The method of claim 1 further comprising scheduling the phoneme to align with at least one of a corresponding mouth shape or facial expression of an animated character.

8. A method for generating a viseme event and corresponding intensity pair, comprising:

generating, by a server, a viseme and corresponding intensity pair based at least on one of a clean vocal track or corresponding transcription;

generating, by the server, a compressed audio file based at least on one of the viseme, the corresponding intensity, music, or visual offset; and

inserting, by the server, a viseme generator based at least on one of a processing buffer or the compressed audio file,

wherein the viseme is scheduled to align with a corresponding phoneme, and scheduled to coincide when a user hears a sound to allow two or more mouth shapes or facial expressions to be in an expected position using at least one visual offset, the at least one visual offset being used to compensate for longer blending between the two or more mouth shapes, and wherein the scheduling is based at least on one of:

a size of a decoder audio buffer for the compressed audio file,

a size of the processing buffer, or

a latency between transferring the decoder audio buffer to a client end application and the sound being heard.

9. The method of claim 8 , wherein the scheduling of the viseme is based on the size of the processing buffer.

10. The method of claim 8 further comprising scheduling the phoneme to align with at least one of a corresponding mouth shape or facial expression of an animated character.

11. The method of claim 10 , wherein the corresponding mouth shape or facial expression of the animated character is created based on the viseme.

12. The method of claim 8 , wherein the frequency of decoding the compressed audio file is based on the size of the processing buffer.

13. The method of claim 8 further comprising feeding, by the server, the compressed audio file to a user equipment.

14. The method of claim 8 further comprising transmitting, by the server, the compressed audio file to a user equipment upon request.

15. A system for generating a viseme event and corresponding intensity pair, comprising:

a processor; and

a non-transitory computer readable storage medium storing programming for execution by the processor, the programming including instructions to:

generate a viseme and corresponding intensity pair based at least on one of a clean vocal track or corresponding transcription;

generate a compressed audio file based at least on one of the viseme, the corresponding intensity, music, or visual offset; and

generate a buffer of raw pulse-code modulated (PCM) data based on decoding at least a part of the compressed audio file,

wherein the viseme is scheduled to align with a corresponding phoneme, and scheduled to coincide when a user hears a sound to allow two or more mouth shapes or facial expressions to be in an expected position using at least one visual offset, the at least one visual offset being used to compensate for longer blending between the two or more mouth shapes, and wherein the scheduling is based at least on one of:

a size of a decoder audio buffer for the compressed audio file,

a size of a processing buffer, or

a latency between transferring the decoder audio buffer to a client end application and the sound being heard.

16. The system of claim 15 , wherein the programming includes further instructions to store the viseme and corresponding intensity pair in an intermediary file.

17. The system of claim 16 , wherein the viseme and corresponding intensity pair is stored in the intermediary file with specific timings for the viseme and intensity.

18. The system of claim 15 , wherein the frequency of decoding the compressed audio file is based on the size of the audio buffer.

19. The system of claim 15 , wherein the programming includes further instructions to feed the PCM data to a user equipment.

20. The system of claim 15 , wherein the programming includes further instructions to transmit the PCM data to a user equipment upon request.

Assignments (8)
RELEASE OF SECURITY INTEREST Recorded Jul 20, 2021
From: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
To: CAMBIUM ASSESSMENT, INC.; KURZWEIL EDUCATION, INC.; LAZEL, INC.; VOYAGER SOPRIS LEARNING, INC.; VKIDZ HOLDINGS INC.; LEXIA LEARNING SYSTEMS LLC; CAMBIUM LEARNING, INC.
Reel/Frame 056921/0286 →
SECURITY INTEREST Recorded Jul 20, 2021
From: CAMBIUM ASSESSMENT, INC.; KURZWEIL EDUCATION, INC.; LAZEL, INC.; VOYAGER SOPRIS LEARNING, INC.; VKIDZ HOLDINGS, INC.; LEXIA LEARNING SYSTEMS LLC; CAMBIUM LEARNING, INC.
To: WILMINGTON TRUST, NATIONAL ASSOCIATION, AS COLLATERAL AGENT
Reel/Frame 056921/0745 →
RELEASE OF SECURITY INTEREST IN PATENTS AT REEL/FRAME NO. 54085/0934 Recorded Mar 12, 2021
From: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
To: ROSETTA STONE LTD.
Reel/Frame 055583/0562 →
RELEASE OF SECURITY INTEREST IN PATENTS AT REEL/FRAME NO. 54085/0920 Recorded Mar 12, 2021
From: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
To: ROSETTA STONE LTD.
Reel/Frame 055583/0555 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 17, 2020
From: ROSETTA STONE, LTD.
To: LEXIA LEARNING SYSTEMS LLC
Reel/Frame 054676/0460 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Oct 15, 2020
From: ROSETTA STONE LTD.; LEXIA LEARNING SYSTEMS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 054085/0934 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Oct 15, 2020
From: ROSETTA STONE LTD.; LEXIA LEARNING SYSTEMS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 054085/0920 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2020
From: WOFFENDEN, CARL ADRIAN
To: ROSETTA STONE, LTD.
Reel/Frame 053859/0953 →