IP Library Patent Application 18089426
Patent Application
App. No. 18/089,426

SYSTEM AND METHOD FOR GENERATING AVATAR OF AN ACTIVE SPEAKER IN A MEETING

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/089,426
Abstract

A method includes receiving a facial data associated with a participant user; generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model; receiving data associated with facial movement; and training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.

Claims (43)

1 . A computer-implemented method comprising:

receiving a facial data associated with a participant user;

generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model;

receiving data associated with facial movement; and

training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.

2 . The computer-implemented method of claim 1 further comprising receiving an audio data associated with the participant user during a meeting session without a video stream of the participant user.

3 . The computer-implemented method of claim 2 further comprising:

identifying the participant user when the user becomes an active speaker at the meeting;

retrieving the trained avatar data associated with the participant user; and

generating an avatar video stream of the participant user based on the trained avatar data mimicking the received audio data associated with the participant user.

4 . The computer-implemented method of claim 3 further comprising:

injecting the generated avatar video stream to data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.

5 . The computer-implemented method of claim 4 , wherein the injecting complements data being transmitted during the meeting session without video stream of the participant user.

6 . The computer-implemented method of claim 3 , wherein the identifying the participant user is through voice recognition processing of the audio data associated with the participant user.

7 . The computer-implemented method of claim 1 , wherein data associated with facial movement comprises video streams of one or more users speaking and wherein the one or more users are different from the participant user.

8 . The computer-implemented method of claim 1 , wherein the facial data associated with the participant user comprises at least one static image of the participant user from an application that is different from an application that facilitates an online meeting for the participant user.

9 . The computer-implemented method of claim 1 , wherein the facial data associated with the participant user comprises at least one static image of the participant user from an that facilitates an online meeting for the participant user.

10 . The computer-implemented method of claim 1 , wherein the facial data associated with the participant user comprises at least a portion of a video stream.

11 . A computer-implemented method comprising:

receiving an audio data associated with a participant user during a meeting session without video stream of the participant user;

receiving a trained avatar data, wherein an avatar data is generated based on facial data associated with the participant user using a first machine learning (ML) model and wherein the avatar data is trained using a second ML model based on data associated with facial movement to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movement associated with audio data for the avatar data; and

generating an avatar video stream of the participant user based on the trained avatar data mimicking facial movements associated with the received audio data associated with the participant user.

12 . The computer-implemented method of claim 11 further comprising:

identifying the participant user when the user becomes an active speaker at the meeting.

13 . The computer-implemented method of claim 12 further comprising retrieving the trained avatar data associated with the participant user based on the identifying the participant user.

14 . The computer-implemented method of claim 12 , wherein the identifying the participant user is through voice recognition processing of the audio data associated with the participant user.

15 . The computer-implemented method of claim 11 further comprising:

injecting the generated avatar video stream to data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.

16 . The computer-implemented method of claim 15 , wherein the injecting complements data being transmitted during the meeting session without video stream of the participant user.

17 . A system, comprising:

a processor;

a memory operatively connected to the processor and storing instructions that, when executed by the processor, cause:

receiving a facial data associated with a participant user;

generating an avatar data of the participant user based on the facial data using a first machine learning (ML) model;

receiving data associated with facial movement; and

training the generated avatar based on the data associated with facial movement using a second ML model to generate a trained avatar data, wherein the trained avatar data mimics appropriate facial movements associated with audio data.

18 . The system of claim 17 , wherein the instructions when executed by the process further cause receiving an audio data associated with the participant user during a meeting session without video stream of the participant user.

19 . The system of claim 18 , wherein the instructions when executed by the process further cause:

identifying the participant user when the user becomes an active speaker at the meeting;

retrieving the trained avatar data associated with the participant user; and

generating an avatar video stream of the participant user based on the trained avatar data mimicking the received audio data associated with the participant user.

20 . The system of claim 19 , wherein the instructions when executed by the process further cause:

injecting the generated avatar video stream to data being transmitted to participants of the meeting when the participant user is speaking as the active speaker.

Assignments (2)
SECURITY INTEREST Recorded Feb 14, 2023
From: RINGCENTRAL, INC.
To: BANK OF AMERICA, N.A., AS COLLATERAL AGENT
Reel/Frame 062973/0194 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2022
From: VENDROW, VLAD; KAMYSHEV, IGOR VLADIMIROVICH
To: RINGCENTRAL, INC.
Reel/Frame 062215/0296 →