IP Library › Granted Patent US 12,724,804
Granted Patent B1
US 12,724,804 · App. 19/442,919 · Granted Sep 1, 2026

Architecture for artificial intelligence avatar deployment across hybrid cloud infrastructure

Inventors: Vasanthakumar Rajendran (New York, NY); Jefferson Okraku (New York, NY); Deepak Kela (New York, NY); Rachit Kumar (New York, NY); Satchel Aviram (New York, NY); Karolina Belwal (New York, NY); Tarak Mehta (New York, NY); Ranjit Kumar Angiya Rameshbabu (Chennai, IN); Robin Jain (Pune, IN); Joseph V. Bonanno, Jr. (Scarsdale, NY); Dipendra Malhotra (Cranbury, NJ); Aravind Pitchai Guruswamy (Whitehouse Station, NJ)
G06F16/33295H04L65/1069
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,804
App. No.
19/442,919
Granted
Sep 1, 2026
Kind
B1
Abstract

The systems and methods disclosed herein provide an architecture for deploying artificial intelligence (AI) avatars that generate audio and video responses to user interactions across various infrastructure environments. The systems and methods disclosed herein establish connections between client applications and a first infrastructure layer that invokes multiple agents trained on domain-specific knowledge bases to retrieve associated data records. The retrieved data and user input are routed to a second infrastructure layer hosting a generative AI model that identifies relevant data fields and generates avatar responses as audio or video streams. The systems and methods disclosed herein are enabled to evaluate the user inputs and/or avatar responses by applying one or more validation criteria. The systems and methods disclosed herein adjust the user inputs and/or avatar responses based on the evaluation. Context data is maintained across communication sessions and multiple client devices to preserve conversational continuity

Claims (96)

1 . A system comprising:

at least one hardware processor; and

at least one non-transitory memory storing instructions, which, when executed by the at least one hardware processor, cause the system to:

receive, at a client application executing on a client device, a user interaction request to initiate a communication session with an artificial intelligence avatar,

wherein the artificial intelligence avatar is configured to present an audiovisual signal responsive to an input received by the client application;

initiate the communication session by establishing a connection between the client application and a first infrastructure layer operating within a first network perimeter,

wherein the connection is configured to bidirectionally transmit data during the communication session;

receive, from the client application through the connection to the first infrastructure layer, user input data comprising at least one of natural language text, audio data, or video data captured during the communication session;

invoke, within the first infrastructure layer, a plurality of agents each configured to execute a command set to retrieve a data record set corresponding to the user interaction request,

wherein the plurality of agents are each trained using different domain-specific knowledge bases to execute a respective command set;

route (a) the user input data and (b) the data record set from the first infrastructure layer operating within a first network perimeter controlled by a first entity to a second infrastructure layer operating within a second network perimeter controlled by a second entity different from the first entity;

cause generation of an avatar response of the artificial intelligence avatar using a generative artificial intelligence model hosted by the second infrastructure layer, wherein the generative artificial intelligence model is configured to:

identify, within the data record set, one or more data fields by comparing a vector representation of the user input data and a vector representation of each data field in the data record set,

map the one or more data fields to one or more corresponding data sources, and

use (a) the one or more data fields and (b) a representation of the one or more corresponding data sources to generate an audio stream and a corresponding video stream within the avatar response,

wherein the audio stream and the corresponding video stream are responsive to the user input data; and

cause transmission of the avatar response from the second infrastructure layer, through the first infrastructure layer and the connection, to the client application for presentation of the avatar response on the client device during the communication session.

2 . The system of claim 1 ,

wherein the user interaction request comprises a service inquiry from a user account, and

wherein the data record set comprises at least one of: account information, transaction history, or product data associated with the user account.

3 . The system of claim 2 , wherein the system is further caused to:

classify the service inquiry as a servicing request or an advice request,

wherein the servicing request comprises a request to execute a transaction in association with the user account, and

wherein the advice request comprises a request for a recommendation generated based on the data record set associated with the user account.

4 . The system of claim 1 , wherein the system is further caused to:

validate the user input data prior to routing the user input data to the second infrastructure layer by applying one or more criteria to remove prohibited content within the user input data,

wherein the one or more criteria are associated with one or more patterns indicative of the prohibited content within the user input data.

5 . The system of claim 1 , wherein the system is further caused to:

validate the avatar response prior to causing transmission of the avatar response to the client application by applying one or more criteria to remove prohibited content within the avatar response,

wherein the one or more criteria are associated with one or more patterns indicative of the prohibited content within the avatar response.

6 . The system of claim 5 ,

wherein the one or more criteria comprise a classification rule configured to determine whether the avatar response includes advice or a transaction request, and

wherein, in response to a determination that the avatar response includes the advice, the system is further caused to modify the avatar response to remove the advice.

7 . The system of claim 1 , wherein the system is further caused to:

apply one or more validation rules to at least one of the user input data or the avatar response to generate a validation result; and

modify the avatar response using the validation result,

wherein modifying the avatar response comprises at least one of: adjusting a conversation topic within the avatar response, requesting additional authentication information, or adding a referral associated with a particular agent within the avatar response.

8 . A non-transitory computer-readable storage medium comprising instructions for deploying an artificial intelligence avatar stored thereon, wherein the instructions when executed by at least one data processor of a system, cause the system to:

obtain, at a client application executing on a client device, a user interaction request to initiate a communication session with an artificial intelligence avatar,

wherein the artificial intelligence avatar is configured to present an audiovisual signal responsive to an input obtained by the client application;

cause initiation of the communication session by establishing a connection between the client application and a first infrastructure layer operating within a first network perimeter controlled by a first entity;

obtain, from the client application through the connection to the first infrastructure layer, user input data comprising at least one of natural language text, audio data, or video data captured during the communication session;

invoke, within the first infrastructure layer, one or more agents each configured to execute a command set to retrieve a data record set corresponding to the user input data;

cause generation of an avatar response of the artificial intelligence avatar using a generative artificial intelligence model hosted by a second infrastructure layer operating within a second network perimeter controlled by a second entity different from the first entity,

wherein the generative artificial intelligence model is configured to use the data record set to generate the avatar response, and

wherein the avatar response is responsive to the user input data; and

cause transmission of the avatar response from the second infrastructure layer, through the first infrastructure layer and the connection, to the client application for presentation of the avatar response on the client device during the communication session.

9 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the system to:

store a record of the avatar response on a distributed ledger,

wherein the distributed ledger comprises at least one of a blockchain or a federated ledger.

10 . The non-transitory computer-readable storage medium of claim 8 ,

wherein the first infrastructure layer operates within a first network perimeter controlled by a first entity, and

wherein the second infrastructure layer operates within a second network perimeter controlled by a second entity different from the first entity.

11 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the system to:

invoke a guardrails service hosted within the first infrastructure layer,

wherein the guardrails service is configured to apply one or more criteria to at least one of the user input data or the avatar response.

12 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the system to:

invoke a guardrails service hosted within the second infrastructure layer,

wherein the guardrails service is configured to apply one or more criteria to at least one of the user input data or the avatar response.

13 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the system to:

invoke a first guardrails service hosted within the first infrastructure layer,

wherein the first guardrails service is configured to trigger one or more function calls to a second guardrails service external to the first infrastructure layer, and

wherein the second guardrails service is configured to apply one or more criteria to at least one of the user input data or the avatar response.

14 . The non-transitory computer-readable storage medium of claim 8 , wherein the instructions further cause the system to:

classify the user input data into a category;

identify a subset of validation rules from a plurality of validation rules based on the category;

apply the subset of validation rules to the avatar response to generate a validation result; and

determine, based on the validation result, whether to transmit the avatar response to the client application or generate a modified avatar response.

15 . A computer-implemented method for deploying an artificial intelligence avatar, the computer-implemented method comprising:

obtaining, at a computer-implemented application, a user interaction request to initiate a communication session with the artificial intelligence avatar configured to present an output responsive to an input obtained by the computer-implemented application;

causing initiation of the communication session by establishing a connection between the computer-implemented application and a first infrastructure layer operating within a first network perimeter controlled by a first entity;

obtaining, from the computer-implemented application through the connection to the first infrastructure layer, input data obtained during the communication session;

causing invocation of, within the first infrastructure layer, one or more agents each configured to execute a command set to retrieve a data record set corresponding to the input data;

causing generation of an avatar response of the artificial intelligence avatar using a generative artificial intelligence model hosted by a second infrastructure layer operating within a second network perimeter controlled by a second entity different from the first entity,

wherein the generative artificial intelligence model is configured to use the data record set to generate the avatar response; and

causing transmission of the avatar response from the second infrastructure layer, through the first infrastructure layer and the connection, to the computer-implemented application for presentation of the avatar response during the communication session.

16 . The computer-implemented method of claim 15 , further comprising:

evaluating the input data to determine one or more subsequent interaction requests to occur within a predetermined number of conversational turns; and

invoking one or more additional agents configured to retrieve supplemental data records corresponding to the one or more subsequent interaction requests.

17 . The computer-implemented method of claim 15 , further comprising:

storing conversation data including one or more of the input data or the avatar response from the communication session in a memory database,

wherein the conversation data is accessible across multiple client devices.

18 . The computer-implemented method of claim 15 , further comprising:

classify conversation data as short-term memory data or long-term memory data based on temporal data associated with the conversation data,

wherein the conversation data includes one or more input data and one or more avatar responses across multiple communication sessions,

storing the short-term memory data in a session memory that is accessible during the communication session; and

storing the long-term memory data in a persistent memory database that is accessible across the multiple communication sessions.

19 . The computer-implemented method of claim 15 , wherein the computer-implemented application is hosted on a first client device, further comprising:

receiving a request to access the communication session from a second client device different from the first client device;

transmitting an authentication test to the second client device;

receiving authentication information from the second client device in response to the authentication test;

validating the authentication information against stored user information associated with the first client device; and

in response to satisfaction of the authentication information with the stored user information, establishing a connection from the second client device to the communication session.

20 . The computer-implemented method of claim 15 , further comprising:

detecting a presence of an unauthorized user during the communication session; and

responsive to detecting the presence of the unauthorized user, modifying the avatar response to remove one or more indicators of the data record set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 8, 2026
From: RAJENDRAN, VASANTHAKUMAR; OKRAKU, JEFFERSON; KELA, DEEPAK; KUMAR, RACHIT; AVIRAM, SATCHEL; BELWAL, KAROLINA; MEHTA, TARAK; ANGIYA RAMESHBABU, RANJIT KUMAR; JAIN, ROBIN; BONANNO, JOSEPH V., JR.; MALHOTRA, DIPENDRA; GURUSWAMY, ARAVIND PITCHAI
To: CITIBANK, N.A.
Reel/Frame 074314/0270 →
References Cited (9)
US 12555008B1 · Singh · 2026 [cited by examiner]
US 20250190460A1 · Madisetti · 2025 [cited by examiner]
US 20250209326A1 · Madisetti · 2025 [cited by examiner]
US 20250321992A1 · Madisetti · 2025 [cited by examiner]
US 20250328560A1 · Madisetti · 2025 [cited by examiner]
US 20250330677A1 · Chomal et al. · 2025 [cited by applicant]
US 20260017386A1 · Ohayon · 2026 [cited by examiner]
Schick, Timo, “Toolformer: Language models can teach themselves to use tools”, Advances in neural information processing systems 36 (2023): 68539-68551. (Year: 2023), 17 pages. [cited by applicant]
Wang, Yuxuan, “Style V Tokens: Unsupervised Style Modeling, Control and Transfer in End-to-End Speech Synthesis.”, (2018) Proceedings of the 35th International Conference on Machine Learning (Year: 2018), 10 pages. [cited by applicant]