IP Library › Granted Patent US 12,748,776
Granted Patent B1
US 12,748,776 · App. 19/540,529 · Granted Sep 29, 2026

Active-active mirrored artificial intelligence architecture

Inventors: Ganesh Prasad Bhat (West Orange, NJ); James Randolph Myers (Clearwater Beach, FL)
Assignee: Citibank, N.A.
G06F16/273G06F16/907
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,776
App. No.
19/540,529
Granted
Sep 29, 2026
Kind
B1
Abstract

Systems and methods for an active-active mirrored artificial intelligence architecture are disclosed. In some implementations, the system identifies, for an inference request from a user, at least two geographically separate compute clusters configured to handle the inference request, each compute cluster storing a version of an artificial intelligence model for processing the inference request. The system determines, from the inference request, a data consistency requirement mapping the inference request to one or more data consistency types. The system determines, for each compute cluster, a composite metric based on, among other things, a consistency freshness score of the compute cluster. The system routes the inference request to a selected compute cluster having a lowest composite metric subject to a constraint that the consistency freshness score for the compute cluster satisfies the data consistency requirement for the inference request. The system returns an inference response for the inference request to the user.

Claims (32)

1 . One or more non-transitory computer-readable storage media storing instructions that, when executed by one or more processors, cause the one or more processors to perform operations that respond to an inference request, the operations comprising:

identifying, for an inference request from a user, at least two compute clusters configured to handle the inference request;

determining, from the inference request, a data consistency requirement that maps the inference request to one or more data consistency types;

routing the inference request to a compute cluster based at least in part on a constraint that a consistency freshness score for the compute cluster satisfies the data consistency requirement for the inference request; and

returning an inference response for the inference request to the user.

2 . The one or more non-transitory computer-readable storage media of claim 1 , wherein the consistency freshness score of the compute cluster for conversation state is computed as a vector clock divergence value.

3 . The one or more non-transitory computer-readable storage media of claim 1 , wherein one or more analytics records generated by the inference request are tagged for asynchronous replication.

4 . The one or more non-transitory computer-readable storage media of claim 1 , wherein a session affinity token accompanies the inference request.

5 . The one or more non-transitory computer-readable storage media of claim 1 , wherein routing the inference request further comprises computing a layer-wise delta between an updated model checkpoint and its previously replicated state and transmitting only those model weights that differ.

6 . The one or more non-transitory computer-readable storage media of claim 5 , wherein the layer-wise delta is singular value decomposition (SVD)-compressed and INT8-quantized before transmission, thereby shortening replication time for subsequent requests.

7 . The one or more non-transitory computer-readable storage media of claim 1 , further storing instructions that replicate model weight and/or configuration update triggered while routing the inference request, the replication being enforced under (i) strong consistency for model weight data and (ii) eventual or causal consistency for non-critical analytics data generated by the inference request.

8 . The one or more non-transitory computer-readable storage media of claim 1 , further storing instructions that capture a conversation state of the inference request and replicate the conversation state to another compute cluster in parallel with the inference to enable a seamless mid-conversation failover.

9 . The one or more non-transitory computer-readable storage media of claim 1 , further storing instructions that, for requests originating from autonomous AI agents, map each agent to a backup compute cluster and pre-stage agent state so that the inference request can be re-issued from the backup compute cluster without data loss.

10 . The one or more non-transitory computer-readable storage media of claim 1 , wherein every model weight replicated during fulfillment of the inference request is cryptographically signed and appended to a distributed ledger, providing a tamper-evident audit trail for the inference request.

11 . A method for responding to an inference request, the method comprising:

identifying, for an inference request from a user, at least two compute clusters configured to handle the inference request;

determining, from the inference request, a data consistency requirement that maps the inference request to one or more data consistency types;

routing the inference request to a compute cluster based at least in part on a constraint that a consistency freshness score for the compute cluster satisfies the data consistency requirement for the inference request; and

returning an inference response for the inference request to the user.

12 . The method of claim 11 , wherein the consistency freshness score of the compute cluster for conversation state is computed as a vector clock divergence value.

13 . The method of claim 11 , wherein one or more analytics records generated by the inference request are tagged for asynchronous replication.

14 . The method of claim 11 , wherein a session affinity token accompanies the inference request.

15 . The method of claim 11 , wherein routing the inference request further comprises computing a layer-wise delta between an updated model checkpoint and its previously replicated state and transmitting only those model weights that differ.

16 . The method of claim 15 , wherein the layer-wise delta is singular value decomposition (SVD)-compressed and INT8-quantized before transmission, thereby shortening replication time for subsequent requests.

17 . The method of claim 11 , further comprising replicating model weight and/or configuration update triggered while routing the inference request, the replication being enforced under (i) strong consistency for model weight data and (ii) eventual or causal consistency for non-critical analytics data generated by the inference request.

18 . The method of claim 11 , further comprising capturing a conversation state of the inference request and replicating the conversation state to another compute cluster in parallel with the inference to enable a seamless mid-conversation failover.

19 . The method of claim 11 , further comprising, for requests originating from autonomous AI agents, mapping each agent to a backup compute cluster and pre-staging agent state so that the inference request can be re-issued from the backup compute cluster without data loss.

20 . A system comprising at least one processor and one or more non-transitory computer-readable media having instructions stored thereon that, when executed by the at least one processor, cause the at least one processor to perform operations comprising:

identifying, for an inference request from a user, at least two compute clusters configured to handle the inference request;

determining, from the inference request, a data consistency requirement that maps the inference request to one or more data consistency types;

routing the inference request to a compute cluster based at least in part on a constraint that a consistency freshness score for the compute cluster satisfies the data consistency requirement for the inference request; and

returning an inference response for the inference request to the user.

Continuity (1)
Continuation 19376656 · Oct 31, 2025
References Cited (10)
US 11422731B1 · Potashnik et al. · 2022 [cited by applicant]
US 12307299B1 · Elbadrashiny et al. · 2025 [cited by applicant]
US 20050203962A1 · Zhou · 2005 [cited by examiner]
US 20170212948A1 · Wang et al. · 2017 [cited by applicant]
US 20180356989A1 · Meister et al. · 2018 [cited by applicant]
US 20200051550A1 · Baker · 2020 [cited by applicant]
US 20250307927A1 · Galvin · 2025 [cited by applicant]
Biswas, Shaikat, “Artificial Intelligence-Enhanced Cybersecurity Frameworks for Real-Time Threat U Detection in Cloud and Enterprise”, ASRC Procedia: Global Perspectives in Science and Scholarship, 1st GRI Conference 20… [cited by applicant]
Cipar, James, et al., “Lazybase: Trading Freshness for Performance in a Scalable Database”, In Proceedings of the 7th ACM European conference on Computer Systems (EuroSys '12), Association for Computing Machinery, New Y… [cited by applicant]
Govindarajulunaidu Sambath Naray, D.B., “Enhancing Data Quality and Consistency in Large-Scale Analytical Systems through AI-Driven Engineering Workflows”, “International Journal of Emerging Trends in Computer Science a… [cited by applicant]