IP Library › Granted Patent US 12,632,588
Granted Patent B2
US 12,632,588 · App. 18/823,115 · Granted May 19, 2026

Integrated private cloud AI platform

Inventors: Prakash Mirji (Spring, TX); Swami Viswanathan (Morgan Hill, CA)
Assignee: Hewlett Packard Enterprise Development LP
G06F21/6227G06F40/40H04L51/02H04L63/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,588
App. No.
18/823,115
Granted
May 19, 2026
Kind
B2
Abstract

Systems and methods are provided for an integrated private cloud AI platform that is deployable at a customer site. The integrated private cloud AI platform can include components deployed at the customer site that utilize an embedding model that accesses/integrates with various knowledge bases and vector data stores locally at the customer site, and permit chatbot-accessible queries to an LLM that utilizes the previously-uploaded embeddings.

Claims (57)

1 . A computer-implemented method comprising:

receiving customer data associated with a customer site and converting the customer data into embeddings using an embedding model;

receiving a search query from a client device;

authenticating the client device;

in response to authenticating, providing the search query to a large language model (LLM);

initiating an authorization process confirming that the client device is allowed to access the customer data accessible by the search query;

in response to the authorization process, accessing the customer data identified by the LLM that is stored as the embeddings; and

generating and providing a response to the search query based on the customer data.

2 . The computer-implemented method of claim 1 , wherein the LLM is reused for different search queries and for different use cases.

3 . The computer-implemented method of claim 1 , wherein the embedding model, the LLM, and the authorization process are implemented in a private cloud stored at the customer site, and the client device is configured to access the private cloud.

4 . The computer-implemented method of claim 1 , further comprising:

storing the embeddings in an embedding data store; and

storing vector representations of the embeddings in a vector data store.

5 . The computer-implemented method of claim 1 , wherein the search query is received from a chatbot module that is accessed by the client device.

6 . The computer-implemented method of claim 1 , further comprising:

receiving the embedding model at the customer site; and

once the embedding model is activated, automatically accessing the customer data at the customer site.

7 . The computer-implemented method of claim 1 , further comprising:

receiving a knowledge base from an administrative user before the search query is received;

converting the knowledge base into second embeddings using the embedding model; and

storing the second embeddings in a vector data store at the customer site.

8 . The computer-implemented method of claim 7 , wherein an automated processing job uploads the second embeddings to a namespace associated with the client device.

9 . The computer-implemented method of claim 7 , wherein the embeddings accessed by the LLM are stored in a vector data store.

10 . A private cloud platform comprising:

a memory storing instructions; and

a processor communicatively coupled to the memory and configured to execute the instructions to:

receive customer data associated with a customer site and converting the customer data into embeddings using an embedding model;

receive a search query from a client device;

authenticate the client device;

in response to authenticating, provide the search query to a large language model (LLM);

initiate an authorization process confirming that the client device is allowed to access the customer data accessible by the search query;

in response to the authorization process, access the customer data identified by the LLM that is stored as the embeddings; and

generate and provide a response to the search query based on the customer data.

11 . The private cloud platform of claim 10 , wherein the LLM is reused for different search queries and for different use cases.

12 . The private cloud platform of claim 10 , wherein the embedding model, the LLM, and the authorization process are implemented in a private cloud stored at the customer site, and the client device is configured to access the private cloud.

13 . The private cloud platform of claim 10 , wherein the processor is further configured to:

store the embeddings in an embedding data store; and

store vector representations of the embeddings in a vector data store.

14 . The private cloud platform of claim 10 , wherein the search query is received from a chatbot module that is accessed by the client device.

15 . The private cloud platform of claim 10 , wherein the processor is further configured to:

receive the embedding model at the customer site; and

once the embedding model is activated, automatically access the customer data at the customer site.

16 . The private cloud platform of claim 10 , wherein the processor is further configured to:

receive a knowledge base from an administrative user before the search query is received;

convert the knowledge base into second embeddings using the embedding model; and

store the second embeddings in a vector data store at the customer site.

17 . The private cloud platform of claim 16 , wherein an automated processing job uploads the second embeddings to a namespace associated with the client device.

18 . A non-transitory computer-readable storage medium storing a plurality of instructions executable by a processor, the plurality of instructions when executed by the processor cause the processor to:

receive customer data associated with a customer site and converting the customer data into embeddings using an embedding model;

receive a search query from a client device;

authenticate the client device;

in response to authenticating, provide the search query to a large language model (LLM);

initiate an authorization process confirming that the client device is allowed to access the customer data accessible by the search query;

in response to the authorization process, access the customer data identified by the LLM that is stored as the embeddings; and

generate and provide a response to the search query based on the customer data.

19 . The non-transitory computer-readable storage medium of claim 18 , wherein the LLM is reused for different search queries and for different use cases.

20 . The non-transitory computer-readable storage medium of claim 18 , wherein the embedding model, the LLM, and the authorization process are implemented in a private cloud stored at the customer site, and the client device is configured to access the private cloud.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 3, 2024
From: MIRJI, PRAKASH; VISWANATHAN, SWAMI
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 068475/0018 →
Continuity (1)
Related Publication 20260064874A1 · Mar 5, 2026
References Cited (22)
US 11928426B1 · Gutzeit · 2024 [cited by examiner]
US 12353469B1 · Mahabadi · 2025 [cited by examiner]
US 20210056169A1 · Bahirwani · 2021 [cited by examiner]
US 20220318518A1 · Mars · 2022 [cited by examiner]
US 20240146734A1 · Southgate · 2024 [cited by examiner]
US 20240355065A1 · Miller · 2024 [cited by examiner]
US 20240394571A1 · Ivaturi · 2024 [cited by examiner]
US 20240419830A1 · Park · 2024 [cited by examiner]
US 20250013436A1 · Mohanty · 2025 [cited by examiner]
US 20250094816A1 · Hu · 2025 [cited by examiner]
US 20250103741A1 · Kishan · 2025 [cited by examiner]
US 20250133111A1 · Neystadt · 2025 [cited by examiner]
US 20250211549A1 · Arunachalam · 2025 [cited by examiner]
US 20250370992A1 · Meteer · 2025 [cited by examiner]
US 20250371187A1 · Han · 2025 [cited by examiner]
Unknown author, Embedding (machine learning)—Wikipedia, printed: Jan. 8, 2026, 2 pages (Year: 2026). [cited by examiner]
AWS, “Amazon Bedrock”, available online at <https://aws.amazon.com/bedrock/>, 2024, 10 pages. [cited by applicant]
Cloud Architecture Center, “Infrastructure for a RAG-capable generative AI application using GKE”, available online at <https://cloud.google.com/architecture/rag-capable-gen-ai-app-using-gke>, 2024, 16 pages. [cited by applicant]
Databricks, “A Compact Guide to Retrieval Augmented Generation (RAG)”, 2024, 38 Pages. [cited by applicant]
HPE, “HPE Ezmeral Unified Analytics Software Documentation”, available online at <https://docs.ezmeral.hpe.com/unified-analytics/14/index.html>, 2024, 1 page. [cited by applicant]
InfiniFlow, “Build Generative AI into Your Business”, available online at <https://ragflow.io/>, 2024, 3 pages. [cited by applicant]
Snowflake, “Large Language Model (LLM) Functions (Snowflake Cortex)”, available online at <https://docs.snowflake.com/en/user-guide/snowflake-cortex/llm-functions>, 2024, 22 pages. [cited by applicant]