Providing customized services to users of data processing systems using retention policies
Methods and systems for providing customized services to users of data processing systems are disclosed. To provide the customized services, a query may be obtained from a user. Contextual information for the query may be obtained by sequentially searching different tiers of information in a first data source, the different tiers being searched in an order defined by ascribed levels of importance specified by the user to the different tiers until the contextual information meets contextual information criteria, at least one portion of the contextual information being stored in the first data source in accordance with a retention plan that limits durations of time that the at least one portion of contextual information is associated with the different tiers. The query may be serviced using the contextual information and the query as input to an artificial intelligence model to obtain an output usable to facilitate provisioning of computer-implemented services.
1 . A method for providing customized services to users of a data processing system, the method comprising:
obtaining a query from a user of the users;
obtaining contextual information for the query by sequentially searching different tiers of information in a first data source comprising a retrieval-augmented generation (RAG) repository associated with the user, the different tiers being searched in an order defined by ascribed levels of importance specified by the user to the different tiers until the contextual information meets contextual information criteria, at least one portion of the contextual information being stored in the first data source in accordance with a retention plan that limits durations of time that the at least one portion of the contextual information is associated with the different tiers, and the retention plan being based on at least one tag applied to the at least one portion of the contextual information that is based on content of the at least one portion of the contextual information;
obtaining, using the query and the contextual information, an ingest data package that is a RAG output from a RAG pipeline process;
generating an output by an artificial intelligence model using the ingest data package;
providing the output to the user as a customized service;
receiving, from the user, an indication of a level of satisfaction based on the output; and
updating, based on the indication, the RAG repository to obtain a modified first data source, the modified first data source comprising modified tiers of the information.
2 . The method of claim 1 , wherein the at least one tag is based on at least one selected from a group consisting of:
anomalousness of the content;
relevance of the content for auditing of the customized services; and
relevance of the content to other users of the data processing system.
3 . The method of claim 2 , wherein the relevance of the content to the other users is based on user defined preferences for ingest of data with respect to other data sources for the other users.
4 . The method of claim 2 , wherein the anomalousness of the content is with respect to other content stored in the data processing system.
5 . The method of claim 2 , wherein the relevance of the content for auditing is based on utility of the content to reconstruct processes performed by the data processing system.
6 . The method of claim 1 , wherein the retention plan specifies:
an initial tier of the different tiers for storing the at least one portion of the contextual information for a period of time until a time threshold for the initial tier has been met;
a reduction in a representation of the content, the representation being a level of fidelity; and
a destination to migrate the at least one portion of the contextual information to when the time threshold has been met.
7 . The method of claim 6 , wherein the destination is a secondary tier of the different tiers to store the at least one portion of the contextual information for a period of time until the time threshold for the secondary tier has been met.
8 . The method of claim 1 , wherein obtaining the contextual information for the query from the first data source comprises:
comparing information obtained from the first data source to the contextual information criteria, the contextual information criteria comprising a level of enhancement threshold to allow the query to be serviced in a manner that is acceptable to the user;
in a first instance of the comparing in which the information meets the contextual information criteria:
concluding that the contextual information is able to be obtained from the first data source;
in a second instance of the comparing in which the information does not meet the contextual information criteria:
concluding that the contextual information is unable to be obtained from the first data source; and
attempting to obtain the contextual information from other data sources.
9 . The method of claim 1 , wherein the artificial intelligence model is a large language model (LLM).
10 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for providing customized services to users of a data processing system, the operations comprising:
obtaining a query from a user of the users;
obtaining contextual information for the query by sequentially searching different tiers of information in a first data source comprising a retrieval-augmented generation (RAG) repository associated with the user, the different tiers being searched in an order defined by ascribed levels of importance specified by the user to the different tiers until the contextual information meets contextual information criteria, at least one portion of the contextual information being stored in the first data source in accordance with a retention plan that limits durations of time that the at least one portion of the contextual information is associated with the different tiers, and the retention plan being based on at least one tag applied to the at least one portion of the contextual information that is based on content of the at least one portion of the contextual information;
obtaining, using the query and the contextual information, an ingest data package that is a RAG output from a RAG pipeline process;
generating an output by an artificial intelligence model using the ingest data package;
providing the output to the user as a customized service;
receiving, from the user, an indication of a level of satisfaction based on the output; and
updating the RAG repository to obtain a modified first data source, the modified first data source comprising modified tiers of the information.
11 . The non-transitory machine-readable medium of claim 10 , wherein at least one tag is based on at least one selected from a group consisting of:
anomalousness of the content;
relevance of the content for auditing of the customized services; and
relevance of the content to other users of the data processing system.
12 . The non-transitory machine-readable medium of claim 11 , wherein the relevance of the content to the other users of the data processing system is based on user defined preferences for ingest of data with respect to other data sources for the other users.
13 . The non-transitory machine-readable medium of claim 11 , wherein the anomalousness of the content is with respect to other content stored in the data processing system.
14 . The non-transitory machine-readable medium of claim 11 , wherein the relevance of the content for auditing is based on utility of the content to reconstruct processes performed by the data processing system.
15 . A data processing system, comprising:
a processor; and
a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for providing customized services to users of the data processing system, the operations comprising:
obtaining a query from a user of the users;
obtaining contextual information for the query by sequentially searching different tiers of information in a first data source comprising a retrieval-augmented generation (RAG) repository associated with the user, the different tiers being searched in an order defined by ascribed levels of importance specified by the user to the different tiers until the contextual information meets contextual information criteria, at least one portion of the contextual information being stored in the first data source in accordance with a retention plan that limits durations of time that the at least one portion of the contextual information is associated with the different tiers, and the retention plan being based on at least one tag applied to the at least one portion of the contextual information that is based on content of the at least one portion of the contextual information;
obtaining, using the query and the contextual information, an ingest data package that is a RAG output from a RAG pipeline process;
generating an output by an artificial intelligence model using the ingest data package;
providing the output to the user as a customized service;
receiving, from the user, an indication of a level of satisfaction based on the output; and
updating the RAG repository to obtain a modified first data source, the modified first data source comprising modified tiers of the information.
16 . The data processing system of claim 15 , wherein at least one tag is based on at least one selected from a group consisting of:
anomalousness of the content;
relevance of the content for auditing of the customized services; and
relevance of the content to other users of the data processing system.
17 . The data processing system of claim 16 , wherein the relevance of the content to the other users of the data processing system is based on user defined preferences for ingest of data with respect to other data sources for the other users.
18 . The data processing system of claim 16 , wherein the anomalousness of the content is with respect to other content stored in the data processing system.
19 . The data processing system of claim 16 , wherein the relevance of the content for auditing is based on utility of the content to reconstruct processes performed by the data processing system.