IP Library Granted Patent US 12706918
Granted Patent B1
US 12706918 · App. 19/240,026 · Granted Aug 11, 2026

Seamless consumer integration of access to a generative response engine

Inventors: David Cummings (San Francisco, CA); Athyuttam Eleti (San Francisco, CA); Miqdad Jaffer (San Francisco, CA)
Assignee: OpenAI OpCo, LLC
H04L63/102G06F16/2457H04L9/3228H04L63/0421H04L63/0428
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12706918
App. No.
19/240,026
Granted
Aug 11, 2026
Kind
B1
Abstract

A device may receive, from a user of a computing device and via an application, a user query. A device may generate a data package comprising the user query and user identity information. A device may transmit, via a company application programming interface (API) between the application and a generative response engine, the data package to the generative response engine. The generative response engine provides resources to process the user query based on entitlements associated with a user account. User anonymity and user entitlements can be fulfilled for the user experience when the user interacts with the generative response engine through the company API, even when the user is not logged into the generative response engine.

Claims (45)

1 . A method of providing access to a generative response engine, the method comprising:

receiving, from a user of a computing device and via an application, a user query;

generating a data package comprising the user query and user identity information;

transmitting, via a company application programming interface between the application and a generative response engine, the data package to the generative response engine, wherein the generative response engine provides resources to process the user query based on entitlements associated with a user account of the user, wherein the entitlements include at least a rate limiting requirement;

managing consumption tracking of the user account without revealing an identity of the user to the application or to a network-based server associated with the application, wherein the application limits access to the generative response engine for the user based on rate-limiting data obtained from the generative response engine that is specific to the user account; and

receiving a response from the generative response engine.

2 . The method of claim 1 , further comprising:

presenting the response to the user via the application.

3 . The method of claim 1 , further comprising:

receiving, from the generative response engine, the rate-limiting data based on the user account of the user.

4 . The method of claim 1 , further comprising:

transmitting the data package to the generative response engine according to an end-to-end encryption protocol.

5 . The method of claim 1 , further comprising:

limiting use of the generative response engine for the user via the application based on the rate-limiting data specific to the user account and obtained from the generative response engine.

6 . The method of claim 1 , wherein the consumption tracking for the user is maintained on the computing device and independent of the application to maintain privacy for the user.

7 . The method of claim 1 , wherein managing, via an anonymous rate-limiting process, the consumption tracking of the user.

8 . The method of claim 1 , wherein an experience of the user with providing the user query and receiving the response from the generative response engine is provided according to the entitlements for the user according to the user account.

9 . The method of claim 1 , wherein privacy is maintained relative to personal data being provided to the application or an associated network server to the application via an authentication flow and via use of temporary refreshable access tokens.

10 . The method of claim 1 , wherein the application enforces throttling locally based on the rate-limiting data returned by the generative response engine.

11 . The method of claim 1 , wherein the entitlement further includes a subscription level associated with the user account of the user and wherein the application enforces a quality of service of the generative response engine consistent with the subscription level.

12 . A system for providing enhanced artificial intelligence engine language assistance, the system comprising:

at least one processor; and

a computer-readable storage device storing instructions, which, when executed by the at least one processor, cause the at least one processor to be configured to:

receive, from a user of a computing device and via an application, a user query;

generate a data package comprising the user query and user identity information; and

transmit, via a company application programming interface between the application and a generative response engine, the data package to the generative response engine, wherein the generative response engine provides resources to process the user query based on entitlements associated with a user account of the user and wherein the entitlements include at least rate liming requirement;

manage consumption tracking of the user without revealing an identity of the user to the application or to a network-based server associated with the application, wherein the application limits access to the generative response engine for the user based on rate-limiting data obtained from the generative response engine that is specific to the user account; and

receive a response from the generative response engine.

13 . The system of claim 12 , wherein the at least one processor is further configured to:

present the response to the user via the application.

14 . The system of claim 12 , wherein the at least one processor is further configured to:

receive, from the generative response engine, the rate-limiting data based on the user account of the user.

15 . The system of claim 12 , wherein the at least one processor is further configured to:

transmit the data package to the generative response engine according to an end-to-end encryption protocol.

16 . The system of claim 12 , wherein the at least one processor is further configured to:

limit use of the generative response engine for the user via the application based on the rate-limiting data specific to the user account and obtained from the generative response engine.

17 . The system of claim 12 , wherein the consumption tracking for the user is maintained on the computing device and independent of the application to maintain privacy for the user.

18 . The system of claim 12 , wherein the at least one processor is configured to manage the consumption tracking for the user, via an anonymous rate-limiting process.

19 . A computer-readable storage device storing instructions, which, when executed by at least one processor, cause the at least one processor to be configured to:

receive, from a user of a computing device and via an application, a user query;

generate a data package comprising the user query and user identity information; and

transmit, via a company application programming interface between the application and a generative response engine, the data package to the generative response engine, wherein the generative response engine provides resources to process the user query based on entitlements associated with a user account of the user, wherein the entitlements include at least rate liming requirement;

manage consumption tracking of the user without revealing an identity of the user to the application or to a network-based server associated with the application, wherein the application limits access to the generative response engine for the user based on rate-limiting data obtained from the generative response engine that is specific to the user account; and

receive a response from the generative response engine.

20 . The computer-readable storage device of claim 19 , wherein privacy is maintained relative to personal data being provided to the application or an associated network server to the application via an authentication flow and via a use of temporary refreshable access tokens.