IP Library › Granted Patent US 11,960,983
Granted Patent B1
US 11,960,983 · App. 18/400,814 · Granted Apr 16, 2024

Pre-fetching results from large language models

Inventors: Ilya Gelfenbeyn (Palo Alto, CA); Mikhail Ermolenko (Mountain View, CA); Kylan Gibbs (San Francisco, CA); Evgenii Shingarev (Mountain View, CA)
Assignee: Theai, Inc.
G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,960,983
App. No.
18/400,814
Filed
Dec 29, 2023
Granted
Apr 16, 2024
Kind
B1
Art Unit
2634
USPC
706/18
Abstract

Systems and methods for pre-fetching results from large language models (LLMs) are provided. The method includes acquiring a context of an interaction between a user and an Artificial Intelligence (AI) character; predicting, based on the context, one or more anticipated words to be uttered by the user; generating, based on the one or more anticipated words, at least one query to an LLM; providing the at least one query to the LLM; generating, based on at least one response obtained from the LLM, an anticipated reply of the AI character model to the one or more anticipated words to be pronounced by the user; receiving one or more words uttered by the user; determining that a level of a discrepancy between the one or more words and the one or more anticipated words is below a predetermined threshold; and providing the anticipated reply to the user.

Claims (70)

1. A method comprising:

acquiring, by at least one processor, a context of an interaction between a user and an Artificial Intelligence (AI) character generated by an AI character model in a virtual environment;

predicting, by the at least one processor and based on the context, one or more anticipated words to be uttered by the user;

generating, by the at least one processor and based on the one or more anticipated words, at least one query to a large language model (LLM);

providing, by the at least one processor, the at least one query to the LLM to obtain at least one response from the LLM;

generating, by the at least one processor and based on the at least one response from the LLM, an anticipated reply of the AI character model to the one or more anticipated words to be pronounced by the user;

receiving, by the at least one processor, one or more words uttered by the user;

determining, by the at least one processor, that a level of a discrepancy between the one or more words and the one or more anticipated words is below a predetermined threshold; and

based on the determination, providing, by the at least one processor, the anticipated reply to the user.

2. The method of claim 1 , wherein the context is determined based on historical data concerning interactions between the user and the AI character.

3. The method of claim 1 , further comprising, prior to the receiving the one or more words uttered by the user:

establishing that the user has started uttering the one or more words;

analyzing a part of the one or more words to update the context;

generate, based on the updated context, at least one updated query to the LLM;

sending the at least one updated query to the LLM to receive at least one updated response from the LLM; and

generating, based on the at least one updated response from the LLM, the anticipated reply of the AI character model to the user.

4. The method of claim 3 , wherein the at least one updated query is sent to the LLM before the user has finished pronouncing the rest of the one or more words.

5. The method of claim 3 , wherein the anticipated reply is provided to the user before the user has finished pronouncing the rest of the one or more words.

6. The method of claim 1 , wherein:

the at least one response from the LLM includes a first response and a second response; and

the generating the anticipated reply includes:

assigning, based on the one or more anticipated words, probabilities to the first response and the second response; and

selecting, based on the probabilities, one of the first response and the second response.

7. The method of claim 1 , wherein:

the at least one response from the LLM includes a first response and a second response;

the first response and the second response are assigned by the LLM with probabilities; and

the generating the anticipated reply includes selecting, based on the probabilities, one of the first response and the second response.

8. The method of claim 1 , wherein the one or more anticipated words are predicted using a pretrained neural network.

9. The method of claim 1 , wherein the context is determined based on one or more of the following: an age of the user, a gender of the user, and a behavioral pattern of the user.

10. The method of claim 1 , wherein the context is determined based on one or more of the following: a gesture of the user, an action performed by the user, and a movement performed by the user.

11. The method of claim 1 , wherein the context is determined based on current parameters of the virtual environment.

12. The method of claim 1 , wherein the level of the discrepancy between the one or more words and the one or more anticipated words is based on a number of differences between the one or more words and the one or more anticipated words.

13. A computing system comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the computing system to:

acquire a context of an interaction between a user and an Artificial Intelligence (AI) character generated by an AI character model in a virtual environment;

predict, based on the context, one or more anticipated words to be uttered by the user;

generate, based on the one or more anticipated words, at least one query to a large language model (LLM);

provide the at least one query to the LLM to obtain at least one response from the LLM;

generate, based on the at least one response from the LLM, an anticipated reply of the AI character model to the one or more anticipated words to be pronounced by the user;

receive one or more words uttered by the user;

determine that a level of a discrepancy between the one or more words and the one or more anticipated words is below a predetermined threshold; and

based on the determination, provide the anticipated reply to the user.

14. The computing system of claim 13 , wherein the context is determined based on historical data concerning interactions between the user and the AI character.

15. The computing system of claim 13 , wherein the instructions further configure the computing system to, prior to the receiving the one or more words uttered by the user:

establish that the user has started uttering the one or more words;

analyze a part of the one or more words to update the context;

generate, based on the updated context, at least one updated query to the LLM;

send the at least one updated query to the LLM to receive at least one updated response from the LLM; and

generate, based on the at least one updated response from the LLM, the anticipated reply of the AI character model to the user.

16. The computing system of claim 15 , wherein the at least one updated query is sent to the LLM before the user has finished pronouncing the rest of the one or more words.

17. The computing system of claim 15 , wherein the anticipated reply is provided to the user before the user has finished pronouncing the rest of the one or more words.

18. The computing system of claim 13 , wherein:

the at least one response from the LLM includes a first response and a second response; and

the generating the anticipated reply includes:

assigning, based on the one or more anticipated words, probabilities to the first response and the second response; and

selecting, based on the probabilities, one of the first response and the second response.

19. The computing system of claim 13 , wherein:

the at least one response from the LLM includes a first response and a second response;

the first response and the second response are assigned by the LLM with probabilities; and

the generating the anticipated reply includes selecting, based on the probabilities, one of the first response and the second response.

20. A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that, when executed by a computing system, cause the computing system to:

acquire a context of an interaction between a user and an Artificial Intelligence (AI) character generated by an AI character model in a virtual environment;

predict, based on the context, one or more anticipated words to be uttered by the user;

generate, based on the one or more anticipated words, at least one query to a large language model (LLM);

provide the at least one query to the LLM to obtain at least one response from the LLM;

generate, based on the at least one response from the LLM, an anticipated reply of the AI character model to the one or more anticipated words to be pronounced by the user;

receive one or more words uttered by the user;

determine that a level of a discrepancy between the one or more words and the one or more anticipated words is below a predetermined threshold; and

based on the determination, provide the anticipated reply to the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2023
From: GELFENBEYN, ILYA; ERMOLENKO, MIKHAIL; GIBBS, KYLAN; SHINGAREV, EVGENII
To: THEAI, INC.
Reel/Frame 065984/0325 →
Continuity (1)
Provisional Application 63436116 · Dec 30, 2022
Cited By (4)
US 12,254,005 US 12,430,370 US 12,541,496 US 12,664,969