IP Library › Granted Patent US 12,406,662
Granted Patent B2
US 12,406,662 · App. 18/456,672 · Granted Sep 2, 2025

Prompting language models to select API calls

Inventors: Hugh Nicholas Perkins (Queens, NY); Michael Griffiths (Brooklyn, NY); Tao Ma (Mountain View, CA); Connor Daniel McNabb (Denver, CO); Theodore David Burke (Brooklyn, NY); Yi Yang (Long Island City, NY)
Assignee: ASAPP, INC.
G10L15/183G06F9/547G10L15/16G10L15/22G10L15/30H04M3/5166H04M3/5191H04M2201/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,662
App. No.
18/456,672
Granted
Sep 2, 2025
Kind
B2
Abstract

A communications session with a user may be automated using a language model. The language model may be instructed to select a next action to be performed where the next action may include transmitting a responsive communication to the user or performing an API call. The prompt used to query the language model may include one or more of the following: a representation of text of the communications session, a list of available API calls, instructions to select a next action, a representation of API calls performed, or a representation of API call responses received. The language model may be sequentially queried to continue the communications session by transmitting responsive communications or performing API calls. In some implementations, a prompt template may be used to generate the prompt and a prompt template may be selected using text of the communications session.

Claims (60)

1. A computer-implemented method, comprising:

starting a communications session with a user;

obtaining session text of the communications session;

creating a first language model prompt, wherein the first language model prompt comprises: (a) a representation of the session text, (b) a list of available API calls, and (c) instructions to select a next action from a plurality of available next actions, wherein the plurality of available next actions comprises (i) transmitting a responsive communication and (ii) performing an API call;

submitting the first language model prompt to a language model to obtain a first language model response comprising a first next action;

determining that the first next action corresponds to performing a first API call;

performing the first API call to obtain a first API call response;

creating a second language model prompt, wherein the second language model prompt comprises: (a) the representation of the session text, (b) the list of available API calls, (c) the instructions to select a next action from the plurality of available next actions, and (d) a text representation of the first API call response;

submitting the second language model prompt to the language model to obtain a second language model response comprising a second next action;

determining that the second next action corresponds to transmitting a first responsive communication; and

transmitting the first responsive communication to the user.

2. The computer-implemented method of claim 1 , wherein obtaining session text comprises performing automatic speech recognition of an audio communication.

3. The computer-implemented method of claim 1 , wherein the second language model prompt comprises a text representation of the first API call.

4. The computer-implemented method of claim 1 , comprising:

selecting the list of available API calls using the session text.

5. The computer-implemented method of claim 4 , wherein selecting the list of available API calls comprises processing at least a portion of the session text with a classifier neural network.

6. The computer-implemented method of claim 4 , wherein selecting the list of available API calls comprises submitting a third language model prompt to the language model or a second language model.

7. The computer-implemented method of claim 1 , wherein:

the plurality of available next actions comprises selecting a language model prompt template;

selecting the language model prompt template; and

creating a third language model prompt using the language model prompt template.

8. The computer-implemented method of claim 1 , wherein the second language model prompt comprises an expected format of an API call response.

9. A system, comprising at least one server computer comprising at least one processor and at least one memory, the at least one server computer configured to:

start a communications session with a user;

obtain session text of the communications session;

create a first language model prompt, wherein the first language model prompt comprises: (a) a representation of the session text, (b) a list of available API calls, and (c) instructions to select a next action from a plurality of available next actions, wherein the plurality of available next actions comprises (i) transmitting a responsive communication and (ii) performing an API call;

submit the first language model prompt to a language model to obtain a first language model response comprising a first next action;

determine that the first next action corresponds to performing a first API call;

perform the first API call to obtain a first API call response;

create a second language model prompt, wherein the second language model prompt comprises: (a) the representation of the session text, (b) the list of available API calls, (c) the instructions to select a next action from the plurality of available next actions, and (d) a text representation of the first API call response;

submit the second language model prompt to the language model to obtain a second language model response comprising a second next action;

determine that the second next action corresponds to transmitting a first responsive communication; and

transmit the first responsive communication to the user.

10. The system of claim 9 , wherein the language model is provided by a third party and submitting the first language model prompt to the language model comprises transmitting the first language model prompt to the third party.

11. The system of claim 9 , the at least one server computer is configured to:

select a language model prompt template using the session text; and

create the first language model prompt using the language model prompt template.

12. The system of claim 11 , wherein selecting the language model prompt template comprises processing at least a portion of the session text with a classifier neural network.

13. The system of claim 11 , wherein the at least one server computer is configured to select the language model prompt template by submitting a third language model prompt to the language model or a second language model.

14. The system of claim 11 , wherein:

the plurality of available next actions comprises selecting a second list of available API calls; and

the at least one server computer is configured to:

select the second list of available API calls, and

create a third language model prompt using the second list of available API calls.

15. The system of claim 9 , wherein the list of available API calls comprises function names and corresponding function arguments.

16. The system of claim 15 , wherein the list of available API calls comprises function argument types.

17. One or more non-transitory, computer-readable media comprising computer-executable instructions that, when executed, cause at least one processor to perform actions comprising:

starting a communications session with a user;

obtaining session text of the communications session;

creating a first language model prompt, wherein the first language model prompt comprises: (a) a representation of the session text, (b) a list of available API calls, and (c) instructions to select a next action from a plurality of available next actions, wherein the plurality of available next actions comprises (i) transmitting a responsive communication and (ii) performing an API call;

submitting the first language model prompt to a language model to obtain a first language model response comprising a first next action;

determining that the first next action corresponds to performing a first API call;

performing the first API call to obtain a first API call response;

creating a second language model prompt, wherein the second language model prompt comprises: (a) the representation of the session text, (b) the list of available API calls, (c) the instructions to select a next action from the plurality of available next actions, and (d) a text representation of the first API call response;

submitting the second language model prompt to the language model to obtain a second language model response comprising a second next action;

determining that the second next action corresponds to transmitting a first responsive communication; and

transmitting the first responsive communication to the user.

18. The one or more non-transitory, computer-readable media of claim 17 , wherein the user is obtaining customer support from a company.

19. The one or more non-transitory, computer-readable media of claim 17 , wherein the first language model prompt does not include a graph corresponding to an interactive voice response system.

20. The one or more non-transitory, computer-readable media of claim 17 , wherein the representation of the session text comprises an automatically-generated summary of at least a portion of the session text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2023
From: PERKINS, HUGH NICHOLAS; GRIFFITHS, MICHAEL; MA, TAO; MCNABB, CONNOR DANIEL; BURKE, THEODORE DAVID; YANG, YI
To: ASAPP, INC.
Reel/Frame 065342/0636 →
Continuity (1)
Related Publication 20250078822A1 · Mar 6, 2025
References Cited (9)
US 11875123B1 · Ben David · 2024 [cited by examiner]
US 20240394285A1 · Cunningham · 2024 [cited by examiner]
US 20250006196A1 · Wang · 2025 [cited by examiner]
US 20250028743A1 · Massoudian · 2025 [cited by examiner]
US 20250068398A1 · Zeng · 2025 [cited by examiner]
“Significant-Gravitas/AutoGPT”, https://github.com/Significant-Gravitas/AutoGPT (accessed on Sep. 25, 2024), 7 pages. [cited by applicant]
Gray, Shiloh , “ASAPP Launches GenerativeAgent to Automate the Majority of Contact Center Interactions”, ASAPP Press Release, https://www.globenewswire.com/news-release/2023/07/26/2711342/0/en/ASAPP-Launches-GenerativeA… [cited by applicant]
Schick, Timo , et al., “Toolformer: Language Models Can Teach Themselves to Use Tools”, arXiv:2302.04761v1 [cs.CL], https://arxiv.org/pdf/2302.04761 (accessed on Sep. 4, 2024), Feb. 9, 2023, 17 pages. [cited by applicant]
Yao, Shunyu , et al., “ReAct: Synergizing Reasoning and Acting in Language Models”, Published as a conference paper at ICLR 2023, arXiv:2210.03629v3 [cs.CL], https://arxiv.org/pdf/2210.03629 (accessed on Sep. 4, 2024), … [cited by applicant]
Cited By (1)
US 12,699,850