IP Library Granted Patent US 12682168
Granted Patent B1
US 12682168 · App. 18/622,557 · Granted Jul 14, 2026

Structured interleaving of generations and external interactions for conversation-based generative artificial intelligence applications using extensible role sets and hidden markers

Inventors: Salvatore Romeo (Alexandria, VA); Yi Zhang (Sammamish, WA); Tamer A N Alkhouli (San Jose, CA); Daniele Bonadiman (Sunnyvale, CA); Zhongyuan Zhu (Bellevue, WA); Monica Lakshmi Sunkara (San Jose, CA); Laura Aina (Barcelona, ES); Sarthak Jain (Seattle, WA); Arshit Gupta (Sunnyvale, CA); Bonan Min (Palo Alto, CA); Kalpit Dixit (Mountain View, CA); Brant Swidler (Kauneonga Lake, NY); Miguel Ballesteros Martinez (New York, NY); Yassine Benajiba (Briarcliff Manor, NY); Katrin Kirchhoff (Seattle, WA); Dan Roth (Philadelphia, PA)
Assignee: Amazon Technologies, Inc.
G06F40/284G06F16/3329
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682168
App. No.
18/622,557
Granted
Jul 14, 2026
Kind
B1
Abstract

An automated conversation intermediary performs several operations during a conversation between end users and chatbots. Prior to adding natural language input from an end user to a chatbot prompt, the intermediary annotates the input to indicate a role of the end user, using hidden markers for the annotation that are not presented to the end user during the conversation. The intermediary receives output generated by the chatbot, which also includes hidden markers in addition to natural language to be presented to the end user, and removes the hidden markers before providing the natural language portion of the output to the end user.

Claims (62)

1 . A system, comprising:

one or more computing devices;

wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices cause the one or more computing devices to:

deploy, using resources of a network-accessible generative artificial intelligence (GAI)-based application management service, a first automated conversation intermediary for interactions among a set of conversation contributor entities, wherein individual ones of the conversation contributor entities have respective pre-defined roles associated with multi-turn conversations between one or more chatbots and one or more end users, wherein the respective pre-defined roles belong to a first extensible role set of a conversation management framework, and wherein the first extensible role set includes at least (a) a first user category role, (b) a second user category role, (c) a natural language output source role, (d) an external action invocator role and (e) an external action implementer role;

perform, by the first automated conversation intermediary at the GAI-based application management service, a plurality of operations to orchestrate a particular conversation between at least a first end user and a first chatbot, wherein the first end user has the first user category role, and wherein the plurality of operations includes:

prior to providing, to the first chatbot as part of a prompt, natural language input obtained from the first end user, annotating the natural language input with (a) an indication that the natural language input was obtained from an entity that has the first user category role and (b) at least a first subset of a collection of hidden markers of the conversation management framework, wherein hidden markers of the collection are not included in output provided to the first end user during the particular conversation, wherein individual ones of the hidden markers of the collection comprise respective bit sequences which do not correspond to Unicode encodings, and wherein the first subset includes a turn start marker, a role marker, and a turn end marker;

receiving, from the first chatbot subsequent to processing of the annotated natural language input by the first chatbot, a first chatbot-generated token sequence which includes (a) an indication of a particular external action to be implemented by a particular external action implementer (b) an indication that the first chatbot-generated token sequence was generated by an entity that has an external action invocator role, and (c) a second subset of the collection of hidden markers, wherein the second subset includes the turn start marker, the role marker, the turn end marker, and an action invocation metadata marker;

obtaining an action result token sequence of the particular external action, wherein the action result token sequence is generated at least in part by the particular external action implementer;

prior to appending the action result token sequence to the prompt of the first chatbot, annotating the action result token sequence with (a) an indication that the action result token sequence was obtained from an entity that has the external action implementer role and (b) a third subset of the collection of hidden markers, wherein the third subset includes the turn start marker, the role marker, the turn end marker, and an action result metadata marker;

receiving, from the first chatbot subsequent to processing of the annotated action result token sequence by the chatbot, a second chatbot-generated token sequence which includes (a) a natural language response generated by the first chatbot for the first end user (b) an indication that the second chatbot-generated token sequence was generated by an entity that has a natural language output source role, and (c) a fourth subset of the collection of hidden markers, wherein the fourth subset includes the turn start marker, the role marker, and the turn end marker; and

presenting an unmarked natural language response token sequence to the first end user, wherein the unmarked natural language response token sequence is generated at least in part by removing the fourth subset of hidden markers from the second chatbot-generated token sequence.

2 . The system as recited in claim 1 , wherein particular conversation comprises interactions between a second end user and the first chatbot, and wherein the plurality of operations includes:

prior to providing, to the first chatbot as part of the prompt, a set of natural language input obtained from the second end user, annotating the set of natural language input with an indication that the set of natural language input was obtained from an entity that has the second user category role.

3 . The system as recited in claim 1 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

prior to initiation of the particular conversation, fine-tune the first chatbot using training input which comprises (a) examples of one or more hidden markers of the collection of hidden markers and (b) examples of one or more roles of the first extensible role set.

4 . The system as recited in claim 1 , wherein the plurality of operations includes:

receiving, from the first chatbot subsequent to processing of additional annotated natural language input of the first end user by the first chatbot, a third chatbot-generated token sequence which includes (a) a representation of a particular intermediate reasoning step completed by the first chatbot, (b) an indication that the third chatbot-generated token sequence was generated by an entity that has a chain-of-thought reasoning indicator role of the first extensible role set, and (c) a fifth subset of the collection of hidden markers, wherein the fifth subset includes the turn start marker, the role marker, and the turn end marker; and

appending the third chatbot-generated token sequence to the prompt of the first chatbot.

5 . The system as recited in claim 4 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:

store the indication of the particular intermediate reasoning step in a repository; and

present, in response to a programmatic request for an explanation of natural language output provided to the first end user during the particular multi-turn conversation, the indication of the particular intermediate reasoning step.

6 . A computer-implemented method, comprising:

performing, by a first automated conversation intermediary, a plurality of operations during a particular conversation between a set of end users and a set of chatbots, wherein the set of end users includes a first end user, wherein the set of chatbots includes a first chatbot, and wherein the plurality of operations includes:

prior to providing, to the first chatbot as part of a prompt, natural language input obtained from the first end user, annotating the natural language input with (a) an indication that the natural language input was obtained from an entity that has a first user category role of a role set defined for conversation contributing entities in a conversation management framework, and (b) at least a first subset of a collection of hidden markers of the conversation management framework, wherein hidden markers of the collection are not included in output provided to the first end user during the particular conversation, and wherein the first subset includes a role marker;

receiving, from the first chatbot subsequent to processing of the prompt by the first chatbot, a first chatbot-generated token sequence which includes (a) a natural language response generated by the first chatbot for the first end user (b) an indication that the first chatbot-generated token sequence was generated by an entity that has a natural language output source role of the role set, and (c) a second subset of the collection of hidden markers, wherein the second subset includes the role marker; and

presenting an unmarked natural language response token sequence to the first end user, wherein the unmarked natural language response token sequence is generated at least in part by removing the second subset of hidden markers from the first chatbot-generated token sequence.

7 . The computer-implemented method as recited in claim 6 , wherein the plurality of operations includes:

in response to (a) receiving additional natural language input from the first end user and (b) detecting that the additional natural language input includes at least one hidden marker of the collection or hidden markers, rejecting the additional natural language input and terminating the particular conversation.

8 . The computer-implemented method as recited in claim 6 , wherein the plurality of operations includes:

in response to (a) receiving a second chatbot-generated token sequence from the first chatbot, and (b) detecting that a natural language output portion of the second chatbot-generated token sequence includes at least one hidden marker of the collection of hidden markers, discarding the second chatbot-generated token sequence without presenting a portion of the second chatbot-generated token sequence to the first end user.

9 . The computer-implemented method as recited in claim 6 , wherein the plurality of operations includes:

receiving, from the first chatbot subsequent to processing of additional annotated natural language input of the first end user by the first chatbot, a second chatbot-generated token sequence which includes (a) an indication of a particular external action to be implemented by a particular external action implementer using a resource other than the first chatbot (b) an indication that the second chatbot-generated token sequence was generated by an entity that has an external action invocator role of the role set, and (c) a third subset of the collection of hidden markers, wherein the second subset includes the role marker and an action invocation metadata marker;

obtaining an action result token sequence of the particular external action, wherein the action result token sequence is generated at least in part by the particular external action implementer; and

prior to appending the action result token sequence to the prompt of the first chatbot, annotating the action result token sequence with (a) an indication that the action result token sequence was obtained from an entity that has an external action implementer role of the role set and (b) a fourth subset of the collection of hidden markers, wherein the fourth subset includes the role marker and an action result metadata marker.

10 . The computer-implemented method as recited in claim 9 , wherein the particular external action includes an operation performed at one or more of: (a) a read-only data source, (b) a search engine, or (c) a transaction processing system.

11 . The computer-implemented method as recited in claim 9 , wherein (a) the action invocation metadata marker indicates a position, within the second chatbot-generated token sequence, of metadata associated with a request for invocation of the particular external action implementer and (b) the action result metadata marker indicates a position, within the action result token sequence, of metadata associated with a result of invoking the particular external action.

12 . The computer-implemented method as recited in claim 11 , wherein the metadata associated with the request for invocation of the particular external action implementer comprises one or more of: (a) a timestamp, (b) a credential, or (c) an indication of whether approval from a specified entity is to be obtained to enable execution of the particular external action implementer.

13 . The computer-implemented method as recited in claim 11 , wherein the metadata associated with the result of invoking the particular external action comprises one or more of: (a) a timestamp, or (b) a status indication of the particular external action.

14 . The computer-implemented method as recited in claim 6 , wherein the plurality of operations includes:

receiving, from the first chatbot subsequent to processing of additional annotated natural language input of the first end user by the first chatbot, a second chatbot-generated token sequence which includes (a) an indication of a particular intermediate reasoning step of the first chatbot (b) an indication that the second chatbot-generated token sequence was generated by an entity that has a reasoning indicator role of the role set, and (c) a third subset of the collection of hidden markers, wherein the third subset includes the role marker; and

appending the second chatbot-generated token sequence to the prompt of the first chatbot; and

wherein the computer-implemented method further comprises:

storing the indication of the particular intermediate reasoning step in a repository; and

presenting, in response to a programmatic request for an explanation of natural language output provided to the first end user during the particular conversation, the indication of the particular intermediate reasoning step.

15 . The computer-implemented method as recited in claim 6 , wherein training of the first chatbot comprises a pre-training phase and at least one fine-tuning phase, the computer-implemented method further comprising:

in response to determining that a particular role is to added to the role set for one or more additional conversations, performing one or more of: (a) deploying an updated version of the first automated conversation intermediary, wherein the updated version includes logic to detect and process conversation turns associated with the particular role or (b) performing an additional fine-tuning phase of the first chatbot, wherein input provided to the first chatbot during the additional fine-tuning phase includes examples of conversation turns associated with the particular role.

16 . One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:

perform, by a first automated conversation intermediary, a plurality of operations during a particular conversation between a set of end users and a set of chatbots, wherein the set of end users includes a first end user, wherein the set of chatbots includes a first chatbot, and wherein the plurality of operations includes:

prior to providing, to the first chatbot as part of a prompt, natural language input obtained from the first end user, annotating the natural language input with (a) an indication that the natural language input was obtained from an entity that has a first user category role of a role set defined for conversation contributing entities in a conversation management framework, and (b) at least a first subset of a collection of hidden markers of the conversation management framework, wherein hidden markers of the collection are not included in output provided to the first end user during the particular conversation, and wherein the first subset includes a role marker;

receiving, from the first chatbot subsequent to processing of the prompt by the first chatbot, a first chatbot-generated token sequence which includes (a) a natural language response generated by the first chatbot for the first end user (b) an indication that the first chatbot-generated token sequence was generated by an entity that has a natural language output source role of the role set, and (c) a second subset of the collection of hidden markers, wherein the second subset includes the role marker; and

presenting an unmarked natural language response token sequence to the first end user, wherein the unmarked natural language response token sequence is generated at least in part by removing the second subset of hidden markers from the first chatbot-generated token sequence.

17 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the plurality of operations includes:

in response to (a) receiving additional natural language input from the first end user and (b) detecting that the additional natural language includes a particular hidden marker of the collection or hidden markers, removing at least the particular hidden marker from the additional natural language prior to providing an annotated version of the additional natural language to the first chatbot.

18 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the plurality of operations includes:

in response to (a) receiving a second chatbot-generated token sequence from the first chatbot, and (b) detecting that a natural language output portion of the second chatbot-generated token sequence includes a particular hidden marker of the collection of hidden markers, removing at least the particular hidden marker from the natural language output portion prior to providing the natural language output portion to the first end user.

19 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the plurality of operations includes:

receiving, from the first chatbot subsequent to processing of additional annotated natural language input of the first end user by the first chatbot, a second chatbot-generated token sequence which includes (a) an indication of a particular external action to be implemented by a particular external action implementer using a resource other than the first chatbot (b) an indication that the second chatbot-generated token sequence was generated by an entity that has an external action invocator role of the role set, and (c) a third subset of the collection of hidden markers, wherein the second subset includes the role marker and an action invocation metadata marker;

obtaining an action result token sequence of the particular external action, wherein the action result token sequence is generated at least in part by the particular external action implementer; and

prior to appending the action result token sequence to the prompt of the first chatbot, annotating the action result token sequence with (a) an indication that the action result token sequence was obtained from an entity that has an external action implementer role of the role set and (b) a fourth subset of the collection of hidden markers, wherein the fourth subset includes the role marker and an action result metadata marker.

20 . The one or more non-transitory computer-accessible storage media as recited in claim 16 , wherein the plurality of operations includes:

receiving, from the first chatbot subsequent to processing of additional annotated natural language input of the first end user by the first chatbot, a second chatbot-generated token sequence which includes (a) an indication of a particular intermediate reasoning step of the first chatbot (b) an indication that the second chatbot-generated token sequence was generated by an entity that has a reasoning indicator role of the role set, and (c) a third subset of the collection of hidden markers, wherein the third subset includes the role marker; and

appending the second chatbot-generated token sequence to the prompt of the first chatbot.