IP Library Granted Patent US 12664986
Granted Patent B2
US 12664986 · App. 18/404,399 · Granted Jun 23, 2026

Interruption response by an artificial intelligence character

Inventors: James R. Kennedy (Glendale, CA); Reshmashree Bangalore Kantharaju (Pasadena, CA); Maike Paetzel-Pruesmann (Zurich, CH); Ronald Cumbal (Stockholm, SE); Komath Naveen Kumar (Los Angeles, CA)
Assignee: Disney Enterprises, Inc.
G10L15/222G06N3/006G06N3/0455G10L13/10G10L15/1807
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664986
App. No.
18/404,399
Filed
Jan 4, 2024
Granted
Jun 23, 2026
Kind
B2
Art Unit
2692
USPC
704/257
Abstract

A system includes a hardware processor, a memory storing software code, and a machine learning (ML) model trained to detect an interruption to a conversation. The system detects, during a conversational turn by an artificial intelligence (AI) character in interaction with a human and/or another AI character, sound produced by the human and/or the other AI character, classifies, using the ML model, the sound as an interruption or irrelevant to the interaction. When the sound is irrelevant to the interaction, the conversational turn of the AI character continues. When the sound is an interruption, the system identifies a response strategy for continuing the interaction, and executes the response strategy including at least one of: (i) retention of the conversational turn, (ii) relinquishment of the conversational turn, or (iii) a negotiation, with the human and/or the other AI character, to determine the retention or the relinquishment of the conversational turn.

Claims (58)

1 . A system comprising:

a hardware processor;

a system memory storing a software code including an interruption response strategy identification engine and an artificial intelligence (AI) character persona database having stored thereon a plurality of character personas; and

a machine learning (ML) model trained to detect an interruption to a conversation by a participant in the conversation;

the hardware processor configured to execute the software code to:

detect, during a conversational turn by a first AI character in an interaction with at least one of a human or a second AI character, sound produced by the at least one of the human or the second AI character;

classify, using the ML model, the sound as one of an interruption to the interaction by the at least one of the human or the second AI character, or as irrelevant to the interaction;

when the sound is classified as irrelevant to the interaction:

continue the conversational turn by the first AI character;

when the sound is classified as the interruption to the interaction:

identify, using the interruption response strategy identification engine, and based on a first character persona of the plurality of the character personas corresponding to the first AI character, a response strategy for continuing the interaction; and

execute, in real-time using word-level time-stamping and streaming playback of speech of the first AI character, the identified response strategy including at least one of: (i) retention, by the first AI character, of the conversational turn, (ii) relinquishment, by the first AI character, of the conversational turn, or (iii) a negotiation, with the at least one of the human or the second AI character, to determine the retention or the relinquishment of the conversational turn by the first AI character.

2 . The system of claim 1 , wherein when the sound is classified as the interruption, the interruption results in a change in a prosody of speech by the first AI character during the conversational turn.

3 . The system of claim 1 , wherein when the sound is classified as the interruption, the interruption results in at least one of a disfluency in speech by the first AI character or a hesitation in speech by the first AI character during the conversational turn.

4 . The system of claim 1 , wherein when the sound is classified as the interruption, the interruption is acknowledged, by the first AI character, with at least one of an utterance, a gesture, a facial expression, or gaze avoidance.

5 . The system of claim 1 , wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn, and wherein the hardware processor is further configured to execute the software code to:

modify, in response to the interruption, one or more lines of planned speech of the conversational turn.

6 . The system of claim 1 , further comprising:

a natural language generator (NLG);

wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn,

wherein the hardware processor is further configured to execute the software code to:

dynamically generate, using the NLG in response to the interruption, one or more lines of dialogue for completing the conversational turn by the first AI character.

7 . A method for use by a system including a hardware processor and a system memory, the system memory storing a software code including an interruption response strategy identification engine and an artificial intelligence (AI) character persona database having stored thereon a plurality of character personas, and a machine learning (ML) model trained to detect an interruption to a conversation by a participant in the conversation, the method comprising:

detecting, by the software code executed by the hardware processor, during a conversational turn by a first AI character in an interaction with at least one of a human or a second AI character, sound produced by the at least one of the human or the second AI character;

classifying, by the software code executed by the hardware processor and using the ML model, the sound as one of an interruption to the interaction by the at least one of the human or the second AI character, or as irrelevant to the interaction;

when the sound is classified as irrelevant to the interaction:

continuing, by the software code executed by the hardware processor, the conversational turn by the first AI character;

when the sound is classified as the interruption:

identifying, by the software code executed by the hardware processor, using the interruption response strategy identification engine and based on a first character persona of the plurality of the character personas corresponding to the first AI character, a response strategy for continuing the interaction; and

executing, by the software code executed by the hardware processor, in real-time using word-level time-stamping and streaming playback of speech of the first AI character, the identified response strategy including at least one of: (i) retention, by the first AI character, of the conversational turn, (ii) relinquishment, by the first AI character, of the conversational turn, or (iii) a negotiation, with the at least one of the human or the second AI character, to determine the retention or the relinquishment of the conversational turn by the first AI character.

8 . The method of claim 7 , wherein when the sound is classified as the interruption, the interruption results in at least one of a change in a prosody of speech by the first AI character, a disfluency in speech by the first AI character, or a hesitation in speech by the first AI character during the conversational turn.

9 . The method of claim 7 , wherein when the sound is classified as the interruption, the first AI character acknowledges the interruption with at least one of an utterance, a gesture, a facial expression, or gaze avoidance.

10 . The method of claim 7 , wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn, and the method further comprises:

modifying, by the software code executed by the hardware processor in response to the interruption, one or more lines of planned speech of the conversational turn.

11 . The method of claim 7 , wherein the system further comprises a natural language generator (NLG),

wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn, and

the method further comprises:

dynamically generating, by the software code executed by the hardware processor and using the NLG in response to the interruption, one or more lines of dialogue for completing the conversational turn by the first AI character.

12 . A computer-readable non-transitory medium having stored thereon instructions, which when executed by a hardware processor, instantiate a method comprising:

detecting during a conversational turn by a first artificial intelligence (AI) character in an interaction with at least one of a human or a second AI character, sound produced by the at least one of the human or the second AI character;

classifying using a machine learning (ML) model trained to detect an interruption to a conversation by a participant in the conversation, the sound as one of an interruption by the at least one of the human or the second AI character, or as irrelevant to the interaction;

when the sound is classified as irrelevant to the interaction:

continuing the conversational turn by the first AI character;

when the sound is classified as the interruption:

identifying, using an interruption response strategy identification engine, and based on a character persona corresponding to the first AI character, a response strategy for continuing the interaction; and

executing, in real-time using word-level time-stamping and streaming playback of speech of the first AI character, the identified response strategy including at least one of: (i) retention, by the first AI character, of the conversational turn, (ii) relinquishment, by the first AI character, of the conversational turn, or (iii) a negotiation, with the at least one of the human or the second AI character, to determine the retention or the relinquishment of the conversational turn by the first AI character.

13 . The computer-readable non-transitory medium of claim 12 , wherein when the sound is classified as the interruption, the interruption results in at least one of a change in a prosody of speech by the first AI character, a disfluency in speech by the first AI character, or a hesitation in speech by the first AI character during the conversational turn.

14 . The computer-readable non-transitory medium of claim 12 , wherein when the sound is classified as the interruption, the interruption is acknowledged by the first AI character with at least one of an utterance, a gesture, a facial expression, or gaze avoidance.

15 . The computer-readable non-transitory medium of claim 12 , wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn, and the method further comprises:

modifying, in response to the interruption, one or more lines of planned speech of the conversational turn.

16 . The computer-readable non-transitory medium of claim 12 , wherein executing the identified response strategy includes retention, by the first AI character, of the conversational turn, and the method further comprises:

dynamically generating, using a natural language generator in response to the interruption, one or more lines of dialogue for completing the conversational turn by the first AI character.

17 . The system of claim 1 , wherein identifying the response strategy for continuing the interaction is further based on at least one of conversational context data, a conversational turn importance score, or an interruption importance score.

18 . The system of claim 4 , further comprising a robot embodying the first AI character, the robot including at least one mechanical actuator, and wherein the hardware processor is further configured to execute the software code to:

produce, using the at least on mechanical actuator, the at least one of the utterance, the gesture, the facial expression, or the gaze avoidance by the robot embodying the first AI character.

19 . The method of claim 7 , wherein identifying the response strategy for continuing the interaction is further based on at least one of conversational context data, a conversational turn importance score, or an interruption importance score.

20 . The method of claim 9 , wherein the system further comprises a robot embodying the first AI character, the robot including at least one mechanical actuator, and the method further comprising:

producing, by the software code executed by the hardware processor and using the at least one mechanical actuator, the at least one of the utterance, the gesture, the facial expression, or the gaze avoidance by the robot embodying the first AI character.