IP Library Granted Patent US 12699856
Granted Patent B1
US 12699856 · App. 18/239,780 · Granted Aug 4, 2026

Automatic translation using an interaction application

Inventors: William Seo (Atlanta, GA); Garrett Paul Simmer (Atlanta, GA); Mallikarjuna Bachu (Cumming, GA); Chinar Dankhara (Atlanta, GA); Vivian Aranha (Atlanta, GA)
Assignee: Delta Air Lines, Inc.
G06F40/58G10L17/02G10L17/04G10L17/14G06Q50/14
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12699856
App. No.
18/239,780
Granted
Aug 4, 2026
Kind
B1
Abstract

Automatic translation using an interaction application can include detecting a conversation between an agent and a customer, capturing audio data associated with the interaction, and identifying a speaker associated with speech represented by the audio data. If the speaker is determined to be the customer, a customer language can be determined, the speech can be translated from the customer language to the agent language, and the translation can be output to the agent. If the speaker is determined to be the agent, the speech can be translated to the customer language and output to the customer.

Claims (83)

1 . A device comprising a processor and a memory that stores computer-executable instructions that, when executed by the processor, cause the processor to perform operations comprising:

detecting, at an agent device comprising sound hardware, a conversation between an agent and a customer, wherein data defining an agent language spoken by the agent and an agent voiceprint is stored at the agent device, and wherein the agent voiceprint comprises a unique pattern of voice characteristics of the agent;

capturing, using the sound hardware and during the conversation, audio data associated with the conversation, wherein the audio data comprises audio signals captured using the sound hardware, and wherein the audio signals represent first speech associated with the agent and second speech associated with the customer;

identifying, based on analysis of the audio signals, a first speaker associated with the first speech and a second speaker associated with the second speech, wherein the first speaker is determined to be the agent and to be speaking the agent language in response to speech characteristics of the first speech matching the agent voiceprint, and wherein the second speaker is determined to be the customer in response to the speech characteristics not matching the agent voiceprint;

determining, based on further analysis of the audio data, a customer language comprising a language spoken by the customer;

translating the second speech represented by the audio data to the agent language;

outputting, at a display of the agent device, first text representing a translation of the second speech in the agent language;

translating the first speech to the customer language; and

outputting, at the display of the agent device, second text representing a translation of the first speech.

2 . The device of claim 1 , wherein identifying the first speaker further comprises determining that the first speech is spoken in the agent language.

3 . The device of claim 1 , wherein outputting the first text and the second text at the display of the agent device comprises:

generating a user interface comprising the first text oriented in a first orientation and the second text oriented in a second orientation that is upside down relative to the first orientation; and

displaying the user interface at the agent device.

4 . The device of claim 1 , wherein the agent voiceprint is generated at the agent device, and wherein generating the agent voiceprint comprises:

obtaining, using the sound hardware, further audio data that represents a voice associated with the agent;

creating, based on the further audio data, the agent voiceprint; and

storing the agent voiceprint with a data association that associates the agent voiceprint with the agent.

5 . The device of claim 1 , wherein determining the customer language comprises:

determining a language likelihood threshold defined for the agent device, the language likelihood threshold comprising a minimum probability that the second speech is spoken in a particular language;

identifying a plurality of languages that meet the language likelihood threshold; and

outputting, on the display of the agent device, a display of a ranked language listing that identifies the plurality of languages in an order from largest probability to smallest probability, wherein a respective identifier of each of the plurality of languages is translated into a corresponding language, and wherein the customer language is selected from the ranked language listing to determine the customer language.

6 . The device of claim 1 , wherein determining the customer language comprises:

determining a first language spoken at a flight origin associated with a flight and a second language spoken at a flight destination associated with the flight,

determining that the second language corresponds to the agent language, and

determining that the first language comprises the customer language.

7 . A method comprising:

detecting, at an agent device comprising a processor and sound hardware, a conversation between an agent and a customer, wherein data defining an agent language spoken by the agent and an agent voiceprint is stored at the agent device, and wherein the agent voiceprint comprises a unique pattern of voice characteristics of the agent;

capturing, by the processor and using the sound hardware and during the conversation, audio data associated with the conversation, wherein the audio data comprises audio signals captured using the sound hardware, and wherein the audio signals represent first speech associated with the agent and second speech associated with the customer;

identifying, by the processor and based on analysis of the audio signals, a first speaker associated with the first speech and a second speaker associated with the second speech, wherein the first speaker is determined to be the agent and to be speaking the agent language in response to speech characteristics of the first speech matching the agent voiceprint, and wherein the second speaker is determined to be the customer in response to the speech characteristics not matching the agent voiceprint;

determining, based on further analysis of the audio data, a customer language comprising a language spoken by the customer;

translating, by the processor, the second speech represented by the audio data to the agent language;

outputting, by the processor and at a display of the agent device, first text representing a translation of the second speech in the agent language;

translating, by the processor, the first speech to the customer language; and

outputting, by the processor and at the display of the agent device, second text representing a translation of the first speech.

8 . The method of claim 7 , wherein identifying the first speaker comprises determining that the first speech is spoken in the agent language.

9 . The method of claim 7 , wherein outputting the first text and the second text at the display of the agent device comprises:

generating a user interface comprising the first text oriented in a first orientation and the second text oriented in a second orientation that is upside down relative to the first orientation; and

displaying the user interface at the agent device.

10 . The method of claim 7 , wherein the agent voiceprint is generated at the agent device, and wherein generating the agent voiceprint comprises:

obtaining, the sound hardware, further audio data that represents a voice associated with the agent;

creating, based on the further audio data, the agent voiceprint; and

storing the agent voiceprint with a data association that associates the agent voiceprint with the agent.

11 . The method of claim 10 , wherein the agent voiceprint defines:

a frequency of the voice associated with the agent;

an intonation of the voice associated with the agent; and

a timbre of the voice associated with the agent.

12 . The method of claim 7 , wherein determining the customer language comprises:

determining a language likelihood threshold that is defined for the agent device, the language likelihood threshold comprising a minimum probability that the second speech is spoken in a particular language;

identifying a language that meets the language likelihood threshold; and

outputting language information that identifies the language.

13 . The method of claim 7 , wherein determining the customer language comprises:

determining a language likelihood threshold that is defined for the agent device, the language likelihood threshold comprising a minimum probability that the second speech is spoken in a particular language;

identifying a plurality of languages that meet the language likelihood threshold; and

outputting, on the display of the agent device, a display of a ranked language listing that identifies the plurality of languages in an order from largest probability to smallest probability, wherein a respective identifier of each of the plurality of languages is translated into a corresponding language, and wherein the customer language is selected from the ranked language listing to determine the customer language.

14 . The method of claim 7 , wherein determining the customer language comprises:

determining a first language spoken at a flight origin associated with a flight and a second language spoken at a flight destination associated with the flight,

determining that the second language corresponds to the agent language, and

determining that the first language comprises the customer language.

15 . A system comprising a processor and a memory that stores computer-executable instructions that, when executed by the processor, cause the processor to perform operations comprising:

detecting, at an agent device comprising sound hardware, a conversation between an agent and a customer, wherein data defining an agent language spoken by the agent and an agent voiceprint is stored at the agent device, and wherein the agent voiceprint comprises a unique pattern of voice characteristics of the agent;

capturing, using the sound hardware and during the conversation, audio data associated with the conversation, wherein the audio data comprises audio signals captured using the sound hardware, and wherein the audio signals represent first speech associated with the agent and second speech associated with the customer;

identifying, based on analysis of the audio signals, a first speaker associated with the first speech and a second speaker associated with the second speech, wherein the first speaker is determined to be the agent and to be speaking the agent language in response to speech characteristics of the first speech matching the agent voiceprint, and wherein the second speaker is determined to be the customer in response to the speech characteristics not matching the agent voiceprint;

determining, based on further analysis of the audio data, a customer language comprising a language spoken by the customer;

translating the second speech represented by the audio data to the agent language;

outputting, at a display of the agent device, first text representing a translation of the second speech in the agent language;

translating the first speech to the customer language; and

outputting, at the display of the agent device, second text representing a translation of the first speech.

16 . The system of claim 15 , wherein identifying the first speaker comprises determining that the first speech is spoken in the agent language.

17 . The system of claim 15 , wherein outputting the first text and the second text at the display of the agent device comprises:

generating a user interface comprising the first text oriented in a first orientation and the second text oriented in a second orientation that is upside down relative to the first orientation; and

displaying the user interface at the agent device.

18 . The system of claim 15 , wherein the agent voiceprint is generated at the agent device, and wherein generating the agent voiceprint comprises:

obtaining, the sound hardware, further audio data that represents a voice associated with the agent;

creating, based on the further audio data, the agent voiceprint; and

storing the agent voiceprint with a data association that associates the agent voiceprint with the agent.

19 . The system of claim 15 , wherein determining the customer language comprises:

determining a language likelihood threshold defined for the agent device, the language likelihood threshold comprising a minimum probability that the second speech is spoken in a particular language;

identifying a plurality of languages that meet the language likelihood threshold; and

outputting, on the display of the agent device, a display of a ranked language listing that identifies the plurality of languages in an order from largest probability to smallest probability, wherein a respective identifier of each of the plurality of languages is translated into a corresponding language, and wherein the customer language is selected from the ranked language listing to determine the customer language.

20 . The system of claim 15 , wherein determining the customer language comprises:

determining a first language spoken at a flight origin associated with a flight and a second language spoken at a flight destination associated with the flight,

determining that the second language corresponds to the agent language, and

determining that the first language comprises the customer language.