IP Library › Granted Patent US 12,651,592
Granted Patent B2
US 12,651,592 · App. 18/199,270 · Granted Jun 9, 2026

System, method, and computer program for real-time language translation using generative artificial intelligence

Inventor: Jean-marc Eric Ohayon (Givat Shmuel, IL)
Assignee: AMDOCS DEVELOPMENT LIMITED
G10L13/086G10L15/005G10L15/01G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,592
App. No.
18/199,270
Granted
Jun 9, 2026
Kind
B2
Abstract

As described herein, a system, method, and computer program provide real-time language translation using generative artificial intelligence. An input in a first spoken language is received. The input in the first spoken language is processed, using a generative artificial intelligence model, to generate a translation of the input in a second spoken language. The translation is output.

Claims (41)

1 . A non-transitory computer-readable media storing computer instructions which when executed by one or more processors of a device cause the device to:

receive, via an application programming interface (API) enabling communications between an application and a language translation service, an input that includes:

audio of a speech that has been generated in a first spoken language by a chatbot process of the application, and

an indication of a second spoken language to which the speech is to be translated;

process the input, using a generative artificial intelligence model of the language translation service, to generate in real time a translation of the speech in the second spoken language; and

output, by the language translation service to the application, an audio of the translation in the second spoken language to cause the chatbot process to operate as a multilingual chatbot that communicates in real time with a customer in the second spoken language.

2 . The non-transitory computer-readable media of claim 1 , wherein the device is further caused to:

generate a training data set for use in training the generative artificial intelligence model.

3 . The non-transitory computer-readable media of claim 2 , wherein the training data set includes multilingual text and speech data in a standardized format.

4 . The non-transitory computer-readable media of claim 3 , wherein the multilingual text and speech data is collected from at least one of:

public databases,

online corpora, or

user generated content.

5 . The non-transitory computer-readable media of claim 2 , wherein the device is further caused to:

train the generative artificial intelligence model using the training data set.

6 . The non-transitory computer-readable media of claim 1 , wherein the generative artificial intelligence model is a transformer model.

7 . The non-transitory computer-readable media of claim 6 , wherein the generative artificial intelligence model is a Long Short-Term Memory network (LSTM).

8 . The non-transitory computer-readable media of claim 1 , wherein the generative artificial intelligence model is a Convolutional Neural Network (CNN).

9 . The non-transitory computer-readable media of claim 1 , wherein the generative artificial intelligence model is an Encoder-Decoder model.

10 . The non-transitory computer-readable media of claim 1 , wherein the generative artificial intelligence model is a neural machine translation (NMT) model.

11 . The non-transitory computer-readable media of claim 1 , wherein the receiving, processing, and outputting is performed in real-time.

12 . The non-transitory computer-readable media of claim 1 , wherein the device is further caused to:

test the translation with real-world data, and

update the generative artificial intelligence model based on a result of the testing.

13 . The non-transitory computer-readable media of claim 12 , wherein the real-world data includes feedback from a user regarding the translation.

14 . The non-transitory computer-readable media of claim 13 , wherein the generative artificial intelligence model is updated using reinforcement learning.

15 . A method, comprising:

at a computer system:

receiving, via an application programming interface (API) enabling communications between an application and a language translation service, an input that includes:

audio of a speech that has been generated in a first spoken language by a chatbot process of the application, and

an indication of a second spoken language to which the speech is to be translated;

processing the input, using a generative artificial intelligence model of the language translation service, to generate in real time a translation of the speech in the second spoken language; and

outputting, by the language translation service to the application, an audio of the translation in the second spoken language to cause the chatbot process to operate as a multilingual chatbot that communicates in real time with a customer in the second spoken language.

16 . A system, comprising:

a non-transitory memory storing instructions; and

one or more processors in communication with the non-transitory memory that execute the instructions to:

receive, via an application programming interface (API) enabling communications between an application and a language translation service, an input that includes:

audio of a speech that has been generated in a first spoken language by a chatbot process of the application, and

an indication of a second spoken language to which the speech is to be translated;

process the input, using a generative artificial intelligence model of the language translation service, to generate in real time a translation of the speech in the second spoken language; and

output, by the language translation service to the application, an audio of the translation in the second spoken language to cause the chatbot process to operate as a multilingual chatbot that communicates in real time with a customer in the second spoken language.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2023
From: OHAYON, JEAN-MARC ERIC
To: AMDOCS DEVELOPMENT LIMITED
Reel/Frame 064379/0834 →
Continuity (1)
Related Publication 20240386879A1 · Nov 21, 2024
References Cited (15)
US 12079587B1 · Radford · 2024 [cited by examiner]
US 12277927B2 · Li · 2025 [cited by examiner]
US 20200044999A1 · Wu · 2020 [cited by examiner]
US 20200184158A1 · Kuczmarski · 2020 [cited by examiner]
US 20210390268A1 · Pandey · 2021 [cited by examiner]
US 20220147721A1 · Galle · 2022 [cited by examiner]
US 20230267285A1 · Wu · 2023 [cited by examiner]
US 20230342561A1 · Zhang · 2023 [cited by examiner]
International Search Report and Written Opinion from PCT Application No. PCT/IB2024/054578, dated Aug. 20, 2024, 11 pages. [cited by applicant]
Dabre et al., “A Survey of Multilingual Neural Machine Translation,” ACM Computing Surveys, vol. 53, Sep. 2020, pp. 1-38. [cited by applicant]
Zhang et al., “Neural Machine Translation: Challenges, Progress and Future,” arXiv, Apr. 2020, 22 pages, retrieved from https://arxiv.org/pdf/2004.05809. [cited by applicant]
Latif et al., “Transformers in Speech Processing: A Survey,” arXiv, 2023, 27 pages, retrieved from https://arxiv.org/pdf/2303.11607. [cited by applicant]
Yuan et al., “A Roadmap for Big Model,” arXiv, 2022, 200 pages, retrieved from https://arxiv.org/pdf/2203.14101v1. [cited by applicant]
Inaguma et al., “Multilingual End-to-End Speech Translation,” Â IEEE Automatic Speech Recognition and Understanding Workshop (ASRU), 2019, pp. 570-577. [cited by applicant]
Sulubacak et al., “Multimodal Machine Translation through Visuals and Speech,” arXiv, 2019, 34 pages, retrived from https://arxiv.org/abs/1911.12798. [cited by applicant]