IP Library Granted Patent US 12,499,332
Granted Patent B2
US 12,499,332 · App. 17/954,845 · Granted Dec 16, 2025

Translating text using generated visual representations and artificial intelligence

Inventors: Rameswar Panda (Cambridge, MA); Yi Li (Cambridge, MA); Richard Chen (Yorktown Heights, NY); Rogerio Schmidt Feris (Yorktown Heights, NY); Yoon Hyung Kim (Cambridge, MA); David Cox (Cambridge, MA)
Assignees: International Business Machines Corporation; Massachusetts Institute of Technology
G06F40/58G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,499,332
App. No.
17/954,845
Granted
Dec 16, 2025
Kind
B2
Abstract

Methods, systems, and computer program products for translating text using generated visual representations and artificial intelligence are provided herein. A computer-implemented method includes generating a tokenized form of at least a portion of input text in a first language; generating at least one visual representation of at least a portion of the input text using a first set of artificial intelligence techniques; generating a tokenized form of at least a portion of the at least one visual representation; and generating an output including a translated version of the input text into at least a second language by processing, using a second set of artificial intelligence techniques, at least a portion of the tokenized form of the at least a portion of the input text and at least a portion of the tokenized form of the at least a portion of the at least one visual representation.

Claims (46)

1 . A computer-implemented method comprising:

generating a tokenized form of at least a portion of input text, wherein the input text is in a first language;

generating, as output from utilizing a first set of one or more artificial intelligence techniques, at least one image representation of at least a portion of the input text by mapping portions of stored image data, the portions of the stored image data selected in connection with processing the input text using the first set of one or more artificial intelligence techniques, into portions of the at least one image representation of the at least a portion of the input text, wherein the first set of one or more artificial intelligence techniques comprises at least one neural network-based autoregressive transformer trained on the stored image data;

generating a tokenized form of at least a portion of the at least one image representation;

generating an output comprising a translated version of the input text into at least a second language by processing, using a second set of one or more artificial intelligence techniques, at least a portion of the tokenized form of the at least a portion of the input text and at least a portion of the tokenized form of the at least a portion of the at least one image representation, wherein the second set of one or more artificial intelligence techniques comprises at least one neural network-based multimodal translation transformer;

automatically training, using feedback related to the generated output, at least one of the at least one neural network-based autoregressive transformer and the at least one neural network- based multimodal translation transformer; and

automatically executing, subsequent to the automatic training, one or more machine translation operations using the at least one of the at least one neural network-based autoregressive transformer and the at least one neural network-based multimodal translation transformer;

wherein the method is carried out by at least one computing device.

2 . The computer-implemented method of claim 1 , further comprising:

automatically training the first set of one or more artificial intelligence techniques by processing a training set of text data and processing image data corresponding to the training set of text data.

3 . The computer-implemented method of claim 2 , wherein automatically training the first set of one or more artificial intelligence techniques comprises using visual representation loss-related techniques.

4 . The computer-implemented method of claim 2 , further comprising:

automatically training the second set of one or more artificial intelligence techniques using (i) a tokenized combination of a visual representation of the training set of text data and the training set of text data, and (ii) a tokenized combination of the image data and the training set of text data.

5 . The computer-implemented method of claim 4 , wherein automatically training the second set of one or more artificial intelligence techniques comprises using translation loss-related techniques and consistency loss-related techniques.

6 . The computer-implemented method of claim 1 , wherein generating the output comprises mapping, using the second set of one or more artificial intelligence techniques, one or more portions of the tokenized form of at least a portion of input text to one or more portions of the tokenized form of at least a portion of the at least one image representation.

7 . The computer-implemented method of claim 1 , wherein software implementing the method is provided as a service in a cloud environment.

8 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computing device to cause the computing device to:

generate a tokenized form of at least a portion of input text, wherein the input text is in a first language;

generate, as output from utilizing a first set of one or more artificial intelligence techniques, at least one image representation of at least a portion of the input text by mapping portions of stored image data, the portions of the stored image data selected in connection with processing the input text using the first set of one or more artificial intelligence techniques, into portions of the at least one image representation of the at least a portion of the input text, wherein the first set of one or more artificial intelligence techniques comprises at least one neural network-based autoregressive transformer trained on the stored image data;

generate a tokenized form of at least a portion of the at least one image representation;

generate an output comprising a translated version of the input text into at least a second language by processing, using a second set of one or more artificial intelligence techniques, at least a portion of the tokenized form of the at least a portion of the input text and at least a portion of the tokenized form of the at least a portion of the at least one image representation, wherein the second set of one or more artificial intelligence techniques comprises at least one neural network-based multimodal translation transformer;

automatically train, using feedback related to the generated output, at least one of the at least one neural network-based autoregressive transformer and the at least one neural network- based multimodal translation transformer; and

automatically execute, subsequent to the automatic training, one or more machine translation operations using the at least one of the at least one neural network-based autoregressive transformer and the at least one neural network-based multimodal translation transformer.

9 . The computer program product of claim 8 , wherein the program instructions is further executable by the computing device to cause the computing device to:

automatically train the first set of one or more artificial intelligence techniques by processing a training set of text data and processing image data corresponding to the training set of text data.

10 . The computer program product of claim 9 , wherein the program instructions is further executable by the computing device to cause the computing device to:

automatically train the second set of one or more artificial intelligence techniques using (i) a tokenized combination of a visual representation of the training set of text data and the training set of text data, and (ii) a tokenized combination of the image data and the training set of text data.

11 . The computer program product of claim 9 , wherein automatically training the first set of one or more artificial intelligence techniques comprises using visual representation loss-related techniques.

12 . The computer program product of claim 10 , wherein automatically training the second set of one or more artificial intelligence techniques comprises using translation loss- related techniques and consistency loss-related techniques.

13 . The computer program product of claim 8 , wherein generating the output comprises mapping, using the second set of one or more artificial intelligence techniques, one or more portions of the tokenized form of at least a portion of input text to one or more portions of the tokenized form of at least a portion of the at least one image representation.

14 . A system comprising:

a memory configured to store program instructions; and

a processor operatively coupled to the memory to execute the program instructions to:

generate a tokenized form of at least a portion of input text, wherein the input text is in a first language;

generate, as output from utilizing a first set of one or more artificial intelligence techniques, at least one image representation of at least a portion of the input text by mapping portions of stored image data, the portions of the stored image data selected in connection with processing the input text using the first set of one or more artificial intelligence techniques, into portions of the at least one image representation of the at least a portion of the input text, wherein the first set of one or more artificial intelligence techniques comprises at least one neural network-based autoregressive transformer trained on the stored image data;

generate a tokenized form of at least a portion of the at least one image representation;

generate an output comprising a translated version of the input text into at least a second language by processing, using a second set of one or more artificial intelligence techniques, at least a portion of the tokenized form of the at least a portion of the input text and at least a portion of the tokenized form of the at least a portion of the at least one image representation, wherein the second set of one or more artificial intelligence techniques comprises at least one neural network-based multimodal translation transformer

automatically train, using feedback related to the generated output, at least one of the at least one neural network-based autoregressive transformer and the at least one neural network-based multimodal translation transformer; and

automatically execute, subsequent to the automatic training, one or more machine translation operations using the at least one of the at least one neural network-based autoregressive transformer and the at least one neural network-based multimodal translation transformer.

15 . The system of claim 14 , wherein the processor is further operatively coupled to the memory to execute the program instructions to:

automatically train the first set of one or more artificial intelligence techniques by processing a training set of text data and processing image data corresponding to the training set of text data.

16 . The system of claim 15 , wherein automatically training the first set of one or more artificial intelligence techniques comprises using visual representation loss-related techniques.

17 . The system of claim 15 , wherein the processor is further operatively coupled to the memory to execute the program instructions to:

automatically train the second set of one or more artificial intelligence techniques using (i) a tokenized combination of a visual representation of the training set of text data and the training set of text data, and (ii) a tokenized combination of the image data and the training set of text data.

18 . The system of claim 17 , wherein automatically training the second set of one or more artificial intelligence techniques comprises using translation loss-related techniques and consistency loss-related techniques.

19 . The system of claim 14 , wherein generating the output comprises mapping, using the second set of one or more artificial intelligence techniques, one or more portions of the tokenized form of at least a portion of input text to one or more portions of the tokenized form of at least a portion of the at least one image representation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2023
From: KIM, YOON HYUNG
To: MASSACHUSETTS INSTITUTE OF TECHNOLOGY
Reel/Frame 063849/0750 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: PANDA, RAMESWAR; LI, YI; CHEN, RICHARD; FERIS, ROGERIO SCHMIDT; COX, DAVID
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061244/0036 →
Continuity (1)
Related Publication 20240127005A1 · Apr 18, 2024
References Cited (25)
US 10902952B2 · Lucas et al. · 2021 [cited by applicant]
US 20180137551A1 · Zheng · 2018 [cited by examiner]
US 20190191044A1 · Ghodke · 2019 [cited by examiner]
US 20200126663A1 · Lucas et al. · 2020 [cited by applicant]
US 20210272341A1 · Swaminathan · 2021 [cited by examiner]
US 20220108685A1 · Petrov et al. · 2022 [cited by applicant]
US 20240104312A1 · Stone · 2024 [cited by examiner]
US 20240338535A1 · Sung · 2024 [cited by examiner]
CN 112257459A · 2021 [cited by examiner]
CN 112257465A · 2021 [cited by examiner]
Chen, Jiacheng, et al. “Learning the best pooling strategy for visual semantic embedding.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2021. https://arxiv.org/abs/2011.04305 (Year:… [cited by examiner]
Li et al., VALHALLA: Visual Hallucination for Machine Translation, May 31, 2022. [cited by applicant]
Li et al., VALHALLA: Visual Hallucination for Machine Translation Supplemental Material, 2022. [cited by applicant]
Raunak et al., The Curious Case of Hallucinations in Neural Network Machine Translation, Jun. 2021. [cited by applicant]
Zhou et al., Detecting Hallucinated Content in Conditional Neural Sequence Generation, Jun. 2021. [cited by applicant]
IP.com, IPCOM000265154D, System and Method for Automated Evaluation and Reconciliation of Regulatory Texts, Mar. 2021. [cited by applicant]
IP.com, IPCOM000262813D, Al Assistance Interaction with Visual Simulation on Edge Devices, Jul. 2020. [cited by applicant]
IP.com, IPCOM000259537D, Automatic Acquisition of Large Annotated Training Corpora for Automated Test Generation and Summarization, Aug. 2019. [cited by applicant]
Nyberg et al., The KANT system: Fast, accurate, high-quality translation in practical domains. In COLING 1992 vol. 3: The 14th International Conference on Computational Linguistics, 1992. [cited by applicant]
Lopez, A. Statistical machine translation. ACM Computing Surveys (CSUR), 40(3):1-49, 2008. [cited by applicant]
Bahdanau et al., Neural machine translation by jointly learning to align and translate. arXiv preprint arXiv:1409.0473, 2014. [cited by applicant]
Sutskever et al., Sequence to Sequence Learning with Neural Networks. In Proceedings ofNeurIPS, 2014. [cited by applicant]
Calixto et al., Doubly-attentive decoder for multi-modal neural machine translation. arXivpreprint arXiv:1702.01287, 2017. [cited by applicant]
In et al., Dynamic context-guided capsule network for multimodal machine translation. In Proceedings of the 28th ACM International Conference onMultimedia, pp. 1320-1329, 2020. [cited by applicant]
Ive et al., Distilling translations with visual awareness. arXiv preprintarXiv:1906.07701, 2019. [cited by applicant]