IP Library Granted Patent US 12,694,234
Granted Patent B2
US 12,694,234 · App. 17/937,250 · Granted Jul 28, 2026

Image-based text translation and presentation

Inventors: Sujith Gunjur Umapathy (Berlin, DE); Nikhil Garg (Berlin, DE); John Mark Hubenthal (Seattle, WA); Jose Luis Baez (Fair Lawn, NJ); Pushpendu Ghosh (Vadodara, IN)
Assignee: Amazon Technologies, Inc.
G06F40/47G06F40/109G06T5/77G06V10/225G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,694,234
App. No.
17/937,250
Granted
Jul 28, 2026
Kind
B2
Abstract

Systems and methods are provided for translation of text in an image, and presentation of a version of the image in which the translated text is displayed a manner consistent with the original image. Text segments are automatically translated from their original source language to a target language. In order to provide presentation of the translated text in a manner that closely matches the source text, various display attributes of the source text are detected (e.g., font size, font color, font style, etc.).

Claims (70)

1 . A system comprising:

an image data store storing a plurality of image files, wherein an image file of the plurality of image files comprises image data, encoded in a pixel-based format, representing an image; and

a server comprising computer-readable memory and one or more processors, wherein the server is configured to:

send a network resource to a client device, wherein the network resource comprises an instruction to display the image based on the image file;

receive, from the client device, a request for a translated version of the image, wherein the request is associated with a target language;

identify first source text in a first portion of image data, wherein the first source text is in a source language, wherein the first portion of image data represents a first two-dimensional region of the image, and wherein the first source text describes a product in a second portion of image data representing a second two-dimensional region of the image;

identify second source text, in the second portion of the image data, to be excluded from translation based on the second source text being located on the product;

translate the first source text to target text in the target language;

determine a font color in which the target text is to be displayed in the translated version, wherein the font color is determined based on color of a first subset of pixels of the image data, the first subset of pixels associated with the first source text;

determine a font size in which the target text is to be displayed in the translated version, wherein the font size is determined based on (a) first quantity of pixels in a first dimension of the first two-dimensional region, (b) a second quantity of pixels in a second dimension of the first two-dimensional region, and (c) a quantity of characters in the target text;

determine a background color of the image based on a clustering of color values associated with a second subset of pixels in a background of the image;

generate translated image data encoded in the pixel-based format, wherein the translated image data represents modified version of the first two-dimensional region comprising the target text formatted for display using the font color and the font size, and wherein a portion of the modified version of the first two-dimensional region comprises the background color of the image in place of at least a portion of the first source text; and

send the translated image data to the client device.

2 . The system of claim 1 , wherein the network resource comprises a second instruction to display the translated version of the image using the translated image data.

3 . The system of claim 1 , wherein to identify the first source text, the server is configured to:

perform optical character recognition on the image data to generate a plurality of text segments, wherein a first text segment comprises:

a first text string within a first subregion of the first two-dimensional region of the image, the first subregion defined by a first set of coordinates; and

a second text string with a second subregion of the first two-dimensional region of the image, the second subregion defined by a second set of coordinates; and

determine that the first source text comprises the first text string and the second text string.

4 . The system of claim 1 , wherein the server is further configured to convert a portion of the first source text associated with a first unit of measure to a portion of the target text associated with a second unit of measure.

5 . A computer-implemented method comprising:

under control of a computing system comprising one or more computer processors configured to execute specific instructions,

identifying first source text in a first portion of image data, wherein the image data is encoded in a pixel-based format, and wherein the first portion of image data represents a first region of an image;

identifying second source text, in a second portion of the image data, to be excluded from translation, wherein the second portion of the image data represents a second region of the image, and wherein the second source text is to be excluded based on a product being recognized in the second region of the image;

translating the first source text in a source language to target text in a target language;

determining, based at least partly on the first portion of the image data, one or more display attributes for display of the target text;

determining a background color of the first region of the image based on a clustering of color values associated with pixels in a background of the first region of the image; and

generating modified image data encoded in the pixel-based format, wherein a portion of the modified image data represents a modified version of the first region of the image comprising the target text formatted for display using the one or more display attributes, and wherein generating the modified image data comprises replacing a font pixel, associated with the first source text in the first region of the image, with a background pixel of the background color.

6 . The computer-implemented method of claim 5 , wherein identifying the first source text comprises:

performing optical character recognition on the image data to generate a plurality of text segments, wherein a first text segment comprises:

a first text string within a first subregion of the first region of the image, the first subregion defined by a first set of coordinates; and

a second text string with a second subregion of the first region of the image, the second subregion defined by a second set of coordinates; and

determining that the first source text comprises the first text string and the second text string.

7 . The computer-implemented method of claim 5 , further comprising:

causing display of a network resource comprising the image and a target language selection control; and

receiving input data representing selection, using the target language selection control, of the target language from a plurality of target languages.

8 . The computer-implemented method of claim 5 , wherein translating the first source text further comprising converting a portion of the first source text associated with a first unit of measure to a portion of the target text associated with a second unit of measure.

9 . The computer-implemented method of claim 5 , wherein determining the one or more display attributes comprises:

determining a first quantity of pixels in a first dimension of a two-dimensional box within which at least a portion of the first source text is identified; and

determining a font size at which the target text is to be displayed based at least partly on the first quantity of pixels.

10 . The computer-implemented method of claim 9 , wherein determining the one or more display attributes further comprises:

determining a second quantity of pixels in a second dimension of the two-dimensional box; and

determining a quantity of characters in the target text, wherein the font size is determined based at least partly on the quantity of characters, the first quantity of pixels, and the second quantity of pixels.

11 . The computer-implemented method of claim 5 , wherein determining the one or more display attributes comprises:

determining a first color of a first subset of pixels in the first region of the image, wherein the first subset of pixels is associated with display of the first source text; and

determining a font color in which the target text is to be displayed based at least partly on the first color.

12 . The computer-implemented method of claim 5 , wherein determining the one or more display attributes comprises determining a font style in which the first source text is presented, wherein the target text is to be displayed based at least partly on the font style.

13 . The computer-implemented method of claim 5 , further comprising:

sending a network resource to a computing device, wherein the network resource comprises an instruction to display the image based on the image data;

sending the image data to the computing device;

receiving a request for the modified image data; and

sending the modified image data to the computing device, wherein the network resource comprises a further instruction to replace display of the image based on the modified image data.

14 . A system comprising computer readable memory and one or more processors, wherein the system is configured to:

identify first source text in a first portion of image data, wherein the image data is encoded in a pixel-based format, and wherein the first portion of image data represents a first region of an image;

identify second source text, in a second portion of the image data, to be excluded from translation, wherein the second portion of the image data represents a second region of the image, and wherein the second source text is to be excluded based on a product being recognized in the second region of the image;

translate the first source text in a source language to target text in a target language;

determine, based at least partly on the first portion of the image data, one or more display attributes for display of the target text;

determine a background color of the first region of the image based on a clustering of color values associated with pixels in a background of the first region of the image; and

generate modified image data encoded in the pixel-based format, wherein a portion of the modified image data represents a modified version of the first region of the image comprising the target text formatted for display using the one or more display attributes, and wherein generating the modified image data comprises replacing a font pixel, associated with the first source text in the first region of the image, with a background pixel of the background color.

15 . The system of claim 14 , wherein to determine the one or more display attributes, the system is further configured to:

determine a first quantity of pixels in a first dimension of a two-dimensional box within which at least a portion of the first source text is identified; and

determine a font size at which the target text is to be displayed based at least partly on the first quantity of pixels.

16 . The system of claim 14 , wherein to determine the one or more display attributes, the system is further configured to:

determine a first color of a first subset of pixels in the first region of the image, wherein the first subset of pixels is associated with display of the first source text; and

determine a font color in which the target text is to be displayed based at least partly on the first color.

17 . The system of claim 14 , wherein to determine the one or more display attributes, the system is further configured to determine a font style in which the first source text is presented, wherein the target text is to be displayed based at least partly on the font style.

18 . The system of claim 14 , wherein the one or more processors are further configured to:

identify one or more strokes of text as distinct from the background or a foreground object; and

based on the one or more strokes of text, generate a map of which pixels in the image data are associated with the first source text.

19 . The system of claim 18 , wherein the one or more processors are further configured to determine the font pixel to be replaced based on the map.