IP Library › Granted Patent US 12,394,235
Granted Patent B2
US 12,394,235 · App. 18/071,371 · Granted Aug 19, 2025

Language-agnostic OCR extraction

Inventors: Osaid Rehman Nasir (New Delhi, IN); Bharat Kumar Jain (Hyderabad, IN); Smitkumar Narotambhai Marvaniya (Bangalore, IN)
Assignee: Microsoft Technology Licensing, LLC
G06V30/2528G06F40/40G06V30/1448G06V30/19147G06V30/274
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,394,235
App. No.
18/071,371
Filed
Nov 29, 2022
Granted
Aug 19, 2025
Kind
B2
Examiner
CRUZ, IRIANA
Art Unit
2681
USPC
382/155
Abstract

Technologies for language agnostic OCR extraction include identifying a word region of an image using optical character recognition, applying a language agnostic machine learning model to the word region, where the language agnostic machine learning model is trained on training data including a set of image-text pairs and a set of multilingual text translation pairs, receiving, from the language agnostic machine learning model, a word region embedding that is associated with the word region, searching a multilingual index for a text embedding that matches the word region embedding, receiving, from the multilingual index, text associated with the text embedding; and outputting at least one of the text or the text embedding to at least one downstream process, application, system, component, or network.

Claims (64)

1. A method, comprising:

identifying a word region of an image using optical character recognition;

the word region comprises a set of bounding box coordinates;

applying a language agnostic machine learning model to the word region;

the language agnostic machine learning model is trained on training data comprising a set of image-text pairs and a set of multilingual text translation pairs;

receiving, from the language agnostic machine learning model, a word region embedding that is associated with the word region;

searching a multilingual index for a text embedding that matches the word region embedding;

receiving, from the multilingual index, text associated with the text embedding; and

using at least one of the text or the text embedding to determine whether or how to display the image via a user interface at a device.

2. The method of claim 1 , further comprising outputting the at least one of the text or the text embedding to at least one of:

a content classification model;

a content ranking model;

a data storage device; or

an output device.

3. The method of claim 2 , further comprising:

based on the text, generating and outputting a caption for the image.

4. The method of claim 1 , further comprising creating the multilingual index by:

applying the language agnostic machine learning model to a multilingual vocabulary;

the multilingual vocabulary comprises a plurality of different words in a plurality of different languages;

receiving, from the language agnostic machine learning model, a set of word embeddings;

the set of word embeddings comprises a word embedding for each word in the multilingual vocabulary; and

indexing the set of word embeddings.

5. The method of claim 4 , further comprising:

using a nearest neighbor algorithm to index the set of word embeddings.

6. The method of claim 4 , further comprising:

creating the multilingual vocabulary based on words extracted from a particular online system.

7. The method of claim 1 , wherein the language agnostic machine learning model embeds both natural language texts and images in a same latent space.

8. The method of claim 1 , wherein the language agnostic machine learning model comprises at least one of a multimodal representation model or a Turing Bletchley model.

9. The method of claim 1 , wherein the word region embedding comprises an image embedding, the image embedding is used to perform a search of the multilingual index, and the text embedding is returned by the search.

10. The method of claim 1 , wherein the word region embedding comprises an image embedding generated based on the image, and the method further comprises:

computing geo-relation data between the image embedding and a text embedding determined based on the image embedding; and

generating a caption for the image based on the geo-relation data.

11. A system comprising:

at least one memory; and

at least one processor coupled to the at least one memory;

wherein the at least one memory comprises instructions that, when executed by the at least one processor cause the at least one processor to perform operations comprising:

identifying a word region of an image using optical character recognition;

the word region comprises a set of bounding box coordinates;

applying a language agnostic machine learning model to the word region;

the language agnostic machine learning model is trained on training data comprising a set of image-text pairs and a set of multilingual text translation pairs;

receiving, from the language agnostic machine learning model, a word region embedding that is associated with the word region;

searching a multilingual index for a text embedding that matches the word region embedding;

receiving, from the multilingual index, text associated with the text embedding; and

using at least one of the text or the text embedding to determine whether or how to display the image via a user interface at a device.

12. The system of claim 11 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

outputting the at least one of the text or the text embedding to at least one of: a content classification model; a content ranking model; a data storage device; or an output device.

13. The system of claim 12 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

based on the text, generating and outputting a caption for the image.

14. The system of claim 11 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising creating the multilingual index by:

applying the language agnostic machine learning model to a multilingual vocabulary;

the multilingual vocabulary comprises a plurality of different words in a plurality of different languages;

receiving, from the language agnostic machine learning model, a set of word embeddings;

the set of word embeddings comprises a word embedding for each word in the multilingual vocabulary; and

indexing the set of word embeddings.

15. The system of claim 14 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

using a nearest neighbor algorithm to index the set of word embeddings.

16. The system of claim 14 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

creating the multilingual vocabulary based on words extracted from a particular online system.

17. The system of claim 11 , wherein the language agnostic machine learning model embeds both natural language texts and images in a same latent space.

18. The system of claim 11 , wherein the language agnostic machine learning model comprises at least one of a multimodal representation model or a Turing Bletchley model.

19. The system of claim 11 , wherein the word region embedding comprises an image embedding, the image embedding is used to perform a search of the multilingual index, and the text embedding is returned by the search.

20. The system of claim 11 , wherein the word region embedding comprises an image embedding generated based on the image, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations further comprising:

computing geo-relation data between the image embedding and a text embedding determined based on the image embedding; and

generating a caption for the image based on the geo-relation data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2022
From: NASIR, OSAID REHMAN; JAIN, BHARAT KUMAR; MARVANIYA, SMITKUMAR NAROTAMBHAI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062072/0772 →
Continuity (1)
Related Publication 20240177513A1 · May 30, 2024
References Cited (18)
US 10936897B2 · Vig · 2021 [cited by examiner]
US 11699275B2 · Kurma · 2023 [cited by examiner]
US 20030200078A1 · Luo et al. · 2003 [cited by applicant]
US 20040260535A1 · Chen et al. · 2004 [cited by applicant]
US 20130039570A1 · Vincent · 2013 [cited by examiner]
US 20160162467A1 · Munro · 2016 [cited by examiner]
US 20160203124A1 · Cuthbert · 2016 [cited by examiner]
US 20160350288A1 · Wick · 2016 [cited by examiner]
US 20190197119A1 · Zhang · 2019 [cited by examiner]
US 20200387677A1 · Kim · 2020 [cited by examiner]
US 20220138439A1 · Tambi et al. · 2022 [cited by applicant]
US 20220350998A1 · Desai · 2022 [cited by examiner]
US 20230016729A1 · Pouran Ben Veyseh · 2023 [cited by examiner]
US 20230073775A1 · Goldstein · 2023 [cited by examiner]
US 20240054294A1 · Sikka · 2024 [cited by examiner]
US 20250036877A1 · Rhatigan · 2025 [cited by examiner]
Tiwary, Saurabh, “Turing Bletchley: A Universal Image Language Representation model by Microsoft”, Retrieved from: https://www.microsoft.com/en-us/research/blog/turing-bletchley-a-universal-image-language-representation… [cited by applicant]
Wang, et al., “Improving OCR-Based Image Captioning by Incorporating Geometrical Relationship”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 19, 2021, pp. 1306-1315. [cited by applicant]
Cited By (1)
US 12,675,652