IP Library › Granted Patent US 12,651,113
Granted Patent B1
US 12,651,113 · App. 19/457,074 · Granted Jun 9, 2026

Document intelligence system

Inventors: Pravesh Agrawal (Mountain View, CA); Sumukh Koteshwara Aithal (Mountain View, CA); Jean-Malo Delignon (Paris, FR); Sandeep Subramanian (San Francisco, CA); Guillaume Lample (Paris, FR)
Assignee: Mistral AI
G06F40/169G06F40/103G06F40/177G06V30/18G06V30/191G06V30/414G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,113
App. No.
19/457,074
Granted
Jun 9, 2026
Kind
B1
Abstract

A method may include obtaining, from a user device, an electronic document to be processed by a model. The method may also include obtaining an optical character recognition output of the electronic document. The method may further include obtaining, from the user device, an output configuration. The method may also include generating a first output and a second output based on the electronic document and the optical character recognition output. The first output and the second output may be defined by the output configuration. The method may further include transmitting the first output and the second output to the user device. The method may also include obtaining a feedback from the user device. The method may further include updating the model using the feedback from the user device.

Claims (40)

1 . A system, comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the system to:

obtain, from a user device, an electronic document to be processed by a model;

obtain an optical character recognition output of the electronic document;

obtain, from the user device, an output configuration;

generate a first output and a second output based on the electronic document and the optical character recognition output, wherein the first output and the second output are defined by the output configuration;

generate a notification that indicates a difference between the electronic document and the first output or a difference between the electronic document and the second output; and

transmit the first output and the second output to the user device.

2 . The system of claim 1 , wherein the optical character recognition output is obtained from a second model.

3 . The system of claim 2 , wherein the optical character recognition output comprises extracted text and extracted bounding boxes from the electronic document.

4 . The system of claim 1 , wherein the model is a vision language model.

5 . The system of claim 1 , wherein the output configuration comprises a bounding box annotation format and a document format.

6 . The system of claim 1 , wherein the output configuration comprises a schema that defines at least one field to be identified in the electronic document.

7 . The system of claim 6 , the instructions to further cause the system to automatically generate the schema using the model to include at least one of a field name, a field type, a description, and a required option toggle.

8 . The system of claim 6 , the instructions to further cause the system to automatically generate the schema using a second model to include at least one of a field name, a field type, a description, and a required option toggle.

9 . The system of claim 1 , wherein the output configuration comprises an image toggle that directs whether an image from the electronic document is included in the first output or the second output.

10 . The system of claim 1 , wherein the first output is a markdown format and the second output is a structured format.

11 . The system of claim 1 , wherein the first output or the second output are transmitted to the user device to be displayed spatially adjacent to the electronic document.

12 . The system of claim 1 , wherein the electronic document comprises a native format, and the native format is one or more of a PDF, a picture, a scan of a document, and a digital photograph.

13 . The system of claim 1 , wherein in response to the model identifying a URL in the electronic document, including the URL in the first output or the second output.

14 . The system of claim 1 , wherein in response to the model identifying a table in the electronic document, including the table in the first output or the second output in an HTML format.

15 . The system of claim 1 , the instructions to further cause the system to:

obtain a feedback from the user device; and

update the model using the feedback from the user device.

16 . A method, comprising:

obtaining, from a user device, an electronic document to be processed by a model;

obtaining an optical character recognition output of the electronic document;

obtaining, from the user device, an output configuration;

generating a first output and a second output based on the electronic document and the optical character recognition output, wherein the first output and the second output are defined by the output configuration;

generating a notification that indicates a difference between the electronic document and the first output or a difference between the electronic document and the second output; and

transmitting the first output and the second output to the user device.

17 . The method of claim 16 , wherein:

the optical character recognition output is obtained from a second model; and

the optical character recognition output comprises extracted text and extracted bounding boxes from the electronic document.

18 . The method of claim 16 , wherein the output configuration comprises a schema that defines at least one field to be identified in the electronic document.

19 . The method of claim 18 , further comprising automatically generating the schema using the model to include at least one of a field name, a field type, a description, and a required option toggle.

20 . The method of claim 18 , further comprising

obtaining a feedback from the user device; and

updating the model using the feedback from the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2026
From: AGRAWAL, PRAVESH; AITHAL, SUMUKH KOTESHWARA; DELIGNON, JEAN-MALO; SUBRAMANIAN, SANDEEP; LAMPLE, GUILLAUME
To: MISTRAL AI
Reel/Frame 074842/0971 →
Continuity (1)
Provisional Application 63952250 · Dec 31, 2025
References Cited (5)
US 20210406451A1 · Iyer · 2021 [cited by examiner]
US 20250348762A1 · Telang · 2025 [cited by examiner]
US 20260017965A1 · Anthony · 2026 [cited by examiner]
Jose D. Bermudez Castro et al., Improvement Optical Character Recognition for Structured Documents using Generative Adversarial Networks, Sep. 1, 2021, International Conference on Computational Science and Its Applicati… [cited by examiner]
Ido Kissos et al., Image and Text Correction Using Language Models, Apr. 1, 2017, International Workshop on Arabic Script Analysis and Recognition, pp. 158-162 (Year: 2017). [cited by examiner]
Cited By (1)
US 12,731,053