IP Library › Granted Patent US 12,547,848
Granted Patent B2
US 12,547,848 · App. 18/319,249 · Granted Feb 10, 2026

One-shot visual language reasoning over graphical depictions of data

Inventors: Julian Martin Eisenschlos (Zürich, CH); Francesco Piccinno (Zürich, CH); Yasemin Altun (Zürich, CH); Syrine Krichene (Zürich, CH); Kenton Chiu Tsun Lee (Kirkland, WA); Fangyu Liu (Cambridge, GB); Mandar Joshi (Seattle, WA); Chenxi Pang (Zürich, CH); Wenhu Chen (Waterloo, CA)
Assignee: GOOGLE LLC
G06F40/40G06T11/206G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,547,848
App. No.
18/319,249
Granted
Feb 10, 2026
Kind
B2
Abstract

Provided is a one-shot solution to visual language reasoning. Example systems described herein decompose the challenge of visual language reasoning into two steps: translation of a graphical depiction of data (e.g., a plot or chart) into text; followed by reasoning over the translated text. In particular, example systems described herein can include a machine-learned visual-to-language conversion model that translates a graphical depiction of a dataset to a set of text descriptive of the dataset. The output of visual-to-language conversion model can then be directly used to prompt a language model, (e.g., a pretrained large language model (LLM)), exploiting the few-shot reasoning capabilities of the language model.

Claims (45)

1 . A computer-implemented method to process graphical depictions of data, the method comprising:

obtaining, by a computing system comprising one or more computing devices, an input comprising a graphical depiction of a dataset;

processing, by the computing system, the graphical depiction of the dataset with a machine-learned visual-to-language conversion model to generate, as an output of the machine-learned visual-to-language conversion model a set of text descriptive of the dataset, the machine-learned visual-to-language conversion model trained using a loss function that measures a relative mapping similarity, the relative mapping similarity comprising a similarity between predicted headers and training headers;

processing, by the computing system, the set of text descriptive of the dataset with a machine-learned language model to generate, as an output of the machine-learned language model, a textual output; and

providing, by the computing system, the textual output as an output.

2 . The computer-implemented method of claim 1 , wherein the graphical depiction of the dataset comprises an image that depicts a chart.

3 . The computer-implemented method of claim 1 , wherein the graphical depiction of the dataset comprises an image that depicts a plot.

4 . The computer-implemented method of claim 1 , wherein the set of text descriptive of the dataset comprises a linearized table.

5 . The computer-implemented method of claim 1 , wherein:

the input further comprises a natural language query;

processing, by the computing system, the set of text descriptive of the dataset with the machine-learned language model comprises jointly processing, by the computing system, the set of text descriptive of the dataset and the natural language query with the machine-learned language model to generate, as the output of the machine-learned language model, the textual output; and

the textual output comprises a textual response that is responsive to the natural language query.

6 . The computer-implemented method of claim 1 , wherein the machine-learned visual-to-language conversion model has been trained separately from the machine-learned language model.

7 . The computer-implemented method of claim 1 , wherein:

the machine-learned visual-to-language conversion model has been trained using a set of supervised training data comprising a plurality of training pairs, each training pair comprising a training graphical depiction of a dataset and a training textual description of the dataset; and

the machine-learned visual-to-language conversion model has been trained to predict the training textual description of the dataset.

8 . The computer-implemented method of claim 7 , wherein the loss function measures the relative mapping similarity between predicted tuples and training tuples, wherein each of the predicted tuples and training tuples comprises a row header, a column header, and a value.

9 . The computer-implemented method of claim 1 , wherein:

the machine-learned visual-to-language conversion model has been trained using a set of supervised training data comprising a plurality of training pairs, each training pair comprising a training graphical depiction of a dataset and training rendering code, wherein the training rendering code comprises rendering code to render the training graphical depiction of a dataset; and

the machine-learned visual-to-language conversion model has been trained to predict the training rendering code.

10 . The computer-implemented method of claim 1 , wherein the machine-learned visual-to-language conversion model has been trained using a math dataset comprising textual math problem inputs rendered as images.

11 . A computing system configured to process graphical depictions of data, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that collectively store:

a machine-learned visual-to-language conversion model configured to convert graphical depictions of data to textual descriptions;

a machine-learned language model configured to process textual input to generate textual output; and

instructions that, when executed by the computing system, cause the computing system to perform operations, the operations comprising:

obtaining, by the computing system, an input comprising a graphical depiction of a dataset;

processing, by the computing system, the graphical depiction of the dataset with a machine-learned visual-to-language conversion model to generate, as an output of the machine-learned visual-to-language conversion model a set of text descriptive of the dataset, the machine-learned visual-to-language conversion model trained using a loss function that measures a relative mapping similarity between predicted headers and training headers; and

processing, by the computing system, the set of text descriptive of the dataset with a machine-learned language model to generate, as an output of the machine-learned language model, a textual output.

12 . The computing system of claim 11 , wherein the graphical depiction of the dataset comprises an image that depicts a chart.

13 . The computing system of claim 11 , wherein the graphical depiction of the dataset comprises an image that depicts a plot.

14 . The computing system of claim 11 , wherein the set of text descriptive of the dataset comprises a linearized table.

15 . The computing system of claim 11 , wherein:

the input further comprises a natural language query;

processing, by the computing system, the set of text descriptive of the dataset with the machine-learned language model comprises jointly processing, by the computing system, the set of text descriptive of the dataset and the natural language query with the machine-learned language model to generate, as the output of the machine-learned language model, the textual output; and

the textual output comprises a textual response that is responsive to the natural language query.

16 . One or more non-transitory computer-readable media that collectively store instructions that, when executed by a computing system, cause the computing system to perform operations, the operations comprising:

obtaining a visual-to-language conversion model;

pre-training the visual-to-language conversion model using one or more pre-training tasks wherein at least one pre-training task of the one or more pre-training tasks comprises a math reasoning task;

after said pre-training, fine-tuning the visual-to-language conversion model on a fine-tuning task, wherein the fine-tuning task comprises converting a graphical depiction of a dataset to a textual description of the dataset; and

after said fine-tuning, deploying the visual-to-language conversion model in combination with a machine-learned language model to perform processing of graphical depictions of data wherein processing the graphical depictions of data comprise rendering a text-based numerical reasoning input as an image and decoding an answer.

17 . The one or more non-transitory computer-readable media of claim 16 , wherein another pre-training task of the one or more pre-training tasks comprises a chart de-rendering task.

18 . The one or more non-transitory computer-readable media of claim 17 , wherein the chart de-rendering task comprises converting a graphical depiction of a dataset to rendering code that, when executed, causes rendering of the graphical depiction of the dataset.

19 . The one or more non-transitory computer-readable media of claim 16 , wherein the math reasoning task comprises processing a textual math reasoning input that has been rendered as an image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2023
From: CHEN, WENHU
To: GOOGLE LLC
Reel/Frame 063871/0986 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: EISENSCHLOS, JULIAN MARTIN; PICCINNO, FRANCESCO; ALTUN, YASEMIN; KRICHENE, SYRINE; LEE, KENTON CHIU TSUN; LIU, FANGYU; JOSHI, MANDAR; PANG, CHENXI
To: GOOGLE LLC
Reel/Frame 063833/0205 →
Continuity (1)
Related Publication 20240386215A1 · Nov 21, 2024
References Cited (43)
US 20120213429A1 · Vasudevan · 2012 [cited by examiner]
US 20220188564A1 · Gudimetla · 2022 [cited by examiner]
US 20240320421A1 · Bursztyn · 2024 [cited by examiner]
US 20240362405A1 · Sultanum · 2024 [cited by examiner]
US 20240386058A1 · Thomas · 2024 [cited by examiner]
Masry et al., “ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning”, Mar. 19, 2022, ACL 22. (Year: 2022). [cited by examiner]
Akhtar et al., “Reading and Reasoning Over Chart Images for Evidence-Based Automated Fact-Checking.”, arXiv:2301.11843v1, Jan. 27, 2023. [cited by applicant]
Andrejczuk et al., “Table-to-Text Generation and Pre-Training with TabT5.”, arXiv:2210.09162v1, Oct. 17, 2022, 9 pages. [cited by applicant]
Biten et al., “ICDAR 2019 Competition on Scene Text Visual Question Answering”, arXiv:1907.00490v1, Jun. 30, 2019, 8 pages. [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners.”, arXiv:2005.14165v4, Jul. 22, 2020, 75 pages. [cited by applicant]
Chen et al., “Evaluating Large Language Models Trained on Code.”, arXiv:2107.03374v2, Jul. 14, 2021, 35 pages. [cited by applicant]
Chen et al., “PaLI: A Jointly-Scaled Multilingual Language-Image Model.”, arXiv:2209.06794v4, Jun. 5, 2023, 33 pages. [cited by applicant]
Chen et al., “Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks.”, arXiv:2211.12588v3, Nov. 29, 2022, 11 pages. [cited by applicant]
Chen., “Large Language Models are few (1)-shot Table Reasoners.”, arXiv:2210.06710v2, Jan. 23, 2023, 11 pages. [cited by applicant]
Cheng et al., “Binding Language Models in Symbolic Languages.”, arXiv:2210.02875v2, Mar. 1, 2023, 27 pages. [cited by applicant]
Cho et al., “Unifying Vision-and-Language Tasks via Text Generation.”, arXiv:2102.02779v2, May 23, 2021, 15 pages. [cited by applicant]
Chowdhery et al., “PaLM: Scaling Language Modeling with Pathways.”, arXiv:2204.02311v5, Oct. 5, 2022, 87 pages. [cited by applicant]
Chung et al., “Scaling Instruction-Finetuned Language Models.”, arXiv:2210.11416v5, Dec. 6, 2022, 54 pages. [cited by applicant]
Dosovitskiy et al., “An Image is Worth 16×16 Words: Transformers for Image Recognition at Scale.”, arXiv:2010.11929v2, Jun. 3, 2021, 22 pages. [cited by applicant]
Gao et al., “PAL: Program-aided Language Models.”, arXiv:2211.10435v2, Jan. 27, 2023. [cited by applicant]
Gehrmann et al., “TaTa: A Multilingual Table-to-Text Dataset for African Languages.”, arXiv:2211.00142v1, Oct. 31, 2022, 24 pages. [cited by applicant]
Herzig et al., “TAPAS: Weakly Supervised Table Parsing via Pre-training.”, arXiv:2004.02349v2, Apr. 21, 2020, 14 pages. [cited by applicant]
Kantharaj et al., “Chart-to-Text: A Large-Scale Benchmark for Chart Summarization.” arXiv:2203.06486v3, Apr. 14, 2022, 19 pages. [cited by applicant]
Kato et al., “Parsing Line Chart Images Using Linear Programming.”, 2022 IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), Waikoloa, Hawaii, United States, Jan. 4-8, 2022, pp. 2553-2562. [cited by applicant]
Lee et al., “Pix2Struct: Screenshot Parsing as Pretraining for Visual Language Understanding.”, arXiv:2210.03347v2, Jun. 15, 2023, 20 pages. [cited by applicant]
Levy et al., “Classification-Regression for Chart Comprehension.”, arXiv:2111.14792v2, Jul. 11, 2022, 27 pages. [cited by applicant]
Liu et al., “DEPLOT: One-shot Visual Language Reasoning by Plot-to-Table Translation.”, arXiv:2212.10505v2, May 23, 2023, 17 pages. [cited by applicant]
Liu et al., “MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering.”, arXiv:2212.09662v2, May 23, 2023, 13 pages. [cited by applicant]
Liu et al., “Mind's eye: Grounded Language Model Reasoning through Simulation.”, arXiv:2210.05359v1, Oct. 11, 2022, 18 pages. [cited by applicant]
Luo et al., “ChartOCR: Data Extraction from Charts Images via a Deep Hybrid Framework.”, 2021 Institute of Electrical and Electronics Engineers Winter Conference on Applications of Computer Vision (WACV), Virtual, Jan. … [cited by applicant]
Masry et al., “ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.”, arXiv:2203.10244v1, Mar. 19, 2022, 17 pages. [cited by applicant]
Methani et al., “PlotQA: Reasoning over Scientific Plots.”, arXiv:1909.00997v3, Feb. 1, 2020, 18 pages. [cited by applicant]
OpenAI, “GPT-4 Technical Report.”, arXiv:2303.08774v3, Mar. 27, 2023, 100 pages. [cited by applicant]
Radford et al., “Learning Transferable Visual Models from Natural Language Supervision.”, arXiv:2103.00020v1, Feb. 26, 2021, 48 pages. [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.”, arXiv:1910.10683v4, Sep. 19, 2023, 67 pages. [cited by applicant]
Rane et al., “ChartReader: Automatic Parsing of Bar-Plots.”, 2021 Institute of Electrical and Electronics Engineers 22nd International Conference on Information Reuse and Integration for Data Science (IRI), Las Vegas, N… [cited by applicant]
Siegel et al., “FigureSeer: Parsing Result-Figures in Research Papers.”, Fourteenth European Conference, Amsterdam, The Netherlands, Oct. 11-14, 2016, pp. 664-680. [cited by applicant]
Su et al., “Language Models Can See: Plugging Visual Controls in Text Generation.”, arXiv:2205.02655v2, May 30, 2022, 21 pages. [cited by applicant]
Wang et al., “Self-Consistency Improves Chain of Thought Reasoning in Language Models.”, arXiv:2203.11171v4, Mar. 7, 2023, 24 pages. [cited by applicant]
Wei et al., “Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.”, arXiv:2201.11903v6, Jan. 10, 2023, 43 pages. [cited by applicant]
Yang et al., “An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQA.”, arXiv:2109.05014v2, Sep. 14, 2022, 10 pages. [cited by applicant]
Yin et al., “TaBERT: Pretraining for Joint Understanding of Textual and Tabular Data.”, arXiv:2005.08314v1, May 17, 2020, 15 pages. [cited by applicant]
Zeng et al., “Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.”, arXiv:2204.00598v2, May 27, 2022, 30 pages. [cited by applicant]