IP Library Granted Patent US 12,596,871
Granted Patent B2
US 12,596,871 · App. 18/354,067 · Granted Apr 7, 2026

Textual encoding and analysis with a large graphical language model

Inventor: Dmitry Valentinovich Kholodkov (Sammamish, WA)
Assignee: MICROSOFT TECHNOLOGY LICENSING, LLC
G06F40/253G06F40/30G06T11/206
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,871
App. No.
18/354,067
Granted
Apr 7, 2026
Kind
B2
Abstract

The techniques discussed herein enhance the operation of content generation and analysis systems. Namely, textual content applications such as technical documentation, creative writing, and content moderation. This is accomplished through generating a graphical representation of a body of text (e.g., a document). The graphical representation can comprise a plurality of nodes representing the words of the document and a plurality of lines that join the nodes representing a level of association between individual words. As such, the graphical representation can capture the semantic and syntactical structure of the associated document while omitting the original textual content. The graphical representation can be subsequently evaluated for complexity based on the density of nodes and lines. Accordingly, the disclosed system can assign a score to a document based on the evaluation of the graphical representation. In addition, various documents can be ranked based on such scores.

Claims (76)

1 . A method comprising:

retrieving a document including a plurality of words;

analyzing the plurality of words included in the document;

generating a graphical representation of the document based on the plurality of words, wherein:

the graphical representation of the document includes a plurality of nodes and a plurality of lines joining the plurality of nodes;

each node of the plurality of nodes represents a word of the plurality of words;

a line joining two nodes represents a level of association between two words represented by the two nodes; and

the plurality of words is omitted from the graphical representation of the document;

determining a number of the plurality of nodes and a number of the plurality of lines included in the graphical representation of the document;

calculating, by utilizing a model without access to the plurality of words included in the document and based on the number of the plurality of nodes and the number of the plurality of lines included in the graphical representation of the document, a numerical evaluation of the graphical representation of the document that quantifies a characteristic of the document;

ranking the graphical representation of the document against a plurality of other graphical representations of other documents based on the numerical evaluation of the graphical representation of the document and a plurality of other numerical evaluations of the plurality of other graphical representations of the other documents; and

generating an identification of a top number of documents based on a corresponding number of top ranked graphical representations.

2 . The method of claim 1 , wherein the document is generated by a multimodal model.

3 . The method of claim 1 , wherein:

a thickness of the line joining the two nodes indicates the level of association; and

the level of association is determined based on at least one of a syntactic connection or a semantic connection.

4 . The method of claim 1 , further comprising:

generating a simplified graphical representation of the document based on the graphical representation;

comparing the simplified graphical representation against a plurality of other simplified graphical representations of the other documents that are generated based on the plurality of other graphical representations; and

determining a writing style for the document based on the plurality of other simplified graphical representations of the other documents.

5 . The method of claim 1 , wherein:

the plurality of words of the document is analyzed by a first multimodal model;

the graphical representation is generated by the first multimodal model; and

the model that calculates the numerical evaluation is a second multimodal model.

6 . The method of claim 1 , wherein the characteristic of the document is a complexity of the document.

7 . The method of claim 1 , wherein the characteristic of the document is a simplicity of the document.

8 . A system comprising:

one or more processing units;

computer-readable storage media storing instructions that, when executed by the one or more processing units, cause the system to perform operations comprising:

analyzing a plurality of words included in a document;

generating a graphical representation of the document based on the plurality of words, wherein:

the graphical representation of the document includes a plurality of nodes and a plurality of lines joining the plurality of nodes;

each node of the plurality of nodes represents a word of the plurality of words;

a line joining two nodes represents a level of association between two words represented by the two nodes; and

the plurality of words is omitted from the graphical representation of the document;

determining a number of the plurality of nodes and a number of the plurality of lines included in the graphical representation of the document;

calculating, by utilizing a model without access to the plurality of words included in the document and based on the number of the plurality of nodes and the number of the plurality of lines included in the graphical representation of the document, a numerical evaluation of the graphical representation of the document that quantifies a characteristic of the document;

ranking the graphical representation of the document against a plurality of other graphical representations of other documents based on the numerical evaluation of the graphical representation of the document and a plurality of other numerical evaluations of the plurality of other graphical representations of the other documents; and

generating an identification of a top number of documents based on a corresponding number of top ranked graphical representations.

9 . The system of claim 8 , wherein:

a thickness of the line joining the two nodes indicates the level of association; and

the level of association is determined based on at least one of a syntactic connection or a semantic connection.

10 . The system of claim 8 , wherein the operations further comprise:

generating a simplified graphical representation of the document based on the graphical representation;

comparing the simplified graphical representation against a plurality of other simplified graphical representations of the other documents that are generated based on the plurality of other graphical representations; and

determining a writing style for the document based on the plurality of other simplified graphical representations of the other documents.

11 . The system of claim 8 , wherein:

the plurality of words of the document is analyzed by a first multimodal model;

the graphical representation is generated by the first multimodal model; and

the model that calculates the numerical evaluation is a second multimodal model.

12 . The system of claim 8 , wherein the characteristic of the document is a complexity of the document.

13 . The system of claim 8 , wherein the characteristic of the document is a simplicity of the document.

14 . A computer-readable storage medium having encoded thereon instructions that, when executed by one or more processing units, cause a system to perform operations comprising:

analyzing a plurality of words included in a document;

generating a graphical representation of the document based on the plurality of words, wherein:

the graphical representation the document includes a plurality of nodes and a plurality of lines joining the plurality of nodes;

each node of the plurality of nodes represents a word of the plurality of words;

a line joining two nodes represents a level of association between two words represented by the two nodes; and

the plurality of words is omitted from the graphical representation of the document;

determining a number of the plurality of nodes and a number of the plurality of lines included in the graphical representation of the document;

calculating, by utilizing a model without access to the plurality of words included in the document and based on the number of the plurality of nodes and the number of the plurality of lines included in the graphical representation of the document, a numerical evaluation of the graphical representation of the document that quantifies a characteristic of the document;

ranking the graphical representation of the document against a plurality of other graphical representations of other documents based on the numerical evaluation of the graphical representation of the document and a plurality of other numerical evaluations of the plurality of other graphical representations of the other documents; and

generating an identification of a top number of documents based on a corresponding number of top ranked graphical representations.

15 . The computer-readable storage medium of claim 14 , wherein:

a thickness of the line joining the two nodes indicates the level of association; and

the level of association is determined based on at least one of a syntactic connection or a semantic connection.

16 . The computer-readable storage medium of claim 14 , wherein the operations further comprise:

generating a simplified graphical representation of the document based on the graphical representation;

comparing the simplified graphical representation against a plurality of other simplified graphical representations of the other documents that are generated based on the plurality of other graphical representations; and

determining a writing style for the document based on the plurality of other simplified graphical representations of the other documents.

17 . The computer-readable storage medium of claim 14 , wherein:

the plurality of words of the document is analyzed by a first multimodal model;

the graphical representation is generated by the first multimodal model; and

the model that calculates the numerical evaluation is a second multimodal model.

18 . The computer-readable storage medium of claim 14 , wherein the characteristic of the document is a complexity of the document.

19 . The computer-readable storage medium of claim 14 , wherein the characteristic of the document is a simplicity of the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2023
From: KHOLODKOV, DMITRY VALENTINOVICH
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064296/0473 →
Continuity (1)
Related Publication 20250028902A1 · Jan 23, 2025
References Cited (17)
US 8296168B2 · Subrahmanian et al. · 2012 [cited by applicant]
US 8467716B2 · Burstein et al. · 2013 [cited by applicant]
US 10324969B2 · Maitra et al. · 2019 [cited by applicant]
US 11334716B2 · Beller · 2022 [cited by examiner]
US 20090254543A1 · Ber · 2009 [cited by examiner]
US 20110267350A1 · Curbera · 2011 [cited by examiner]
US 20190018843A1 · Och et al. · 2019 [cited by applicant]
US 20210004432A1 · Li et al. · 2021 [cited by applicant]
US 20220103872A1 · Liu · 2022 [cited by examiner]
CN 116127046A · 2023 [cited by applicant]
KR 20230047849A · 2023 [cited by applicant]
Y. Hou, W. Zhao, Y. Li and J. Wen, “Privacy-Preserved Neural Graph Similarity Learning”, Nov. 28, 2022, IEEE Xplore, 2022 IEEE International Conference on Data Mining (ICDM), 191-200 (Year: 2022). [cited by examiner]
Kumar, Ajitesh, “Large language models: Concepts & Examples”, Retrieved from: https://vitalflux.com/large-language-models-concepts-examples/, May 1, 2023, 14 Pages. [cited by applicant]
Schulman, et al., “Introducing ChatGPT”, Retrieved from: https://openai.com/blog/chatgpt, Nov. 30, 2022, 11 Pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/036474, mailed on Nov. 29, 2024, 14 pages. [cited by applicant]
Zhang, et al., “Evolution Analysis of Information Retrieval based on co-word network”, 3rd International Conference On Electronic Information Technology And Computer Engineering (EITCE), Oct. 18, 2019, pp. 1837-1840. [cited by applicant]
International Preliminary Report on Patentability (Chapter I) received for PCT Application No. PCT/US2024/036474, mailed on Jan. 29, 2026, 09 pages. [cited by applicant]