IP Library › Granted Patent US 12,406,135
Granted Patent B2
US 12,406,135 · App. 17/549,270 · Granted Sep 2, 2025

Assisted review of text content using a machine learning model

Inventors: Navita Goyal (College Park, MD); Ani Nenkova Nenkova (Philadelphia, PA); Natwar Modani (Bengaluru, IN); Ayush Maheshwari (Kota, IN); Inderjeet Jayakumar Nair (Indore, IN)
Assignee: ADOBE INC.
G06F40/205G06F40/20G06F40/279G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,406,135
App. No.
17/549,270
Granted
Sep 2, 2025
Kind
B2
Abstract

Techniques described herein are directed to assisting review of documents. In one embodiment, one or more text segments and one or more subjects in a document are identified. A text segment in the document is associated with a corresponding subject identified in the document. The text segment is classified with a content type value corresponding to a relation of the text segment to the corresponding subject. Thereafter, information is provided for the text segment associated with the corresponding subject for display on a user interface. Such information can include a representation of the content type value for the text segment.

Claims (87)

1. A method for assisted review of a document, the method comprising:

identifying two or more similar reference text segments, from a reference corpus of text content, that are similar to a text segment of the document by:

converting the text segment to a dense vector representation of the text segment after replacing numerical text in the text segment with a corresponding token representing the numerical text;

converting reference text segments from the reference corpus to corresponding dense vector representations of the reference text segments after replacing corresponding numerical text in the reference text segments with corresponding tokens representing the corresponding numerical text;

computing corresponding similarity scores between the dense vector representation of the text segment and the corresponding dense vector representations of the reference text segments using a machine learning model trained using the reference corpus to identify similar text segments; and

subsequent to determining the two or more similar reference text segments based on each corresponding similarity score between the dense vector representation to the corresponding dense vector representations above a threshold level of similarity:

accessing the corresponding numerical text from each of the two or more similar reference text segments; and

determining a computed numerical value from the two or more similar reference text segments based on computing at least one of an average value, a median value, minimum value, or a maximum value of the corresponding numerical text from each of the two or more similar reference text segments; and

providing information for the text segment for display on a user interface, the information including the computed numerical value from the two or more similar reference text segments.

2. The method of claim 1 , further comprising:

training the machine learning model by converting the reference text segments to the corresponding dense vector representations after replacing the corresponding numerical text in the reference text segments with the corresponding tokens representing the numerical text.

3. The method of claim 1 , further comprising:

further providing the information for the text segment for display on the user interface, wherein the information comprises a corresponding textual portion of at least one of the two or more similar segments that differ from the text segment to provide suggested language.

4. The method of claim 1 , further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment; and

providing the information for the text segment associated with the subject for display on the user interface.

5. The method of claim 1 , further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

computing the corresponding similarity scores by converting the text segment to the dense vector representation after replacing the alias in the text segment with the subject token; and

providing the information for the text segment associated with the subject for display on the user interface.

6. The method of claim 1 , further comprising:

accessing the document, wherein the document comprises a contract;

mapping an alias to a subject token mapped to a subject after identifying the subject from a preamble of contract;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment; and

providing the information for the text segment associated with the subject for display on the user interface.

7. The method of claim 1 , further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

classifying the text segment with a content type value corresponding to a relation of the text segment to the subject; and

providing the information for the text segment associated with the subject for display on the user interface, wherein the information includes a representation of the content type value for the text segment.

8. The method of claim 1 , further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

classifying the text segment with a content type value corresponding to a relation of the text segment to the subject, the content type value corresponding to the relation of the text segment to the subject comprises one of an obligation type value indicating an obligation of the subject, a prohibition type indicating a prohibition of the subject, and an entitlement type indicating an entitlement of the subject; and

providing the information for the text segment associated with the subject for display on the user interface, wherein the information includes a representation of the content type value for the text segment.

9. One or more non-transitory computer-readable media having computer executable instructions stored thereon which, when executed by one or more processors, cause the one or more processors to execute operations comprising:

identifying two or more similar reference text segments, from a reference corpus of text content, that are similar to a text segment of the document by:

converting the text segment to a dense vector representation of the text segment after replacing numerical text in the text segment with a corresponding token representing the numerical text;

converting reference text segments from the reference corpus to corresponding dense vector representations of the reference text segments after replacing corresponding numerical text in the reference text segments with corresponding tokens representing the corresponding numerical text;

computing corresponding similarity scores between the dense vector representation of the text segment and the corresponding dense vector representations of the reference text segments using a machine learning model trained using the reference corpus to identify similar text segments; and

subsequent to determining the two or more similar reference text segments based on each corresponding similarity score between the dense vector representation to the corresponding dense vector representations above a threshold level of similarity:

accessing the corresponding numerical text from each of the two or more similar reference text segments; and

determining a computed numerical value from the two or more similar reference text segments based on computing at least one of an average value, a median value, minimum value, or a maximum value of the corresponding numerical text of each of the two or more similar reference text segments; and

providing information for the text segment for display on a user interface, the information including the computed numerical value from the two or more similar reference text segments.

10. The method of claim 9 , further comprising:

training the machine learning model by converting the reference text segments to the corresponding dense vector representations after replacing the corresponding numerical text in the reference text segments with the corresponding tokens representing the numerical text.

11. The computer-readable media of claim 9 , the operations further comprising:

further providing the information for the text segment for display on the user interface, wherein the information comprises a corresponding textual portion of at least one of the two or more similar segments that differ from the text segment to provide suggested language.

12. The computer-readable media of claim 9 , the operations further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment; and

providing the information for the text segment associated with the subject for display on the user interface.

13. The computer-readable media of claim 9 , the operations further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

computing the corresponding similarity scores by converting the text segment to the dense vector representation after replacing the alias in the text segment with the subject token; and

providing the information for the text segment associated with the subject for display on the user interface.

14. The computer-readable media of claim 9 , the operations further comprising:

accessing the document, wherein the document comprises a contract;

mapping an alias to a subject token mapped to a subject after identifying the subject from a preamble of contract;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment; and

providing the information for the text segment associated with the subject for display on the user interface.

15. The computer-readable media of claim 9 , the operations further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

classifying the text segment with a content type value corresponding to a relation of the text segment to the subject; and

providing the information for the text segment associated with the subject for display on the user interface, wherein the information includes a representation of the content type value for the text segment.

16. The computer-readable media of claim 9 , the operations further comprising:

mapping an alias to a subject token mapped to a subject from the document;

associating the text segment from the document to the subject from the document based on identifying the alias in the text segment;

classifying the text segment with a content type value corresponding to a relation of the text segment to the subject, the content type value corresponding to the relation of the text segment to the subject comprises one of an obligation type value indicating an obligation of the subject, a prohibition type indicating a prohibition of the subject, and an entitlement type indicating an entitlement of the subject; and

providing the information for the text segment associated with the subject for display on the user interface, wherein the information includes a representation of the content type value for the text segment.

17. A system for identifying atypical text in a natural language text document, the system comprising:

one or more processors; and

one or more memory devices in communication with the one or more processors, the one or more processors to execute operations comprising:

identifying two or more similar reference text segments, from a reference corpus of text content, that are similar to a text segment of the document by:

converting the text segment to a dense vector representation of the text segment after replacing numerical text in the text segment with a corresponding token representing the numerical text;

converting reference text segments from the reference corpus to corresponding dense vector representations of the reference text segments after replacing corresponding numerical text in the reference text segments with corresponding tokens representing the corresponding numerical text;

computing corresponding similarity scores between the dense vector representation of the text segment and the corresponding dense vector representations of the reference text segments using a machine learning model trained using the reference corpus to identify similar text segments; and

subsequent to determining the two or more similar reference text segments based on each corresponding similarity score between the dense vector representation to the corresponding dense vector representations above a threshold level of similarity;

accessing the corresponding numerical text from each of the two or more similar reference text segments; and

determining a computed numerical value from the two or more similar reference text segments based on computing at least one of an average value, a median value, minimum value, or a maximum value of the corresponding numerical text of each of the two or more similar reference text segments; and

providing information for the text segment for display on a user interface, the information including the computed numerical value from the two or more similar reference text segments.

18. The system of claim 17 , the operations further comprising:

training the machine learning model by converting the reference text segments to the corresponding dense vector representations after replacing the corresponding numerical text in the reference text segments with the corresponding tokens representing the numerical text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2021
From: GOYAL, NAVITA; NENKOVA, ANI NENKOVA; MODANI, NATWAR; MAHESHWARI, AYUSH; NAIR, INDERJEET JAYAKUMAR
To: ADOBE INC.
Reel/Frame 058374/0202 →
Continuity (1)
Related Publication 20230186667A1 · Jun 15, 2023
References Cited (20)
US 20090327269A1 · Paparizos · 2009 [cited by examiner]
US 20170366568A1 · Narasimhan · 2017 [cited by examiner]
US 20210383070A1 · Hunter · 2021 [cited by examiner]
WO WO2022213197A1 · 2022 [cited by examiner]
Angelidis, I., et al., “Named Entity Recognition, Linking and Generation for Greek Legislation”, In JURIX, pp. 1-10 (Sep. 2018). [cited by applicant]
Bommarito, M.J. II, et al., “LexNLP: Natural language processing and information extraction for legal and regulatory texts”, Preprint submitted to SSRN—Version 1.01, pp. 1-7 (Jun. 12, 2018). [cited by applicant]
Borchmann, Ł., et al., “Contract discovery: Dataset and a few-shot semantic retrieval challenge with competitive baselines”, arXiv preprint arXiv:1911.03911v2, pp. 15 (Oct. 8, 2020). [cited by applicant]
Chalkidis, I., et al., “Legal-Bert: The muppets straight out of law school”, arXiv preprint arXiv:2010.02559v1, pp. 7 (Oct. 6, 2020). [cited by applicant]
Chalkidis, I., et al., “Extracting contract elements”, ICAIL '17: Proceedings of the 16th edition of the International Conference on Articial Intelligence and Law, pp. 1-10 (Jun. 12-15, 2017). [cited by applicant]
Dragoni, M., et al., “Combining NLP approaches for rule extraction from legal documents”, In 1st Workshop on Mining and REasoning with Legal texts (MIREL 2016), pp. 13 (Dec. 2016). [cited by applicant]
Dua, D., et al., “DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs”, arXiv preprint arXiv:1903.00161v2, pp. 12 (Apr. 16, 2019). [cited by applicant]
Gregory, K., et al., “Siamese neural networks for one-shot image recognition”, In ICML deep learning workshop, pp. 8 (2015). [cited by applicant]
Hendrycks, D., et al., “Cuad: An expert-annotated nlp dataset for legal contract review”, 35th Conference on Neural Information Processing Systems (NeurlPS 2021) Track on Datasets and Benchmarks, arXiv preprint arXiv:21… [cited by applicant]
Jiang, C., et al., “Learning numeral embedding”, arXiv preprint arXiv:2001.00003v1, pp. 16 (Dec. 28, 2019). [cited by applicant]
MacQueen, J., “Some methods for classification and analysis of multivariate observations”, Proceedings of the fifth Berkeley symposium on mathematical statistics and probability, vol. 1, No. 14, pp. 281-297 (1967). [cited by applicant]
Miller, G. A., “WordNet: a lexical database for English”, Communications of the ACM, vol. 38, No. 11, 39-41 (Nov. 1995). [cited by applicant]
Schuster, M., and Paliwal, K.K., “Bidirectional recurrent neural networks”, IEEE transactions on Signal Processing, vol. 45, No. 11, pp. 2673-2681 (Nov. 1997). [cited by applicant]
Tecuci, D.G., et al., “DICR: AI Assisted, Adaptive Platform for Contract Review”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 09, pp. 13638-13639 (2020). [cited by applicant]
Wallace, E., et al., “Do nlp models know numbers? probing numeracy in embeddings”, arXiv preprint arXiv:1909.07940v2, pp. 12 (Sep. 18, 2019). [cited by applicant]
Wolf, T., et al., “Huggingface's transformers: State-of-the-art natural language processing”, arXiv preprint arXiv:1910.03771v3, pp. 11 (Oct. 16, 2019). [cited by applicant]