IP Library › Granted Patent US 12,293,843
Granted Patent B1
US 12,293,843 · App. 18/815,494 · Granted May 6, 2025

Generating, filtering, and combing structured data records using machine learning

Inventors: Zachary Michael Ziegler (Cambridge, MA); Jonas Sebastian Wulff (Glendale, CA); Evan Hernandez (Wimauma, FL); Daniel Joseph Nadler (Nassau, BS)
Assignee: Xyla Inc.
G16H50/70G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,843
App. No.
18/815,494
Granted
May 6, 2025
Kind
B1
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for processing a large corpus of unstructured data using an extraction neural network to generate a corresponding collection of structured data records. According to one aspect, there is provided a method that includes obtaining a set of input text sequences, and generating a collection of structured data records from the set of input text sequences using an extraction neural network. Each structured data record defines a structured representation of a corresponding input text sequence with reference to a predefined schema of semantic categories. The collection of structured data records can be filtered to identify and remove structured data records that are predicted to be unreliable. The collection of structured data records can then be processed to generate an output that is directed to a selected topic and that aggregates information from across multiple structured data records.

Claims (79)

1. A method performed by one or more computers, the method comprising:

obtaining a set of input text sequences;

generating a collection of structured data records from the set of input text sequences using an extraction neural network, wherein each structured data record defines a structured representation of a corresponding input text sequence with reference to a predefined schema of semantic categories, and wherein generating each structured data record comprises:

processing an input text sequence using the extraction neural network to generate an output text sequence that defines a corresponding structured data record, comprising, for each position in the output text sequence:

processing a sequence of embeddings representing the input text sequence and any part of the output text sequence preceding the position in the output text sequence in accordance with trained values of a set of extraction neural network parameters to generate a score distribution over a set of tokens; and

selecting a token, in accordance with the score distribution over the set of tokens, to occupy the position in the output text sequence;

wherein the extraction neural network has been trained by a machine learning training technique to perform a natural language understanding task;

filtering the collection of structured data records to identify and remove structured data records that are predicted to be unreliable; and

processing the collection of structured data records to generate an article that is directed to a selected topic and that aggregates information from across multiple structured data records.

2. The method of claim 1 , wherein each input text sequence is extracted from a document describing a clinical trial.

3. The method of claim 1 , wherein for each structured data record:

the output text sequence generated by the extraction neural network includes delimiters that, for each semantic category in the schema of semantic categories, designate a respective text string from the output text sequence as being included in the semantic category; and

for each semantic category in the schema of semantic categories, the text string from the output text sequence that is designated as being included in the semantic category expresses information from the input text sequence that is relevant to the semantic category.

4. The method of claim 1 , wherein the predefined schema includes respective semantic categories corresponding to one or more of: a size of a population studied in a clinical trial, an age group of a population studied in a clinical trial, a medical intervention applied to a population in a clinical trial, a variable under study in a population in a clinical trial, or a result of a clinical trial.

5. The method of claim 1 , wherein filtering the collection of structured data records to identify and remove structured data records that are predicted to be unreliable comprises, for each structured data record:

processing the structured data record to evaluate whether the structured data record satisfies each of one or more reliability criteria; and

generating a reliability prediction characterizing a predicted reliability of information included in the structured data record based on a result of evaluating whether the structured data record satisfies the reliability criteria.

6. The method of claim 5 , wherein processing the structured data record to evaluate whether the structured data record satisfies each of one or more reliability criteria comprises:

selecting a semantic category from the schema of semantic categories;

determining whether the text string included in semantic category in the structured data record is found in the corresponding input text sequence;

determining whether a reliability criterion is satisfied based on whether the text string included in the semantic category in the structured data record is found in the corresponding input text sequence.

7. The method of claim 5 , wherein processing the structured data record to evaluate whether the structured data record satisfies each of one or more reliability criteria comprises:

selecting a semantic category from the schema of semantic categories; and

determining a confidence of the extraction neural network in generating the text string included in the semantic category in the structured data record; and

determining whether a reliability criterion is satisfied based on the confidence of the extraction neural network in generating the text string included in the semantic category in the structured data record.

8. The method of claim 5 , wherein processing the structured data record to evaluate whether the structured data record satisfies each of one or more reliability criteria comprises:

generating a measure of semantic consistency between the structured data record and the corresponding input text sequence; and

determining whether a reliability criterion is satisfied based on the measure of semantic consistency between the structured data record and the corresponding input text sequence.

9. The method of claim 8 , wherein generating the measure of semantic consistency between the structured data record and the corresponding input text sequence comprises:

generating a summary text sequence, based on the structured data record, that summarizes at least some of the information included in the structured data record;

generating an augmented text sequence by combining: (i) the input text sequence, and (ii) the summary text sequence based on the structured data record;

generating a likelihood value for the augmented text sequence using a natural language processing neural network; and

determining the measure of semantic consistency between the structured data record and the input text sequence based on the likelihood value for the augmented text sequence.

10. The method of claim 9 , wherein generating the likelihood value for the augmented text sequence using a natural language processing neural network comprises:

processing the augmented text sequence using the natural language processing neural network to generate a respective score distribution over a set of tokens for each position in the augmented text sequence; and

generating the likelihood value for the augmented text sequence based on, for each position in the augmented text sequence, a score for the token at the position in the augmented text sequence under the score distribution generated by the natural language processing neural network for the position.

11. The method of claim 1 , further comprising:

selecting a semantic category from the schema of semantic categories;

clustering text strings included in the semantic category across the collection of structured data records; and

updating the collection of structured data records based on the clustering.

12. The method of claim 11 , wherein clustering the text strings included in the semantic category across the collection of structured data records comprises:

generating a set of text strings that comprises, for each structured data record in the collection of structured data records, the text string included in the semantic category in the structured data record; and

clustering the set of text strings to generate a partition of the set of text strings into a plurality of clusters.

13. The method of claim 1 , wherein processing the collection of structured data records to generate the article directed to the topic comprises:

generating an article that includes natural language text, one or more graphical visualizations, or both.

14. The method of claim 1 , wherein a request to generate the article specifies a style for the article from a set of possible styles, and wherein processing the structured data records to generate the article directed to the topic comprises:

mapping the style specified in the request to a corresponding set of programmatic instructions, expressed in a template programming language, that is associated with the style; and

applying the set of programmatic instructions associated with the style to the structured data records to generate the article directed to the topic.

15. The method of claim 14 , wherein the set of possible styles includes respective styles corresponding to different education levels.

16. The method of claim 1 , wherein generating the article directed to the topic comprises:

generating a textual summary of the topic; and

including the textual summary of the topic in the article directed to the topic.

17. The method of claim 16 , wherein generating the textual summary of the topic comprises:

obtaining a text string defining the topic;

generating one or more search queries that include the text string defining the topic;

extracting a plurality of text sequences from search results obtained by querying a search engine using the search queries;

classifying one or more of the text sequences as being relevant to the topic using a classification neural network; and

generating the textual summary of the topic using the text sequences classified as being relevant to the topic.

18. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining a set of input text sequences;

generating a collection of structured data records from the set of input text sequences using an extraction neural network, wherein each structured data record defines a structured representation of a corresponding input text sequence with reference to a predefined schema of semantic categories, and wherein generating each structured data record comprises:

processing an input text sequence using the extraction neural network to generate an output text sequence that defines a corresponding structured data record, comprising, for each position in the output text sequence:

processing a sequence of embeddings representing the input text sequence and any part of the output text sequence preceding the position in the output text sequence in accordance with trained values of a set of extraction neural network parameters to generate a score distribution over a set of tokens; and

selecting a token, in accordance with the score distribution over the set of tokens, to occupy the position in the output text sequence;

wherein the extraction neural network has been trained by a machine learning training technique to perform a natural language understanding task;

filtering the collection of structured data records to identify and remove structured data records that are predicted to be unreliable; and

processing the collection of structured data records to generate an article that is directed to a selected topic and that aggregates information from across multiple structured data records.

19. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining a set of input text sequences;

generating a collection of structured data records from the set of input text sequences using an extraction neural network, wherein each structured data record defines a structured representation of a corresponding input text sequence with reference to a predefined schema of semantic categories, and wherein generating each structured data record comprises:

processing an input text sequence using the extraction neural network to generate an output text sequence that defines a corresponding structured data record, comprising, for each position in the output text sequence:

processing a sequence of embeddings representing the input text sequence and any part of the output text sequence preceding the position in the output text sequence in accordance with trained values of a set of extraction neural network parameters to generate a score distribution over a set of tokens; and

selecting a token, in accordance with the score distribution over the set of tokens, to occupy the position in the output text sequence;

wherein the extraction neural network has been trained by a machine learning training technique to perform a natural language understanding task;

filtering the collection of structured data records to identify and remove structured data records that are predicted to be unreliable; and

processing the collection of structured data records to generate an article that is directed to a selected topic and that aggregates information from across multiple structured data records.

20. The one or more non-transitory computer storage media of claim 19 , wherein each input text sequence is extracted from a document describing a clinical trial.

Assignments (2)
CHANGE OF NAME Recorded Jul 25, 2025
From: XYLA INC.
To: OPENEVIDENCE INC.
Reel/Frame 072243/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2024
From: ZIEGLER, ZACHARY MICHAEL; WULFF, JONAS SEBASTIAN; HERNANDEZ, EVAN; NADLER, DANIEL JOSEPH
To: XYLA INC.
Reel/Frame 068517/0343 →
Continuity (6)
Continuation 18814294 · Aug 23, 2024
Continuation 18812375 · Aug 22, 2024
Continuation 18810153 · Aug 20, 2024
Continuation 18810328 · Aug 20, 2024
Continuation 18219027 · Jul 6, 2023
Provisional Application 63368434 · Jul 14, 2022
References Cited (28)
US 6446061B1 · Doerre · 2002 [cited by examiner]
US 6629097B1 · Keith · 2003 [cited by examiner]
US 10770180B1 · Kemp · 2020 [cited by examiner]
US 11487942B1 · Senthivel et al. · 2022 [cited by applicant]
US 12094018B1 · O'Malley · 2024 [cited by examiner]
US 20100293451A1 · Carus · 2010 [cited by examiner]
US 20110196704A1 · Mansour · 2011 [cited by examiner]
US 20180082197A1 · Aravamudan et al. · 2018 [cited by applicant]
US 20200126663A1 · Lucas et al. · 2020 [cited by applicant]
US 20200176098A1 · Lucas · 2020 [cited by examiner]
US 20200184278A1 · Zadeh et al. · 2020 [cited by applicant]
US 20210090694A1 · Colley et al. · 2021 [cited by applicant]
US 20220115100A1 · Barve et al. · 2022 [cited by applicant]
US 20220197961A1 · Baek · 2022 [cited by examiner]
US 20220398374A1 · Chowdhury · 2022 [cited by examiner]
US 20230101817A1 · Sinha · 2023 [cited by examiner]
US 20230116115A1 · Shukla · 2023 [cited by examiner]
US 20230117206A1 · Venkateshwaran · 2023 [cited by examiner]
Chen et al., A general approach for improving deep learning-based medical relation extraction using a pre-trained model and fine-tuning, Database, vol. 2019, 2019, baz116, https://doi.org/10.1093/database/baz116 (Year: … [cited by examiner]
Uthealth, CLAMP—Clinical Language Annotation, Modeling, and Processing Toolkit, https://clamp.uth.edu/manual.php, Feb. 1, 2018, pp. 1-69, Center For Computational Biomedicine, School of Biomedical Informatics, The Unive… [cited by examiner]
Li et al., Neural Natural Language Processing for Unstructured Data in Electronic Health Records: a Review, https://arxiv.org/pdf/2107.02975, Yale University, 2021 (Year: 2021). [cited by examiner]
Karim et al., Deep learning-based clustering approaches for bioinformatics, Briefings in Bioinformatics, vol. 22, Issue 1, Jan. 2021, pp. 393-415, https://doi.org/10.1093/bib/bbz170 (Year: 2021). [cited by examiner]
Vaswani et al., “Attention is all you need,” CoRR, submitted on Dec. 6, 2017, arXiv: 1706.03762v5, 15 pages. [cited by applicant]
U.S. Appl. No. 18/219,027, filed Jul. 6, 2023, Ziegler et al. [cited by applicant]
Chen et al., “A General Approach for Improving Deep Learning-based Medical Relation Extraction Using a Pre-trained Model and Fine-tuning,” Database, Mar. 24, 2019, 2019:1-15. [cited by applicant]
Clamp.uth.edu [online], UTHealth, “CLAMP: Clinical Language Annotation, Modeling, and Processing Toolkit,” Feb. 1, 2018, retrieved on Nov. 26, 2024, retrieved from URL <https://clamp.uth.edu/manual.php>, pp. 1-69. [cited by applicant]
U.S. Appl. No. 18/810,153, filed Aug. 20, 2024, Ziegler et al. [cited by applicant]
U.S. Appl. No. 18/814,294, filed Aug. 23, 2024, Ziegler et al. [cited by applicant]