IP Library › Granted Patent US 12,243,653
Granted Patent B1
US 12,243,653 · App. 18/810,153 · Granted Mar 4, 2025

Generating structured data records using an extraction neural network

Inventors: Zachary Michael Ziegler (Cambridge, MA); Jonas Sebastian Wulff (Glendale, CA); Evan Hernandez (Wimauma, FL); Daniel Joseph Nadler (Nassau, BS)
Assignee: Xyla Inc.
G16H50/70G06F16/30G06F16/35G06F40/10G06F40/205G06F40/279G06F40/284G06F40/30G16H10/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,653
App. No.
18/810,153
Granted
Mar 4, 2025
Kind
B1
Abstract

A computer-implemented method is disclosed for generating structured data representations from unstructured input text sequences through a neural network-based extraction pipeline. The method involves acquiring input text sequences and processing them with an extraction neural network to produce output text sequences. Each output text sequence includes embedded delimiter tokens that annotate specific text segments, mapping them to semantic categories defined within a structured schema. The annotated text segments encapsulate semantically relevant information extracted from the input sequence. The output text sequences are subsequently processed to generate structured data records.

Claims (47)

1. A method performed by one or more computers, the method comprising:

obtaining an input text sequence;

processing the input text sequence using an extraction neural network, in accordance with a set of extraction neural network parameters, to generate a corresponding output text sequence, wherein for each semantic category in a predefined schema of semantic categories:

the output text sequence includes delimiters that designate a respective text string from the output text sequence as being included in the semantic category; and

the text string from the output text sequence that is designated as being included in the semantic category expresses information from the input text sequence that is relevant to the semantic category; and

processing the output text sequence to generate a structured data record that defines a structured representation of information in the input text sequence with reference to the predefined schema of semantic categories.

2. The method of claim 1 , wherein the input text sequence is extracted from a document.

3. The method of claim 2 , wherein the document is a medical paper describing a clinical trial.

4. The method of claim 3 , wherein the predefined schema of semantic categories includes respective semantic categories corresponding to one or more of: a size of a population studied in the clinical trial, an age group of the population studied in the clinical trial, a medical intervention applied to the population in the clinical trial, a variable under study in the population in the clinical trial, or a result of the clinical trial.

5. The method of claim 1 , wherein processing the output text sequence to generate a structured data record comprises, for each semantic category in the schema:

processing the output text sequence to identify the respective text string from the output text sequence that is designated as being included in the semantic category; and

populating the semantic category in the structured data record with the text string from the output text sequence that is designated as being included in the semantic category.

6. The method of claim 5 , wherein for each semantic category in the schema, processing the output text sequence to identify the respective text string from the output text sequence that is designated as being included in the semantic category comprises:

identifying a position of a delimiter in the output text sequence that defines a start or an end of the text string from the output text sequence that is designated as being included in the semantic category.

7. The method of claim 1 , wherein the delimiters included in the output text sequence define a partition of some or all of the output text sequence into respective text strings that are each designated as being included in a respective semantic category.

8. The method of claim 1 , wherein the output text sequence includes delimiters that define a partition of the output text sequence into a plurality of subsequences, wherein each subsequence corresponds to a respective structured data record.

9. The method of claim 8 , wherein processing the output text sequence to generate the structured data record comprises:

processing the output text sequence to generate a respective structure data record corresponding to each of the plurality of subsequences of the output text sequences delineated by the delimiters.

10. The method of claim 1 , wherein the schema includes a plurality of semantic categories.

11. The method of claim 1 , wherein the extraction neural network generates the output text sequence autoregressively.

12. The method of claim 1 , wherein the extraction neural network generates each token in the output text sequence based on: (i) the input text sequence, and (ii) any preceding tokens in the output text sequence.

13. The method of claim 12 , wherein to generate each token in the output text sequence, the extraction neural network performs operations comprising:

processing: (i) the input text sequence, and (ii) any preceding tokens in the output text sequence, to generate a score distribution over a set of tokens; and

selecting a token for position in the output text sequence in accordance with the score distribution over the set of tokens.

14. The method of claim 13 , wherein selecting the token for the position in the output text sequence in accordance with the score distribution over the set of tokens comprises:

sampling from a probability distribution over the set of tokens, wherein the probability distribution over the set of tokens is based on the score distribution over the set of tokens.

15. The method of claim 13 , wherein selecting the token for the position in the output text sequence in accordance with the score distribution over the set of tokens comprises:

selecting the token having a highest score, from among the set of tokens, under the score distribution over the set of tokens.

16. The method of claim 1 , wherein the extraction neural network includes one or more self-attention neural network layers.

17. The method of claim 1 , wherein the extraction neural network has been trained on a set of training examples, wherein each training example comprises: (i) a training text sequence, and (ii) a target text sequence that should be generated by the extraction neural network by processing the training text sequence.

18. The method of claim 17 , wherein for each training example, training the extraction neural network on the training example comprises:

processing an input text sequence based on the training example using the extraction neural network to generate, for each position in the target text sequence, a respective score distribution over a set of tokens;

determining gradients of an objective function that, for each position in the target text sequence, measures an error between: (i) the score distribution over the set of tokens generated by the extraction neural network for the position, and (ii) a token at the position in the target text sequence.

19. A system comprising:

one or more computers; and

one or more storage devices communicatively coupled to the one or more computers, wherein the one or more storage devices store instructions that, when executed by the one or more computers, cause the one or more computers to perform operations comprising:

obtaining an input text sequence;

processing the input text sequence using an extraction neural network, in accordance with a set of extraction neural network parameters, to generate a corresponding output text sequence, wherein for each semantic category in a predefined schema of semantic categories:

the output text sequence includes delimiters that designate a respective text string from the output text sequence as being included in the semantic category; and

the text string from the output text sequence that is designated as being included in the semantic category expresses information from the input text sequence that is relevant to the semantic category; and

processing the output text sequence to generate a structured data record that defines a structured representation of information in the input text sequence with reference to the predefined schema of semantic categories.

20. One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations comprising:

obtaining an input text sequence;

processing the input text sequence using an extraction neural network, in accordance with a set of extraction neural network parameters, to generate a corresponding output text sequence, wherein for each semantic category in a predefined schema of semantic categories:

the output text sequence includes delimiters that designate a respective text string from the output text sequence as being included in the semantic category; and

the text string from the output text sequence that is designated as being included in the semantic category expresses information from the input text sequence that is relevant to the semantic category; and

processing the output text sequence to generate a structured data record that defines a structured representation of information in the input text sequence with reference to the predefined schema of semantic categories.

Assignments (2)
CHANGE OF NAME Recorded Jul 25, 2025
From: XYLA INC.
To: OPENEVIDENCE INC.
Reel/Frame 072243/0065 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2024
From: ZIEGLER, ZACHARY MICHAEL; WULFF, JONAS SEBASTIAN; HERNANDEZ, EVAN; NADLER, DANIEL JOSEPH
To: XYLA INC.
Reel/Frame 068517/0343 →
Continuity (2)
Continuation 18219027 · Jul 6, 2023
Provisional Application 63368434 · Jul 14, 2022
References Cited (25)
US 6446061B1 · Doerre et al. · 2002 [cited by applicant]
US 6629097B1 · Keith · 2003 [cited by applicant]
US 10770180B1 · Kemp · 2020 [cited by examiner]
US 11487942B1 · Senthivel et al. · 2022 [cited by applicant]
US 12094018B1 · O'Malley · 2024 [cited by applicant]
US 20100293451A1 · Carus · 2010 [cited by applicant]
US 20110196704A1 · Mansour · 2011 [cited by examiner]
US 20180082197A1 · Aravamudan et al. · 2018 [cited by applicant]
US 20200126663A1 · Lucas et al. · 2020 [cited by applicant]
US 20200176098A1 · Lucas · 2020 [cited by examiner]
US 20200184278A1 · Zadeh et al. · 2020 [cited by applicant]
US 20210090694A1 · Colley et al. · 2021 [cited by applicant]
US 20220115100A1 · Barve et al. · 2022 [cited by applicant]
US 20220197961A1 · Baek · 2022 [cited by examiner]
US 20220398374A1 · Chowdhury · 2022 [cited by examiner]
US 20230101817A1 · Sinha · 2023 [cited by examiner]
US 20230116115A1 · Shukla · 2023 [cited by examiner]
US 20230117206A1 · Venkateshwaran et al. · 2023 [cited by applicant]
Chen et al., A general approach for improving deep learning-based medical relation extraction using a pre-trained model and fine-tuning, Database, vol. 2019, 2019, baz116, https://doi.org/10.1093/database/baz116 (Year: … [cited by examiner]
Uthealth, CLAMP—Clinical Language Annotation, Modeling, and Processing Toolkit, https://clamp.uth.edu/manual.php, Feb. 1, 2018, pp. 1-69, Center For Computational Biomedicine, School of Biomedical Informatics, The Unive… [cited by examiner]
Vaswani et al., “Attention is all you need,” CoRR, submitted on Dec. 6, 2017, arXiv:1706.03762v5, 15 pages. [cited by applicant]
U.S. Appl. No. 18/219,027, filed Jul. 6, 2023, Ziegler et al. [cited by applicant]
U.S. Appl. No. 18/814,294, filed Aug. 23, 2024, Ziegler et al. [cited by applicant]
Karim et al., “Deep learning-based clustering approaches for bioinformatics,” Briefings in Bioinformatics, Feb. 2, 2020, 22(1):393-415. [cited by applicant]
Li et al., “Neural Natural Language Processing for Unstructured Data in Electronic Health Records: A Review,” CS.CL, Jul. 7, 2021, https://arxiv.org/pdf/2107.02975, 33 pages. [cited by applicant]
Cited By (2)
US 12,412,047 US 12,585,644