IP Library Granted Patent US 12,475,324
Granted Patent B2
US 12,475,324 · App. 17/940,019 · Granted Nov 18, 2025

Artificial intelligence-enabled system and method for authoring a scientific document

Inventor: Ilango Ramanujam (Plainsboro, NJ)
Assignee: ZYLIQ INC.
G06F40/40G06F16/345G06F40/186G06V30/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,475,324
App. No.
17/940,019
Granted
Nov 18, 2025
Kind
B2
Abstract

A system and a method for automatically authoring a scientific document using a machine learning model and natural language processing (NLP) with minimal user intervention are provided. The system configures a scientific document template including multiple sections based on scientific document requirements. The system maps the sections in the scientific document template with content from the source documents by executing a section mapping algorithm and automatically generates the scientific document. The mapping includes matching the sections of the scientific document template with sections extracted from the source documents, and predicting appropriate sections in the scientific document template for rendering the content from the source documents based on the matching using the machine learning model and historical scientific document information. The system executes one or more content editing functions, for example, tense conversion, additional information fetch and display, post-text to in-text conversion, etc., on the scientific document using NLP.

Claims (64)

1 . A system for automatically authoring a scientific document using a machine learning model and natural language processing, the system comprising:

at least one processor;

a non-transitory, computer-readable storage medium operably and communicatively coupled to the at least one processor and configured to store computer program instructions executable by the at least one processor; and

an automated authoring engine defining the computer program instructions, which when executed by the at least one processor, cause the at least one processor to:

configure a scientific document template comprising a plurality of sections based on scientific document requirements, wherein one or more of the plurality of sections are configured as feedback to retrain the machine learning model;

receive a plurality of source documents from a user and store the received source documents in a source database;

automatically extract and pre-process content from the plurality of source documents using natural language processing;

map section names in the scientific document template with table of contents in the plurality of source documents by executing a section mapping algorithm, wherein the mapping comprises:

matching the section names in the scientific document template with the table of contents extracted from the plurality of source documents; and

predicting appropriate sections from among the plurality of sections in the scientific document template for rendering the content from the plurality of source documents into the scientific document template based on the matching of the section names to the table of contents, using the machine learning model and historical scientific document information acquired from the user;

automatically generate the scientific document by rendering the content from the plurality of source documents into the predicted sections of the scientific document template; and

execute one or more of a plurality of content editing functions on the automatically generated scientific document using the natural language processing.

2 . The system of claim 1 , wherein the plurality of sections of the scientific document template comprises fixed sections and user-configurable sub-sections.

3 . The system of claim 1 , wherein the plurality of content editing functions comprises:

automatically converting tenses of the content in the automatically generated scientific document based on user preferences by executing a natural language generation algorithm;

highlighting data fields in the automatically generated scientific document that require attention and editing from the user; and

executing post-text to in-text conversion.

4 . The system of claim 1 , wherein one or more of the computer program instructions defined by the automated authoring engine, when executed by the at least one processor, cause the at least one processor to interpret in-text tables from the plurality of source documents and generate an in-text table summary by executing a natural language understanding algorithm.

5 . The system of claim 1 , wherein one or more of the computer program instructions defined by the automated authoring engine, when executed by the at least one processor, cause the at least one processor to fetch and display, in response to a user input, additional information from the plurality of source documents for selection and rendering into one or more of the plurality of sections in the scientific document template, wherein the user input is configured as additional feedback to retrain the machine learning model.

6 . The system of claim 1 , wherein one or more of the computer program instructions defined by the automated authoring engine, when executed by the at least one processor, cause the at least one processor to provide selective access of one of: an entirety of the automatically generated scientific document and one or more sections of the automatically generated scientific document, to one or more co-authors of the automatically generated scientific document for performing one or more actions on the automatically generated scientific document.

7 . The system of claim 1 , wherein one or more of the computer program instructions defined by the automated authoring engine, when executed by the at least one processor, cause the at least one processor to generate and render a preview of the automatically generated scientific document on a preview screen of a user interface for subsequent editing and automatic regeneration of the scientific document.

8 . The system of claim 1 , wherein one or more of the computer program instructions defined by the automated authoring engine, when executed by the at least one processor, cause the at least one processor to generate and render one or more of a plurality of reports comprising:

a traceability report configured to display the mapping of the section names with the table of contents of the source documents containing the rendered content;

an audit report configured to record and display actions performed on the automatically generated scientific document; and

a version history report configured to display versions of the automatically generated scientific document.

9 . The system of claim 1 , wherein the scientific document is a clinical study report, and wherein the scientific document requirements based on which the scientific document template is configured comprise regulatory authority guidelines, and wherein the regulatory authority guidelines comprise the International Council for Harmonisation of Technical Requirements for Pharmaceuticals for Human Use (ICH) E3 guidelines defined by the ICH.

10 . The system of claim 1 , wherein the plurality of source documents comprises a protocol document, a statistical analysis plan document, a case report form, safety narratives, in-text tables, post-text tables, summary reports, and tables, listings, and figures.

11 . A method employing an automated authoring engine defining computer program instructions executable by at least one processor for automatically authoring a scientific document using a machine learning model and natural language processing, the method comprising:

configuring a scientific document template comprising a plurality of sections based on scientific document requirements, wherein one or more of the plurality of sections are configured as feedback to retrain the machine learning model;

receiving a plurality of source documents from a user and storing the received source documents in a source database;

automatically extracting and pre-processing content from the plurality of source documents using natural language processing;

mapping section names in the scientific document template with table of contents in the plurality of source documents by executing a section mapping algorithm, wherein the mapping comprises:

matching the section names in the scientific document template with the table of contents extracted from the plurality of source documents; and

predicting appropriate sections from among the plurality of sections in the scientific document template for rendering the content from the plurality of source documents into the scientific document template based on the matching of the section names to the table of contents, using the machine learning model and historical scientific document information acquired from the user;

automatically generating the scientific document by rendering the content from the plurality of source documents into the predicted sections of the scientific document template; and

executing one or more of a plurality of content editing functions on the automatically generated scientific document using natural language processing.

12 . The method of claim 11 , wherein the plurality of sections of the scientific document template comprises fixed sections and user-configurable sub-sections.

13 . The method of claim 11 , wherein the plurality of content editing functions comprises:

automatically converting tenses of the content in the automatically generated scientific document based on user preferences by executing a natural language generation algorithm;

fetching and displaying, in response to a user input, additional information from the plurality of source documents for selection and rendering into one or more of the plurality of sections in the scientific document template, wherein the user input is configured as additional feedback to retrain the machine learning model;

highlighting data fields in the automatically generated scientific document that require attention and editing from the user; and

executing post-text to in-text conversion.

14 . The method of claim 11 , further comprising interpreting in-text tables from the plurality of source documents and generating an in-text table summary by executing a natural language understanding algorithm.

15 . The method of claim 11 , further comprising providing selective access of one of: an entirety of the automatically generated scientific document and one or more sections of the automatically generated scientific document, to one or more co-authors of the automatically generated scientific document for performing one or more actions on the automatically generated scientific document.

16 . The method of claim 11 , further comprising generating and rendering a preview of the automatically generated scientific document on a preview screen of a user interface for subsequent editing and automatic regeneration of the scientific document.

17 . The method of claim 11 , further comprising generating and rendering one or more of a plurality of reports comprising:

a traceability report configured to display the mapping of the sections with the source documents containing the rendered content;

an audit report configured to record and display actions performed on the automatically generated scientific document; and

a version history report configured to display versions of the automatically generated scientific document.

18 . A non-transitory, computer-readable storage medium having embodied thereon, computer program instructions executable by at least one processor for automatically authoring a scientific document using a machine learning model and natural language processing, the computer program instructions when executed by the at least one processor cause the at least one processor to:

configure a scientific document template comprising a plurality of sections based on scientific document requirements, wherein one or more of the plurality of sections are configured as feedback to retrain the machine learning model;

receive a plurality of source documents from a user and store the received source documents in a source database;

automatically extract and pre-process content from the plurality of source documents using natural language processing;

map section names in the scientific document template with table of contents in the plurality of source documents by executing a section mapping algorithm, wherein the mapping comprises:

matching the section names in the scientific document template with the table of contents extracted from the plurality of source documents; and

predicting appropriate sections from among the plurality of sections in the scientific document template for rendering the content from the plurality of source documents into the scientific document template based on the matching of the section names to the table of contents, using the machine learning model and historical scientific document information acquired from the user;

automatically generate the scientific document by rendering the content from the plurality of source documents into the predicted sections of the scientific document template; and

execute one or more of a plurality of content editing functions on the automatically generated scientific document using natural language processing.

19 . The non-transitory, computer-readable storage medium of claim 18 , wherein the plurality of content editing functions comprises:

automatically converting tenses of the content in the automatically generated scientific document based on user preferences by executing a natural language generation algorithm;

fetching and displaying, in response to a user input, additional information from the plurality of source documents for selection and rendering into one or more of the plurality of sections in the scientific document template, wherein the user input is configured as additional feedback to retrain the machine learning model;

highlighting data fields in the automatically generated scientific document that require attention and editing from the user; and

executing post-text to in-text conversion.

20 . The non-transitory, computer-readable storage medium of claim 18 , wherein one or more of the computer program instructions when executed by the at least one processor further cause the at least one processor to interpret in-text tables from the plurality of source documents and generate an in-text table summary by executing a natural language understanding algorithm.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2025
From: SYMBIANCE INC
To: ZYLIQ INC.,
Reel/Frame 071572/0070 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2023
From: RAMANUJAM, ILANGO
To: SYMBIANCE INC.
Reel/Frame 063012/0741 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2023
From: RAMANUJAM, ILANGO
To: SYMBIANCE INC.
Reel/Frame 062969/0296 →
Continuity (1)
Related Publication 20240086647A1 · Mar 14, 2024
References Cited (18)
US 7386442B2 · Dehlinger · 2008 [cited by examiner]
US 10664558B2 · Mahamood · 2020 [cited by examiner]
US 20070203869A1 · Ramsey et al. · 2007 [cited by applicant]
US 20110060761A1 · Fouts · 2011 [cited by applicant]
US 20120016678A1 · Gruber et al. · 2012 [cited by applicant]
US 20180267950A1 · de Mello Brandao · 2018 [cited by examiner]
US 20190260764A1 · Humphrey · 2019 [cited by examiner]
US 20190311003A1 · Tonkin · 2019 [cited by examiner]
US 20210182486A1 · Yoo · 2021 [cited by examiner]
US 20220050960A1 · Sanyasi · 2022 [cited by examiner]
US 20220366344A1 · Shi · 2022 [cited by examiner]
US 20230005573A1 · Ramanujam · 2023 [cited by examiner]
US 20230325590A1 · Shevchenko · 2023 [cited by examiner]
US 20240086647A1 · Ramanujam · 2024 [cited by examiner]
US 20240394422A1 · Kumar · 2024 [cited by examiner]
Buchkremer et al. “The application of artificial intelligence technologies as a substitute for reading and to support and enhance the authoring of scientific review articles”. IEEE Access 7 (2019): 65263-65276. May 20, … [cited by applicant]
Getahun. “After an AI bot wrote a scientific paper on itself, the researcher behind the experiment says she hopes she didn't open a ‘Pandora's box’.” Insider. Jul. 9, 2022 (Jul. 9, 2022) Retrieved on Jan. 7, 2023 (Jan. … [cited by applicant]
Marr. “Artificial Intelligence Can Now Write Amazing Content—What Does That Mean For Humans?” Forbes. Mar. 29, 2019 (Mar. 29, 2019) Retrived on Jan. 7, 2023 (Jan. 7, 2023) from <https://www.forbes.com/sites/bernardmarr/… [cited by applicant]