IP Library Granted Patent US 11,373,423
Granted Patent B2
US 11,373,423 · App. 17/070,533 · Granted Jun 28, 2022

Automated classification and interpretation of life science documents

Inventors: Gary Shorter (Danbury, CT); Barry Ahrens (Danbury, CT)
Assignee: IQVIA Inc.
G06V30/413G06F40/279G06F40/30G06V30/414G06V2201/09G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,373,423
App. No.
17/070,533
Granted
Jun 28, 2022
Kind
B2
Abstract

A computer-implemented tool for automated classification and interpretation of documents, such as life science documents supporting clinical trials, is configured to perform a combination of raw text, document construct, and image analyses to enhance classification accuracy by enabling a more comprehensive machine-based understanding of document content. The combination of analyses provides context for classification by leveraging relative spatial relationships among text and image elements, identifying characteristics and formatting of elements, and extracting additional metadata from the documents as compared to conventional automated classification tools.

Claims (40)

1. A computer-implemented method for classifying and interpreting a plurality of life science documents, the method comprising steps of:

receiving a digitized representation of each of the plurality of life science documents, wherein the digitized representation comprises a plurality of document elements having one or more text and/or one or more images

performing a text analysis of the digitized representation of each of the plurality of life science documents, wherein the text analysis includes identifying raw words from the one or more text of each of the plurality of document elements;

performing a construct analysis of the digitized representation of each of the plurality of life science documents, wherein the construct analysis includes, identifying a document context such that the identified document context describes one or more characteristic of each of the plurality of document elements and a relative spatial position of each of the plurality of document elements on one or more pages of each of the plurality of life science documents;

performing an image analysis of the digitized representation of each of the plurality of life science documents, wherein the image analysis includes, identifying one or more images and processing the identified one or more images to extract one or more additional characteristics of each of the plurality of document elements;

collectively utilizing results of the text analysis, the construct analysis, and the image analysis to classify each of the plurality of life science documents into one or more predefined classes; and

tagging each of the plurality of life science documents with one or more class tags and one or more event tags, wherein the one or more class tags represents a class and/or a subclass for the each of the plurality of life science documents, wherein the one or more event tags represents one of, an event, an action, a trigger, or a combination thereof.

2. The computer-implemented method of claim 1 , wherein the relative spatial position of each of the plurality of document elements comprises at least one of, a header, a footer, a caption, a footnote, a title, or a combination thereof.

3. The computer-implemented method of claim 1 , further comprising a step of identifying the document context by identifying a formatting of each of the plurality of life science documents.

4. The computer-implemented method of claim 1 , wherein the image analysis is performed by identifying one or more logos, one or more graphics, one or more diagrams, one or more diagram text, one or more captions, or a combination thereof.

5. The computer-implemented method of claim 4 , further comprising a step of extracting an additional document context from the identified one or more logos, the one or more graphics, the one or more diagrams, the one or more diagram text, and the one or more captions.

6. The computer-implemented method of claim 1 , further comprising a step of performing the image analysis by converting the one or more images of each of the plurality of documents elements into a text for extracting the text in digitized form from the one or more images.

7. The computer-implemented method of claim 1 , further comprising a step of performing the image analysis by tracking a sequence of the one or more text of each of the plurality of document elements in each of the plurality of the life science documents.

8. The computer-implemented method of claim 1 , further comprising a step of performing the construct analysis by tracking a text neighboring each of the plurality of document elements.

9. The computer-implemented method of claim 1 , further comprising a step of classifying a content in each of the plurality of life science documents into the one or more predefined classes.

10. The computer-implemented method of claim 1 , further comprising a step of generating a metadata associated with each of the plurality of life science documents such that the generated metadata is utilized, at least in part, to perform the classification of each of the plurality of life science documents, wherein the metadata is generated using one of, the text analysis, the construct analysis, the image analysis, or a combination thereof.

11. The computer-implemented method of claim 1 , further comprising a step of generating the digitized representation of each of the plurality of life science documents using an image capture device, wherein the image capture device is selected from one of, a camera, a scanner, or a combination thereof.

12. The computer-implemented method of claim 1 , further comprising a step of providing a real-time classification feedback of each of the plurality of life science documents to a human operator, wherein the real-time classification feedback comprises one of, a suggested classification for each of the plurality of life science documents, an associated metadata, or a combination thereof.

13. The computer-implemented method of claim 12 , further comprising a step of enabling the human operator to review the real-time classification feedback of each of the plurality of life science documents through a User Interface (UI).

14. The computer-implemented method of claim 13 , wherein the reviewed real-time classification feedback of each of the plurality of life science documents is used as an input to a Machine learning process for enhancing an accuracy of classification and interpretation of each of the plurality of life science documents.

15. The computer-implemented method of claim 1 , wherein the characteristic of each of the plurality of document elements include one of, a font of the text, a size of the text, a format of the text, or a combination thereof.

16. The computer-implemented method of claim 1 , wherein the one or more predefined classes includes classes defined by a Drug Information Association (DIA).

17. A computing device configured to operate as a computer-implemented automated classification and interpretation tool, comprising:

one or more processors; and

one or more non-transitory computer-readable storage media storing instructions which, when executed by the one or more processors, cause the computing device to:

deconstruct one or more life science documents into a standardized data structure to generate a plurality of document elements comprising one or more images and one or more digitized text as an input to the computer-implemented automated classification and interpretation tool;

perform a combination of a text analysis, a construct analysis, and an image analysis on each of the plurality of document elements to create a context-based representations of each of the plurality of life science documents such that a spatial relationship between each of the plurality of document elements is identified;

extract a metadata associated with each of the plurality of life science documents such that the metadata describes at least one of the plurality of document elements;

utilize the context-based representations and the extracted metadata to assist the classification of each of the plurality of life science documents into one or more predefined classes; and

interpret each of the plurality of life science documents into one or more class tags and/or one or more event tags wherein the one or more class tags represents a class and/or a subclass for the each of the plurality of life science documents, wherein the one or more event tags represents one of, an event, an action, a trigger, or a combination thereof.

18. The computing device of claim 17 , wherein the executed instructions further cause the computing device to classify each of the plurality of life science documents using a machine learning process, wherein the machine learning process is adjustable according to an input from a human operator.

19. One or more non-transitory computer-readable storage media storing executable instructions which, when executed by one or more processors in a computing device, implement a computer-implemented automated classification tool configured to perform a method including the steps of:

identifying a raw text in a plurality of digitized life science documents;

identifying construction of each of the plurality of digitized life science documents to identify a relative spatial position of one or more text and one or more image elements in each of the plurality of digitized life science documents;

identifying the one or more images to extract text in digitized form;

identifying one or more characteristics of the raw text and the extracted text;

utilizing results of each of the identification steps in combination to generate a metadata;

classifying each of the plurality of life science documents utilizing the generated metadata; and

tagging each of the plurality of life science documents with one or more class tags and one or more event tags, wherein the one or more class tags represents a class and/or a subclass for the each of the plurality of life science documents, wherein the one or more event tags represents one of, an event, an action, a trigger, or a combination thereof.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the classification utilizes one or more of weighting the results, application of latent semantic analysis, or non-parametric Analysis Of Covariance (ANCOVA).

Assignments (7)
SECURITY INTEREST Recorded Mar 12, 2026
From: IMS SOFTWARE SERVICES LTD.; IQVIA INC.; IQVIA RDS INC.; RULES-BASED MEDICINE, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 075047/0061 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTIES INADVERTENTLY NOT INCLUDED IN FILING PREVIOUSLY RECORDED AT REEL: 065709 FRAME: 618. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY AGREEMENT. Recorded Dec 6, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065790/0781 →
SECURITY INTEREST Recorded Nov 29, 2023
From: IQVIA INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065709/0618 →
SECURITY INTEREST Recorded Nov 29, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065710/0253 →
SECURITY INTEREST Recorded Jul 12, 2023
From: IQVIA INC.; IMS SOFTWARE SERVICES, LTD.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 064258/0577 →
SECURITY INTEREST Recorded May 24, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 063745/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 14, 2020
From: SHORTER, GARY; AHRENS, BARRY
To: IQVIA INC.
Reel/Frame 054054/0907 →