IP Library Granted Patent US 11,574,491
Granted Patent B2
US 11,574,491 · App. 17/112,322 · Granted Feb 7, 2023

Automated classification and interpretation of life science documents

Inventors: Gary Douglas Shorter (Cary, NC); Barry Matthew Ahrens (Fayetteville, NC); Cara Elizabeth Willoughby (Chapel Hill, NC); Yatesh Dass Midha (Raleigh, NC)
Assignee: IQVIA Inc.
G06V30/413G06F40/279G06F40/30G06V30/414G06V2201/09G06V2201/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,574,491
App. No.
17/112,322
Granted
Feb 7, 2023
Kind
B2
Abstract

A computer-implemented tool for automated classification and interpretation of documents, such as life science documents supporting clinical trials, is configured to perform a combination of raw text, document construct, and image analyses to enhance classification accuracy by enabling a more comprehensive machine-based understanding of document content. The combination of analyses provides context for classification by leveraging relative spatial relationships among text and image elements, identifying characteristics and formatting of elements, and extracting additional metadata from the documents as compared to conventional automated classification tools, wherein natural language processing (NLP) is applied to associate text with tokens, and relevant differences and similarities between protocols are identified.

Claims (40)

1. A computer-implemented method for classifying and interpreting a plurality of life science documents, the method comprising:

receiving a digitized representation of the plurality of life science documents, wherein the digitized representation includes a plurality of text and a plurality of images;

performing a text analysis of the digitized representation of the plurality of life science documents, wherein the text analysis includes analyzing raw text without respect to any set formatting and using a text sequence to provide additional context on the raw text;

performing a construct analysis of the digitized representation of the plurality of life science documents, wherein the construct analysis includes analyzing a relative spatial position of the plurality of text and the plurality of images, analyzing element characteristics, and analyzing context connections;

performing an image analysis of the digitized representation of the plurality of life science documents, wherein the image analysis includes, identifying locations of images in the life science documents and applying image to text conversions to create digitization of text elements;

collecting key metadata from the image data, wherein natural language processing (NLP) to associate relevant text with tokens is utilized, and differences and similarities between protocols and risks involved with one or more amendments are identified;

classifying the plurality of life science documents based on the text analysis, construct analysis, and image analysis; and

tagging the plurality of life science documents with one or more class tags based on the classifying of the plurality of life science documents, wherein the one or more class tags represents a class and/or a subclass for the plurality of life science documents.

2. The computer-implemented method of claim 1 , wherein the context connections include keeping a connected location of an element in relation to the text information that occurs immediately before and immediately after the text information.

3. The computer-implemented method of claim 1 , wherein the text sequence is used to identify specific text to provide the additional context.

4. The computer-implemented method of claim 1 , wherein the element characteristics are analyzed using a text font, size and format.

5. The computer-implemented method of claim 1 , wherein the context connections is analyzed using a nearest neighbor.

6. The computer-implemented method of claim 1 , wherein the construct analysis provides a content connection to determine a relevance of a document element with regard to its position in a document.

7. The computer-implemented method of claim 1 , wherein the construct analysis includes obtaining metadata to provide additional context with regard to a document.

8. The computer-implemented method of claim 1 , wherein a protocol synopsis document is distinguished from an informed assent document.

9. The computer-implemented method of claim 1 , wherein the image analysis is further configured generate metadata to enable document classification and document interpretation.

10. The computer-implemented method of claim 1 , further comprising a step of performing latent semantic analysis to weight one or more document characteristics for relevance.

11. The computer-implemented method of claim 1 , wherein the construct analysis includes distinguishing between different text fonts.

12. The computer-implemented method of claim 1 , wherein the construct analysis is configured to identify how the plurality of life science documents are constructed from constituent elements and relationships.

13. The computer-implemented method of claim 1 , wherein the image analysis includes identifying the one or more images that provide additional metadata.

14. The computer-implemented method of claim 1 , wherein the text analysis includes collecting a plurality of words from each of the life science documents.

15. The computer-implemented method of claim 1 , wherein the spatial information includes one of, or a combination of, a location for each image on each page and a location of text in headers and footers on each page of the life science documents.

16. The computer-implemented method of claim 1 , wherein the construct analysis further includes obtaining one or more connections among document elements.

17. A computing device configured to operate as a computer-implemented automated classification and interpretation tool, comprising:

one or more processors; and

one or more non-transitory computer-readable storage media storing instructions which, when executed by the one or more processors, cause the computing device to:

deconstruct a plurality of life science documents into a standardized data structure to generate a plurality of document elements comprising images and digitized text as an input to the computer-implemented automated classification and interpretation tool;

perform a text analysis, a construct analysis, and an image analysis on the plurality of document elements, wherein the text analysis is performed to identify a bag of words in each of the life science documents, wherein the construct analysis is performed to determine how each of the life science documents are constructed, and wherein the image analysis is performed to acquire metadata to enable for machine-based understanding of each of the life science documents, wherein one or more new protocols are identified, sections of text are identified, and modular similarities and differences are identified;

extract the metadata obtained from a combination of the text analysis, construct analysis, and image analysis in relation to the plurality of life science documents such that the metadata describes at least one of the plurality of document elements;

utilize context-based representations and the extracted metadata to allow for a classification of the plurality of life science documents into one or more predefined classes; and

identify one or more class tags and/or one or more event tags for the plurality of life science documents, wherein the one or more class tags represents a class and/or a subclass for the plurality of life science documents, and wherein the one or more event tags represents one of, an event, an action, a trigger, or a combination thereof.

18. The computing device of claim 17 , wherein the text analysis, construct analysis and image analysis includes at identifying raw text without formatting, a spatial position of text and images, and analyzing the images to generate the metadata.

19. One or more non-transitory computer-readable storage media storing executable instructions which, when executed by one or more processors in a computing device, implement a computer-implemented automated classification tool configured to perform a method including the steps of:

identifying a bag of words using text analysis in a plurality of digitized life science documents;

analyzing a spatial position of document elements and element characteristics using construct analysis in each of the plurality of digitized life science documents;

analyzing on or more images in the plurality of digitized life science documents to generate metadata to aid with document classification;

utilizing results from the text analysis, construct analysis, and image analysis to generate additional metadata, wherein the key metadata from the image data is identified, natural language processing (NLP) to associate relevant text with tokens is utilized, and differences and similarities between protocols and risks involved with one or more amendments are identified;

classifying each of the plurality of life science documents utilizing the generated metadata; and

tagging each of the plurality of life science documents with one or more class tags and one or more event tags, wherein the one or more class tags represents a class and/or a subclass for the each of the plurality of life science documents, wherein the one or more event tags represents one of, an event, an action, a trigger, or a combination thereof.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the classification is adjusted based on a review of the class tags and the event tags.

Assignments (7)
SECURITY INTEREST Recorded Mar 12, 2026
From: IMS SOFTWARE SERVICES LTD.; IQVIA INC.; IQVIA RDS INC.; RULES-BASED MEDICINE, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 075047/0061 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTIES INADVERTENTLY NOT INCLUDED IN FILING PREVIOUSLY RECORDED AT REEL: 065709 FRAME: 618. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY AGREEMENT. Recorded Dec 6, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065790/0781 →
SECURITY INTEREST Recorded Nov 29, 2023
From: IQVIA INC.
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065709/0618 →
SECURITY INTEREST Recorded Nov 29, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 065710/0253 →
SECURITY INTEREST Recorded Jul 12, 2023
From: IQVIA INC.; IMS SOFTWARE SERVICES, LTD.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 064258/0577 →
SECURITY INTEREST Recorded May 24, 2023
From: IQVIA INC.; IQVIA RDS INC.; IMS SOFTWARE SERVICES LTD.; Q SQUARED SOLUTIONS HOLDINGS LLC
To: U.S. BANK TRUST COMPANY, NATIONAL ASSOCIATION
Reel/Frame 063745/0279 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: SHORTER, GARY DOUGLAS; AHRENS, BARRY MATTHEW; WILLOUGHBY, CARA ELIZABETH; MIDHA, YATESH DASS
To: IQVIA INC.
Reel/Frame 054550/0372 →
Continuity (3)
Continuation In Part 17070533 · Oct 14, 2020
Continuation 16289729 · Mar 1, 2019
Related Publication 20210089764A1 · Mar 25, 2021