IP Library Granted Patent US 11,055,525
Granted Patent B2
US 11,055,525 · App. 16/450,490 · Granted Jul 6, 2021

Determining experiments represented by images in documents

Inventors: Anshuman Sahoo (Toronto, CA); Thomas Kai Him Leung (Toronto, CA); David Qixiang Chen (Toronto, CA); Elvis Mboumien Wianda (Toronto, CA)
Assignee: Scinapsis Analytics Inc.
G06K9/00456G06K9/00463G06K9/6256G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,055,525
App. No.
16/450,490
Granted
Jul 6, 2021
Kind
B2
Abstract

A method may include acquiring one or more image texts from an image of a document, segmenting the image into one or more sub-images using the one or more image texts, determining, by applying a machine learning model, one or more experimental techniques of one or more experiments for the one or more sub-images, and adding, to a knowledge base, one or more mappings of the one or more sub-images to the one or more experiments.

Claims (64)

1. A method, comprising:

acquiring, from a document, (i) one or more image texts of an image and (ii) a plurality of sub-legend texts, wherein the image is a visual representation of one or more experiments;

segmenting the image into one or more sub-images by:

recognizing a plurality of sub-image labels in the one or more image texts, wherein the plurality of sub-legend texts describe one or more experiments corresponding to the one or more sub-images, and

for each sub-legend text of the first plurality of sub-legend texts, attempting a match between the sub-legend text and a sub-image label of the plurality of sub-image labels by:

determining whether the number of sub-legend texts equals the number of sub-image labels, and

determining whether the sub-legend text comprises the sub-image label;

for each sub-image of the one or more sub-images, determining, by applying a machine learning model, that the sub-image is a visual representation of an experimental technique used in the one or more experiments; and

adding, to a knowledge base, one or more mappings of the one or more sub-images to the one or more experiments.

2. The method of claim 1 , further comprising:

determining, using body text of the document, materials for the one or more experiments.

3. The method of claim 1 , further comprising:

classifying, by the machine learning model, a sub-image of the one or more sub-images as a filter class; and

in response to classifying the sub-image as the filter class, determining that the sub-image has no recognized experimental technique.

4. The method of claim 1 , further comprising:

determining one or more bounding boxes within the image for the one or more sub-images.

5. The method of claim 4 , further comprising:

obtaining one or more bounding boxes within the image for the one or more image texts, wherein the first match is attempted using the one or more bounding boxes for the one or more sub-images and the one or more bounding boxes for the one or more sub-image labels.

6. The method of claim 5 , further comprising:

determining macromolecules and experimental contexts for the one or more experiments using the one or more image texts, the one or more bounding boxes for the one or more sub-images, and the one or more bounding boxes for the one or more image texts.

7. A system, comprising:

a memory coupled to a computer processor;

a repository configured to store:

a document comprising (i) an image comprising one or more image texts and (ii) a plurality of sub-legend texts, wherein the image is a visual representation of one or more experiments,

a machine learning model, and

a knowledge base; and

an image analyzer, executing on the computer processor and using the memory, configured to:

acquire, from the document, (i) the one or more image texts of the image and (ii) the plurality of sub-legend texts,

segment the image into one or more sub-images by:

recognizing a plurality of sub-image labels in the one or more image texts, wherein the plurality of sub-legend texts describe one or more experiments corresponding to the one or more sub-images, and

for each sub-legend text of the first plurality of sub-legend texts, attempting a match between the sub-legend text and a sub-image label of the plurality of sub-image labels by:

determining whether the number of sub-legend texts equals the number of sub-image labels, and

determining whether the sub-legend text comprises the sub-image label,

for each sub-image of the one or more sub-images, determine, by applying the machine learning model, that the sub-image is a visual representation of an experimental technique used in the one or more experiments, and

add, to the knowledge base, one or more mappings of the one or more sub-images to the one or more experiments.

8. The system of claim 7 , wherein the image analyzer is further configured to:

determine, using body text of the document, materials for the one or more experiments.

9. The system of claim 7 , wherein the image analyzer is further configured to:

classify, by the machine learning model, a sub-image of the one or more sub-images as a filter class; and

in response to classifying the sub-image as the filter class, determine that the sub-image has no recognized experimental technique.

10. The system of claim 7 , wherein the image analyzer is further configured to:

determine one or more bounding boxes within the image for the one or more sub-images.

11. The system of claim 10 , wherein the image analyzer is further configured to:

obtain one or more bounding boxes within the image for the one or more image texts, wherein the first match is attempted using the one or more bounding boxes for the one or more sub-images and the one or more bounding boxes for the one or more sub-image labels.

12. The system of claim 11 , wherein the image analyzer is further configured to:

determine macromolecules and experimental contexts for the one or more experiments using the one or more image texts, the one or more bounding boxes for the one or more sub-images, and the one or more bounding boxes for the one or more image texts.

13. A non-transitory computer readable medium comprising instructions that, when executed by a computer processor, perform:

acquiring, from a document, (i) one or more image texts of an image and (ii) a plurality of sub-legend texts, wherein the image is a visual representation of one or more experiments;

segmenting the image into one or more sub-images by:

recognizing a plurality of sub-image labels in the one or more image texts, wherein the plurality of sub-legend texts describe one or more experiments corresponding to the one or more sub-images, and

for each sub-legend text of the first plurality of sub-legend texts, attempting a match between the sub-legend text and a sub-image label of the plurality of sub-image labels by:

determining whether the number of sub-legend texts equals the number of sub-image labels, and

determining whether the sub-legend text comprises the sub-image label;

for each sub-image of the one or more sub-images, determining, by applying a machine learning model, that the sub-image is a visual representation of an experimental technique used in the one or more experiments; and

adding, to a knowledge base, one or more mappings of the one or more sub-images to the one or more experiments.

14. The non-transitory computer readable medium of claim 13 , wherein the instructions further perform:

classifying, by the machine learning model, a sub-image of the one or more sub-images as a filter class; and

in response to classifying the sub-image as the filter class, determining that the sub-image has no recognized experimental technique.

15. The non-transitory computer readable medium of claim 13 , wherein the instructions further perform:

determining one or more bounding boxes within the image for the one or more sub-images.

16. The non-transitory computer readable medium of claim 15 , wherein the instructions further perform:

obtaining one or more bounding boxes within the image for the one or more image texts, wherein the first match is attempted using the one or more bounding boxes for the one or more sub-images and the one or more bounding boxes for the one or more sub-image labels.

17. The non-transitory computer readable medium of claim 16 , wherein the instructions further perform:

determining macromolecules and experimental contexts for the one or more experiments using the one or more image texts, the one or more bounding boxes for the one or more sub-images, and the one or more bounding boxes for the one or more image texts.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Mar 12, 2025
From: BANC OF CALIFORNIA F/K/A PACIFIC WESTERN BANK D/B/A PACIFIC WESTERN BUSINESS FINANCE
To: RIVAL DOWNHOLE TOOLS LLC
Reel/Frame 070484/0567 →
SECURITY INTEREST Recorded Oct 28, 2022
From: RIVAL DOWNHOLE TOOLS LLC
To: PACIFIC WESTERN BANK D/B/A PACIFIC WESTERN BUSINESS FINANCE
Reel/Frame 061586/0149 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2019
From: SAHOO, ANSHUMAN; LEUNG, THOMAS KAI HIM; CHEN, DAVID QIXIANG; WIANDA, ELVIS MBOUMIEN
To: SCINAPSIS ANALYTICS INC. DBA BENCHSCI
Reel/Frame 049570/0692 →
Continuity (1)
Related Publication 20200401799A1 · Dec 24, 2020