IP Library Patent Application 16783906
Patent Application
App. No. 16/783,906

SYSTEM AND METHOD FOR USING ARTIFICIAL INTELLIGENCE TO DEDUCE THE STRUCTURE OF PDF DOCUMENTS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/783,906
Abstract

An architecture for generating accessible documents divides content into blocks and sub-blocks and uses multiple Artificial Intelligence or other machine learning processes to predict the structure type and additional Meta data in PDF documents. The processes base their predictions on user selectable models generated from previously learned well-tagged documents using different classification algorithms and metrics.

Claims (32)

1 . A system for generating accessible documents, the system being of the type that uses a predictive model generated in response to learning from a set of tagged documents, including rendering images within the tagged documents, extracting features from the tagged documents and saving the extracted features along with any provided alternative text in a repository, the system comprising:

a memory configured to store the repository; and

at least one processor configured to perform operations comprising:

dividing a previously tagged, untagged or partially tagged document into blocks;

checking the blocks for well-known patterns to subdivide the blocks into sub-blocks;

using an artificial intelligence tag predictor to predict, based at least in part on the predictive model, a type of tag for the document;

providing meta data for the document to meet accessibility standards;

adjusting heading levels for heading tags to provide compliance with accessibility standards;

using an image artificial intelligence and/or machine learning based algorithm to compare images of said document to other images and if a match is found, assigning meta data to said images; and

using a text tag meta data artificial intelligence and/or machine learning based algorithm to determine other meta data for attachment to tags of said document.

2 . A method for generating an accessible PDF file comprising:

(a) generating a model in response to learning from a set of similar, well tagged PDF files and user input to select metrics to be used for prediction in the model, including rendering images within the well tagged PDF files, extracting features from the well tagged PDF documents and saving the extracted features in a repository along with any provided alternative text;

(b) opening previously tagged, untagged or partially tagged files and dividing the previously tagged, untagged or partially tagged files into blocks;

(c) checking blocks for well-known patterns to subdivide the blocks into sub-blocks;

(d) using an artificial intelligence tag predictor to predict, based at least in part on the generated model, the type of tag including providing additional meta data as required by accessibility standards;

(e) adjusting heading levels for heading tags to ensure compliance with accessibility standards;

(f) using a figure artificial intelligence and/or machine learning based algorithm to compare figures to other figures stored in the database and if a match is found, assigning meta data provided in the database; and

(g) using a text tag meta data artificial intelligence and/or machine learning based algorithm to determine other meta data that may be attached to the tag.

3 . The method of claim 2 wherein the text tag meta data algorithm uses language, alternative text, actual text or expansion text based on the textual content of the tag in addition to the text case such has all caps or mixed case.

4 . The method of claim 2 wherein the model is generated in response to varying the metrics used and the algorithms used for classification.

5 . The method of claim 2 wherein the features of the images are extracted for comparison with new images in documents.

6 . The method of claim 2 wherein the documents are divided into blocks and sub-blocks.

7 . The method of claim 6 wherein the order of the pattern matching is Table, Figure, TOC, List, Index then other tags.

8 . The method of claim 2 wherein the heading level is adjusted to ensure compliance with accessibility standards.

9 . The method of claim 2 wherein the figures are matched to those stored in the repository and the best match (above a certain user-adjustable threshold) is selected and the corresponding Meta data is set for the figure.

10 . The method in claim 2 wherein additional Meta data is assigned to the tag.

11 . The method of claim 2 wherein multiple repositories of learning data of documents derived from the same template or similar templates can be created, stored and then utilized to increase the accuracy of tagging.

12 . A system for generating an accessible file comprising:

(a) generating a model in response to machine learning;

(b) enabling a user to select metrics to be used for prediction in the model;

(c) using an artificial intelligence tag predictor to predict, based at least in part on the generated model, a type of tag; and

(d) assigning metadata to the tag.

Assignments (4)
SECURITY INTEREST Recorded Nov 2, 2021
From: NETCENTRIC TECHNOLOGIES INC.
To: AUDAX PRIVATE DEBT LLC, AS AGENT
Reel/Frame 057998/0685 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: EL-RAYES, FERASS
To: NETCENTRIC TECHNOLOGIES INC. DBA COMMONLOOK®
Reel/Frame 057549/0781 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: NETCENTRIC TECHNOLOGIES INC. DBA COMMONLOOK®
To: NETCENTRIC TECHNOLOGIES INC.
Reel/Frame 057550/0418 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2020
From: EL-RAYES, FERASS
To: NETCENTRIC TECHNOLOGIES INC. DBA COMMONLOOK
Reel/Frame 051873/0565 →