IP Library Granted Patent US 11,568,664
Granted Patent B2
US 11,568,664 · App. 17/108,985 · Granted Jan 31, 2023

Separating documents based on machine learning models

Inventor: Adithya Kumar (Redmond, WA)
Assignee: SAP SE
G06V30/413G06K9/6217G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,664
App. No.
17/108,985
Granted
Jan 31, 2023
Kind
B2
Abstract

Some embodiments provide a non-transitory machine-readable medium that stores a program executable by a device. The program receives a request to process a file. The file includes a set of images of text. The program further converts the text in each image in the set of images into a set of machine-readable text. The program also uses a machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

Claims (49)

1. A non-transitory machine-readable medium storing a program executable by at least one processing unit of a device, the program comprising sets of instructions for:

receiving a request to process a file, the file comprising a set of images of text;

converting the text in each image in the set of images into a set of machine-readable text; and

using a machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

2. The non-transitory machine-readable medium of claim 1 , wherein the request is further to process a plurality of files that includes the file, wherein each file in the plurality of files comprises a set of images of text, wherein the program further comprises a set of instruction for:

for each file in the plurality of files other than the file, converting the text in each image in the set of images into machine-readable text and using the machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

3. The non-transitory machine-readable medium of claim 2 , wherein the program further comprises set of instructions for:

for each file in the plurality of files, determining a score associated with the file that indicates a confidence of the prediction;

converting the scores associated with the plurality of files to a plurality of single document scores, wherein each single document score indicates a confidence that the set of images of the file are images of pages that belong to a single document; and

determining a lowest single document score in the plurality of single document scores.

4. The non-transitory machine-readable medium of claim 1 , wherein the request is received from a document manager service, wherein the program further comprises sets of instructions for:

determining a score associated with the file that indicates a confidence of the prediction; and

sending the score to the document manager service.

5. The non-transitory machine-readable medium of claim 4 , wherein the machine learning model is further used to determine the score associated with the file.

6. The non-transitory machine-readable medium of claim 1 , wherein the request comprises a unique identifier (ID) for identifying the file, wherein the program further comprises a set of instructions for storing the ID associated with the file, the set of machine-readable text, and the prediction in a storage.

7. The non-transitory machine-readable medium of claim 6 , wherein converting the text in each image in the set of images comprises sending the ID to a queue for processing by a service configured to convert the text in the image into the set of machine-readable text.

8. A method comprising:

receiving a request to process a file, the file comprising a set of images of text;

converting the text in each image in the set of images into a set of machine-readable text; and

using a machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

9. The method of claim 8 , wherein the request is further to process a plurality of files that includes the file, wherein each file in the plurality of files comprises a set of images of text, wherein the method further comprises:

for each file in the plurality of files other than the file, converting the text in each image in the set of images into machine-readable text and using the machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

10. The method of claim 9 further comprising:

for each file in the plurality of files, determining a score associated with the file that indicates a confidence of the prediction;

converting the scores associated with the plurality of files to a plurality of single document scores, wherein each single document score indicates a confidence that the set of images of the file are images of pages that belong to a single document; and

determining a lowest single document score in the plurality of single document scores.

11. The method of claim 8 , wherein the request is received from a document manager service, wherein the method further comprises:

determining a score associated with the file that indicates a confidence of the prediction; and

sending the score to the document manager service.

12. The method of claim 11 , wherein the machine learning model is further used to determine the score associated with the file.

13. The method of claim 8 , wherein the request comprises a unique identifier (ID) for identifying the file, wherein the method further comprises storing the ID associated with the file, the set of machine-readable text, and the prediction in a storage.

14. The method of claim 13 , wherein converting the text in each image in the set of images comprises sending the ID to a queue for processing by a service configured to convert the text in the image into the set of machine-readable text.

15. A system comprising:

a set of processing units; and

a non-transitory machine-readable medium storing instructions that when executed by at least one processing unit in the set of processing units cause the at least one processing unit to:

receive a request to process a file, the file comprising a set of images of text;

convert the text in each image in the set of images into a set of machine-readable text; and

use a machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

16. The system of claim 15 , wherein the request is further to process a plurality of files that includes the file, wherein each file in the plurality of files comprises a set of images of text, wherein the instructions further cause the at least one processing unit to:

for each file in the plurality of files other than the file, convert the text in each image in the set of images into a set of machine-readable text and use the machine learning model to predict, based on the set of machine-readable text, whether the set of images of the file are images of pages that belong to a single document or images of pages that belong to different documents.

17. The system of claim 16 , wherein the instructions further cause the at least one processing unit to:

for each file in the plurality of files, determine a score associated with the file that indicates a confidence of the prediction;

convert the scores associated with the plurality of files to a plurality of single document scores, wherein each single document score indicates a confidence that the set of images of the file are images of pages that belong to a single document; and

determine a lowest single document score in the plurality of single document scores.

18. The system of claim 15 , wherein the request is received from a document manager service, wherein the instructions further cause the at least one processing unit to:

determine a score associated with the file that indicates a confidence of the prediction; and

send the score to the document manager service.

19. The system of claim 18 , wherein the machine learning model is further used to determine the score associated with the file.

20. The system of claim 15 , wherein the request comprises a unique identifier (ID) for identifying the file, wherein the instructions further cause the at least one processing unit to store the ID associated with the file, the set of machine-readable text, and the prediction in a storage.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2020
From: KUMAR, ADITHYA
To: SAP SE
Reel/Frame 054611/0113 →
Continuity (1)
Related Publication 20220171965A1 · Jun 2, 2022