IP Library Granted Patent US 10,621,676
Granted Patent B2
US 10,621,676 · App. 15/013,284 · Granted Apr 14, 2020

System and methods for extracting document images from images featuring multiple documents

Inventors: Isaac Saft (Kfar Neter, IL); Noam Guzman (Ramat Hasharon, IL)
Assignee: VatBox, Ltd.
G06Q40/123G06K9/00442G06K9/325G06K9/3233
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,621,676
App. No.
15/013,284
Granted
Apr 14, 2020
Kind
B2
Abstract

A system and method for extracting document images from images featuring multiple documents are presented. The method includes receiving a multiple-document image including a plurality of document images, wherein each document image is associated with a document; extracting a plurality of visual identifiers from the multiple-document image, wherein each visual identifier is associated with one of the plurality of document images; analyzing the plurality of visual identifiers to identify each document image; determining, based on the analysis, an image area of each document image; extracting each document image based on its image area.

Claims (41)

1. A method for extracting document images from images featuring multiple documents, comprising:

receiving a multiple-document image including a plurality of document images, wherein each document image is associated with a document;

extracting a plurality of visual identifiers from the multiple-document image, wherein each visual identifier is text indicating information related to one of the plurality of document images;

analyzing the plurality of visual identifiers to identify each document image, wherein each document image is identified based on at least one threshold visual identifier requirement representing a portion of the plurality of visual identifiers that need to be included in each of the identified document image;

identifying, for each identified document image that meets the at least one threshold visual identifier requirement, a boundary based on the analysis, the boundary occupying a textless border around the respective identified document image and enclosing all of the plurality of visual identifiers that need to be included within the document image as represented by the at least one threshold visual identifier requirement;

determining, based on the analysis, an image area of each document image, wherein the image area of the document image is defined by the boundary; and

extracting each document image based on its image area, wherein extracting each document image further comprises generating a file including the document image.

2. The method of claim 1 , wherein analyzing the plurality of visual identifiers further comprises:

executing at least one machine imaging process to identify metadata associated with each visual identifier.

3. The method of claim 1 , wherein each boundary is identified based on portions of the multiple-document image in which no text appears.

4. The method of claim 1 , further comprising:

generating a plurality of files, each file including one of the extracted document images.

5. The method of claim 1 , wherein extracting each document image further comprises at least one of: cutting the document image, copying the document image, and cropping the document image.

6. The method of claim 1 , wherein the visual identifier threshold is any of: a number of visual identifiers, a particular visual identifier, and a combination of visual identifiers.

7. The method of claim 6 , further comprising:

determining, for each document image, whether any required visual identifiers have not been extracted; and

upon determining that at least one required visual identifier has not been extracted, retrieving the at least one required visual identifier.

8. The method of claim 7 , further comprising:

determining, for each document image, an eligibility for a potential value-added tax (VAT) refund based on the visual identifiers.

9. A non-transitory computer readable medium having stored thereon instructions for causing one or more processing units to execute the method according to claim 1 .

10. A system for extracting document images from images featuring multiple documents, comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

receive a multiple-document image including a plurality of document images, wherein each document image is associated with a document;

extract a plurality of visual identifiers from the multiple-document image, wherein each visual identifier is text indicating information related to one of the plurality of document images;

analyze the plurality of visual identifiers to identify each document image, wherein each document image is identified based on at least one threshold visual identifier requirement representing a portion of the plurality of visual identifiers that need to be included in each of the identified document image;

identify, for each identified document image that meets the at least one threshold visual identifier requirement, a boundary based on the analysis, the boundary occupying a textless border around the respective identified document image and enclosing all visual identifiers that need to be included within the document image as represented by the at least one threshold visual identifier requirement;

determine, based on the analysis, an image area of each document image, wherein the image area of the document image is defined by the boundary; and

extract each document image based on its image area, wherein extracting each document image further comprises generating a file including the document image.

11. The system of claim 10 , wherein the system is further configured to:

execute at least one machine imaging process to identify metadata associated with each visual identifier.

12. The system of claim 10 , wherein each boundary is identified based on portions of the multiple-document image in which no text appears.

13. The system of claim 10 , wherein the system is further configured to:

generate a plurality of files, each file including one of the extracted document images.

14. The system of claim 10 , wherein the system is further configured to perform at least one of: cut the document image, copy the document image, and crop the document image.

15. The system of claim 10 , wherein the visual identifier threshold is any of: a number of visual identifiers, a particular visual identifier, and a combination of visual identifiers.

16. The system of claim 15 , wherein the system is further configured to:

determine, for each document image, whether any required visual identifiers have not been extracted; and

retrieve the at least one required visual identifier, upon determining that at least one required visual identifier has not been extracted.

17. The system of claim 16 , wherein the system is further configured to:

determine, for each document image, an eligibility for a potential value-added tax (VAT) refund based on the visual identifiers.

Assignments (4)
SECURITY INTEREST Recorded Sep 11, 2023
From: VATBOX LTD
To: BANK HAPOALIM B.M.
Reel/Frame 064863/0721 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Dec 4, 2019
From: VATBOX LTD
To: SILICON VALLEY BANK
Reel/Frame 051187/0764 →
NUNC PRO TUNC ASSIGNMENT Recorded Feb 8, 2017
From: VATBOX, LTD.
To: VATBOX, LTD.
Reel/Frame 041661/0092 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 2, 2016
From: SAFT, ISAAC; GUZMAN, NOAM
To: VATBOX, LTD.
Reel/Frame 037644/0841 →
Continuity (2)
Provisional Application 62111690 · Feb 4, 2015
Related Publication 20160225101A1 · Aug 4, 2016