IP Library Granted Patent US 10,685,224
Granted Patent B1
US 10,685,224 · App. 16/131,740 · Granted Jun 16, 2020

Unsupervised removal of text from form images

Inventor: Rohan Kekatpure (Mountain View, CA)
Assignee: INTUIT, INC.
G06K9/00449G06F40/186G06K9/00463G06K9/4652
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,685,224
App. No.
16/131,740
Granted
Jun 16, 2020
Kind
B1
Abstract

The present disclosure relates to language agnostic unsupervised removal of text from form images. According to one embodiment, a method comprises generating a spectral domain representation of an image by applying a transformation, where the image depicts form layout elements and text elements, applying a first filter to the spectral domain representation to remove a portion of the frequency domain corresponding to the text element, and applying a transformation to the filtered spectral domain representation of the image to generate a reconstructed image. The text elements are not depicted in the reconstructed image.

Claims (64)

1. A method for removing text from form images comprising:

generating a frequency domain representation of an image by applying a transformation to the image, wherein the image depicts at least a portion of a form and wherein the form includes layout elements and text elements;

removing a portion of the frequency domain representation that represents the text elements by applying a first filter to the frequency domain representation, producing a first filtered frequency domain representation; and

generating a reconstructed image of at least the portion of the form by applying a transformation to the first filtered frequency domain representation, wherein the reconstructed image does not depict the text elements;

removing a portion of the frequency domain representation that represents the layout elements by applying a layout filter to the frequency domain representation, producing a second filtered frequency domain representation;

generating a reconstructed text image by applying a transformation to the second filtered frequency domain representation, wherein the reconstructed text image does not depict the layout elements;

matching the reconstructed image to a form template;

generating text from the reconstructed text image; and

mapping a portion of the text to a corresponding field of the form template based on a position of the portion of the text in the reconstructed text image.

2. The method of claim 1 , wherein the portion of the text comprises a text element of the text elements.

3. The method of claim 2 , wherein the reconstructed image is matched to the form template using a neural network.

4. The method of claim 2 , wherein the layout filter comprises a two dimensional high-pass filter.

5. The method of claim 1 , wherein the transformation applied to the image is a two dimensional discrete Fourier transform.

6. The method of claim 1 , wherein the first filter comprises a two dimensional low-pass filter.

7. The method of claim 1 , further comprising:

upon determining an amount of layout elements is below a threshold, adjusting the portion of the frequency domain representation removed by the modifying the filter; and

generating a modified reconstructed image by applying a transformation to the frequency domain representation after being filtered with the modified filter.

8. The method of claim 1 , further comprising:

generating a frequency domain representation of a second image by applying a transformation, wherein the second image depicts at least a second portion of the form;

removing a portion of the frequency domain representation that represents the text elements of the second image by applying a second filter to the frequency domain representation of the second image, producing a filtered frequency domain representation of the second image; and

generating a reconstructed second image by applying a transformation to the filtered frequency domain representation of the second image, wherein the text element is not depicted in the reconstructed second image; and

generating a composite form image by combining the reconstructed image and the reconstructed second image.

9. A system, comprising:

a processor; and

memory storing instructions which, when executed on the processor, operate to remove text from form images, the operation comprising:

generating a frequency domain representation of an image by applying a transformation to the image, wherein the image depicts at least a portion of a form and wherein the form includes layout elements and text elements;

removing a portion of the frequency domain representation that represents the text elements by applying a first filter to the frequency domain representation, producing a first filtered frequency domain representation; and

generating a reconstructed image of at least the portion of the form by applying a transformation to the first filtered frequency domain representation, wherein the reconstructed image does not depict the text elements;

removing a portion of the frequency domain representation that represents the layout elements by applying a layout filter to the frequency domain representation, producing a second filtered frequency domain representation;

generating a reconstructed text image by applying a transformation to the second filtered frequency domain representation, wherein the reconstructed text image does not depict the layout elements;

matching the reconstructed image to a form template;

generating text from the reconstructed text image; and

mapping a portion of the text to a corresponding field of the form template based on a position of the portion of the text in the reconstructed text image.

10. The system of claim 9 , wherein the portion of the text comprises a text element of the text elements.

11. The system of claim 10 , wherein the reconstructed image is matched to the form template using a neural network.

12. The system of claim 10 , wherein the layout filter comprises a two dimensional high-pass filter.

13. The system of claim 9 , wherein the transformation applied to the image is a two dimensional discrete Fourier transform.

14. The system of claim 9 , wherein the first filter comprises a two dimensional low-pass filter.

15. The system of claim 9 , wherein the operation further comprises:

upon determining an amount of layout elements is below a threshold, adjusting the portion of the frequency domain representation removed by the modifying the filter; and

generating a modified reconstructed image by applying a transformation to the frequency domain representation after being filtered with the modified filter.

16. The system of claim 9 , wherein the operation further comprises:

generating a frequency domain representation of a second image by applying a transformation, wherein the second image depicts at least a second portion of the form;

removing a portion of the frequency domain representation that represents the text elements of the second image by applying a second filter to the frequency domain representation of the second image, producing a filtered frequency domain representation of the second image; and

generating a reconstructed second image by applying a transformation to the filtered frequency domain representation of the second image, wherein the text element is not depicted in the reconstructed image; and

generating a composite form image by combining the reconstructed image and the reconstructed second image.

17. A computer-readable medium comprising instructions which, when executed by one or more processors, performs an operation for removing text from form images, the operation comprising:

generating a frequency domain representation of an image by applying a transformation to the image, wherein the image depicts at least a portion of a form and wherein the form includes layout elements and text elements;

removing a portion of the frequency domain representation that represents the text elements by applying a first filter to the frequency domain representation, producing a first filtered frequency domain representation; and

generating a reconstructed image of at least the portion of the form by applying a transformation to the first filtered frequency domain representation, wherein the reconstructed image does not depict the text elements;

removing a portion of the frequency domain representation that represents the layout elements by applying a layout filter to the frequency domain representation, producing a second filtered frequency domain representation;

generating a reconstructed text image by applying a transformation to the second filtered frequency domain representation, wherein the reconstructed text image does not depict the layout elements;

matching the reconstructed image to a form template;

generating text from the reconstructed text image; and

mapping a portion of the text to a corresponding field of the form template based on a position of the portion of the text in the reconstructed text image.

18. The computer-readable medium of claim 17 , wherein the portion of the text comprises a text element of the text elements.

19. The computer-readable medium of claim 17 , wherein the operation further comprises:

upon determining an amount of layout elements is below a threshold, adjusting the portion of the frequency domain representation removed by the modifying the filter; and

generating a modified reconstructed image by a transformation to the frequency domain representation after being filtered with the modified filter.

20. The computer-readable medium of claim 19 , wherein the operation further comprises:

generating a frequency domain representation of a second image by applying a transformation, wherein the second image depicts at least a second portion of the form;

removing a portion of the frequency domain representation that represents the text elements of the second image by applying a second filter to the frequency domain representation of the second image, producing a filtered frequency domain representation of the second image; and

generating a reconstructed second image by applying a transformation to the filtered frequency domain representation of the second image, wherein the text element is not depicted in the reconstructed image; and

generating a composite form image by combining the reconstructed image and the reconstructed second image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2018
From: KEKATPURE, ROHAN
To: INTUIT, INC.
Reel/Frame 046880/0611 →
Continuity (1)
Continuation 15395728 · Dec 30, 2016
Cited By (2)
US 12,198,293 US 12,314,657