IP Library Granted Patent US 7,734,636
Granted Patent B2
US 7,734,636 · App. 11/094,415 · Granted Jun 8, 2010

Systems and methods for electronic document genre classification using document grammars

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,734,636
App. No.
11/094,415
Granted
Jun 8, 2010
Kind
B2
Abstract

A system for classifying a genre of an electronic document may include a network processor configured to receive an electronic document and convert the electronic document to rich text format (RTF). The processor may be configured to parse the RTF document into lines of text ordered from top to bottom and left to right and assign tokens to each line of text based on content of the line and to line separators based on space between blocks of lines. The network processor may be configured to sequence the tokens, parse the tokenized document with a number of pre-defined document grammars, determine a probability for each genre corresponding to the electronic document, and classify the electronic document as the genre with the highest probability.

Claims (21)

1. A system for classifying a genre of an electronic document, comprising:

a network processor configured to receive an electronic document; convert the electronic document to rich text format (RTF) using optical character recognition technology; parse the RTF document into lines of text ordered from top to bottom and left to right, and into line separators based on space between blocks of the lines of text; assign tokens to each of the lines of text based on content of the line of text and to each of the line separators; sequence the tokens; parse the tokenized document with a number of pre-defined document grammars; determine a probability for each genre corresponding to the electronic document based on the parsed tokenized document; classify the electronic document as the genre with the highest probability; and route the electronic document to at least one output device based on the genre classification.

2. The system of claim 1 , wherein the at least one output device comprises one of a server, a personal computer, or a personal digital assistant.

3. The system of claim 1 , further comprising a personal computer configured to send electronic documents to the network processor.

4. The system of claim 1 , wherein the network processor is configured to receive the electronic document via email, ftp, or facsimile.

5. The system of claim 1 , wherein the network processor is configured to parse the tokenized document with a first pre-defined document grammar representative of a business card and a second pre-defined document grammar representative of a business letter.

6. A method for classifying a genre of an electronic document, comprising:

receiving an electronic document;

converting the electronic document to rich text format (RTF) using optical character recognition technology;

parsing the RTF document into lines of text ordered from top to bottom and left to right, and into line separators based on space between blocks of the lines of text;

assigning tokens to each of the lines of text based on the content the line of text and to each of the line separators;

sequencing the tokens;

parsing the tokenized document with a number of pre-defined document grammars;

determining a probability for each genre corresponding to the electronic document based on the parsed tokenized document;

classifying the electronic document as the genre with the highest probability; and

routing the electronic document to at least one output device based on the classification.

7. The method of claim 6 , wherein the output device is one of a server, a personal computer, or a personal digital assistant.

8. The method of claim 6 , wherein the electronic document is received via email, ftp, or facsimile.

9. The method of claim 6 , wherein the electronic document is an image file.

10. The method of claim 6 , wherein the electronic document is a text file.

11. The method of claim 6 , wherein said parsing the tokenized document comprises parsing the tokenized document with a first pre-defined document grammar representative of a business card and a second pre-defined document grammar representative of a business letter.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Aug 31, 2022
From: JPMORGAN CHASE BANK, N.A. AS SUCCESSOR-IN-INTEREST ADMINISTRATIVE AGENT AND COLLATERAL AGENT TO BANK ONE, N.A.
To: XEROX CORPORATION
Reel/Frame 061360/0628 →
SECURITY AGREEMENT Recorded Jun 30, 2005
From: XEROX CORPORATION
To: JP MORGAN CHASE BANK
Reel/Frame 016761/0158 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 31, 2005
From: HANDLEY, JOHN C.
To: XEROX CORPORATION
Reel/Frame 016445/0998 →