System and method for performing optical character recognition
Techniques including a system and method for optical character recognition. The techniques may involve the use of a system. The system may include a plurality of optical character recognition engines configured to process, in parallel, at least one document or portion thereof, and produce output results for each of the optical character recognition engines. The system may include a component adapted to combine the output results of each of the optical character recognition engines and produce a single unified view of the at least one document or portion thereof.
1 . A system comprising:
at least one computer hardware processor; and
at least one non-transitory computer readable storage medium, storing processor-executable instructions, that, when executed by the at least one computer hardware processor, cause the at least one computer hardware processor to perform a method comprising:
processing, using a plurality of optical character recognition engines in parallel, at least one document or portion thereof;
combining output results of each of the optical character recognition engines to produce a single unified view of the at least one document or portion thereof; and
producing, from the output results of each of the optical character recognition engines, a respective interval tree for each of the respective output results of each of the optical character recognition engines.
2 . The system according to claim 1 , wherein each respective interval tree comprises a cluster of word entities identified within the at least one document or portion thereof, and wherein each respective interval tree is arranged based on positional characteristics of the word entities identified within the at least one document or portion thereof.
3 . The system according to claim 2 , wherein the at least one non-transitory computer readable storage medium stores further instructions that cause the at least one computer hardware processor to perform:
joining the respective interval trees into a graph of entities using a connected component analysis.
4 . The system according to claim 3 , wherein the plurality of optical character recognition engines includes at least three optical character recognition engines, and wherein combining the output results of each of the optical character recognition engines comprises resolving consensus between the output results of each of the optical character recognition engines.
5 . The system according to claim 4 , wherein resolving consensus between the output results of each of the optical character recognition engines comprises determining consensus at least in part based on a distance between words within cluster group.
6 . The system according to claim 5 , wherein the distance between words within a cluster group is determined based on a determination of a Levenshtein distance.
7 . The system according to claim 1 , wherein the at least one non-transitory computer readable storage medium stores further instructions that cause the at least one computer hardware processor to perform:
outputting the single unified view of the at least one document or portion thereof.
8 . A method comprising:
using at least one computer hardware processor to perform:
processing, using a plurality of optical character recognition engines in parallel, at least one document or portion thereof;
combining output results of each of the optical character recognition engines to produce a single unified view of the at least one document or portion thereof; and
producing, from the output results of each of the optical character recognition engines, a respective interval tree for each of the respective output results of each of the optical character recognition engines.
9 . The method according to claim 8 , wherein producing a respective interval tree for each of the respective output results of each of the optical character recognition engines comprises:
identifying a cluster of word entities within the at least one document or portion thereof for the respective interval tree; and
arranging the respective interval tree based on positional characteristics of the word entities identified within the at least one document or portion thereof.
10 . The method according to claim 9 , further comprising joining the respective interval trees into a graph of entities using a connected component analysis.
11 . The method according to claim 10 , wherein the plurality of optical character recognition engines includes at least three optical character recognition engines, and wherein combining the output results of each of the optical character recognition engines comprises resolving consensus between the output results of each of the optical character recognition engines.
12 . The method according to claim 11 , wherein resolving consensus between the output results of each of the optical character recognition engines comprises determining consensus at least in part based on a distance between words within cluster group.
13 . The method according to claim 12 , wherein determining consensus at least in part based on a distance between words within cluster group comprises:
determining a Levenshtein distance between words within a cluster group; and
determining the distance between words within the cluster group based on the Levenshtein distance.
14 . The method according to claim 8 , further comprising outputting the single unified view of the at least one document or portion thereof.
15 . At least one non-transitory computer-readable storage medium storing processor-executable instructions that, when executed by at least one computer hardware processor, cause the at least one computer hardware processor to perform a method comprising:
processing, using a plurality of optical character recognition engines in parallel, at least one document or portion thereof;
combining output results of each of the optical character recognition engines to produce a single unified view of the at least one document or portion thereof; and
producing, from the output results of each of the optical character recognition engines, a respective interval tree for each of the respective output results of each of the optical character recognition engines.
16 . The at least one non-transitory computer-readable storage medium according to claim 15 , wherein producing a respective interval tree for each of the respective output results of each of the optical character recognition engines comprises:
identifying a cluster of word entities within the at least one document or portion thereof for the respective interval tree; and
arranging the respective interval tree based on positional characteristics of the word entities identified within the at least one document or portion thereof.
17 . The at least one non-transitory computer-readable storage medium according to claim 15 , wherein the plurality of optical character recognition engines includes at least three optical character recognition engines, and wherein combining the output results of each of the optical character recognition engines comprises resolving consensus between the output results of each of the optical character recognition engines.