IP Library › Granted Patent US 12,505,688
Granted Patent B2
US 12,505,688 · App. 17/740,683 · Granted Dec 23, 2025

Character-based representation learning for information extraction using artificial intelligence techniques

Inventors: Saurabh Jha (Bangalore, IN); Atul Kumar (Bangalore, IN)
Assignee: Dell Products L.P.
G06V30/19147G06F17/14G06V10/454G06V30/41
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,505,688
App. No.
17/740,683
Granted
Dec 23, 2025
Kind
B2
Abstract

Methods, apparatus, and processor-readable storage media for character-based representation learning for information extraction using artificial intelligence techniques are provided herein. An example computer-implemented method includes identifying, from unstructured documents, words and corresponding document position information using artificial intelligence-based text extraction techniques; generating an intermediate output by implementing at least one character embedding with respect to the unstructured documents using at least one artificial intelligence-based encoder; determining structure-related information for at least a portion of the unstructured documents using one or more artificial intelligence-based graph-related techniques; generating a character-based representation of at least a portion of the unstructured documents using at least one artificial intelligence-based decoder; classifying one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and performing one or more automated actions based on the classifying.

Claims (58)

1 . A computer-implemented method comprising:

identifying, from at least one set of unstructured documents comprising digital documents with varied layouts, one or more words and corresponding document position information by processing at least a portion of the at least one set of unstructured documents using one or more artificial intelligence-based text extraction techniques;

generating an intermediate output by implementing character embeddings with respect to the at least one set of unstructured documents by processing at least a portion of the one or more identified words and corresponding document position information using at least one artificial intelligence-based encoder, wherein implementing the character embeddings comprises:

implementing at least one textual embedding by processing the at least a portion of the one or more identified words and corresponding document position information using at least one integral transform in connection with multiple named-entity recognition prefix-based text labels; and

implementing at least one visual embedding by processing the at least a portion of the one or more identified words and the corresponding document position information using one or more convolution layers with at least one dilation rate;

determining structure-related information for at least a portion of the at least one set of unstructured documents by processing the intermediate output using one or more artificial intelligence-based graph-related techniques;

generating a character-based representation of at least a portion of the at least one set of unstructured documents by processing at least a portion of the intermediate output in connection with the determined structure-related information using at least one artificial intelligence-based decoder;

classifying one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and

performing one or more automated actions based at least in part on the classifying of the one or more portions of the character-based representation;

wherein the method is performed by at least one processing device comprising a processor coupled to a memory.

2 . The computer-implemented method of claim 1 , wherein performing one or more automated actions comprises:

generating one or more inferences based at least in part on the classifying; and

extracting information from at least a portion of the at least one set of unstructured documents based at least in part on the one or more inferences.

3 . The computer-implemented method of claim 1 , wherein processing the intermediate output using one or more artificial intelligence-based graph-related techniques comprises learning, using a graph convolution layer, a two-dimensional layout of at least a portion of the at least one set of unstructured documents, and learning, using the graph convolution layer, information pertaining to how one or more words within at least a portion of the at least one set of unstructured documents relate to one or more other words within at least a portion of the at least one set of unstructured documents.

4 . The computer-implemented method of claim 1 , wherein processing the intermediate output using one or more artificial intelligence-based graph-related techniques comprises processing the intermediate output using one or more two-dimensional connection learning layers.

5 . The computer-implemented method of claim 1 , wherein implementing the at least one textual embedding comprises processing the at least a portion of the one or more identified words and corresponding document position information using at least one Fourier transform.

6 . The computer-implemented method of claim 1 , wherein generating a character-based representation comprises modifying one or more sequences within the intermediate output using at least one union layer of the at least one artificial intelligence-based decoder.

7 . The computer-implemented method of claim 1 , further comprising:

generating at least one hidden state for the one or more portions of the character-based representation by processing at least a portion of the character-based representation using at least one bidirectional long short-term memory model.

8 . The computer-implemented method of claim 7 , further comprising:

processing at least a portion of the generated hidden states using at least one conditional random field layer.

9 . The computer-implemented method of claim 1 , wherein processing at least a portion of the at least one set of unstructured documents using one or more artificial intelligence-based text extraction techniques comprises processing at least a portion of the at least one set of unstructured documents using at least one optical character recognition technique.

10 . The computer-implemented method of claim 1 , wherein identifying one or more words and corresponding document position information comprises identifying one or more words and coordinates of corresponding bounding boxes within the at least one set of unstructured documents.

11 . The computer-implemented method of claim 1 , wherein the at least one set of unstructured documents comprises one or more multi-lingual documents.

12 . The computer-implemented method of claim 1 , further comprising:

labeling at least a portion of the at least one set of unstructured documents;

generating one or more training datasets based at least in part on the at least one labeled portion of the unstructured documents; and

generating one or more validation datasets based at least in part on the at least one labeled portion of the unstructured documents.

13 . A non-transitory processor-readable storage medium having stored therein program code of one or more software programs, wherein the program code when executed by at least one processing device causes the at least one processing device:

to identify, from at least one set of unstructured documents comprising digital documents with varied layouts, one or more words and corresponding document position information by processing at least a portion of the at least one set of unstructured documents using one or more artificial intelligence-based text extraction techniques;

to generate an intermediate output by implementing character embeddings with respect to the at least one set of unstructured documents by processing at least a portion of the one or more identified words and corresponding document position information using at least one artificial intelligence-based encoder, wherein implementing the character embeddings comprises:

implementing at least one textual embedding by processing the at least a portion of the one or more identified words and corresponding document position information using at least one integral transform in connection with multiple named-entity recognition prefix-based text labels; and

implementing at least one visual embedding by processing the at least a portion of the one or more identified words and the corresponding document position information using one or more convolution layers with at least one dilation rate;

to determine structure-related information for at least a portion of the at least one set of unstructured documents by processing the intermediate output using one or more artificial intelligence-based graph-related techniques;

to generate a character-based representation of at least a portion of the at least one set of unstructured documents by processing at least a portion of the intermediate output in connection with the determined structure-related information using at least one artificial intelligence-based decoder;

to classify one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and

to perform one or more automated actions based at least in part on the classifying of the one or more portions of the character-based representation.

14 . The non-transitory processor-readable storage medium of claim 13 , wherein performing one or more automated actions comprises:

generating one or more inferences based at least in part on the classifying; and

extracting information from at least a portion of the at least one set of unstructured documents based at least in part on the one or more inferences.

15 . The non-transitory processor-readable storage medium of claim 13 , wherein processing the intermediate output using one or more artificial intelligence-based graph-related techniques comprises learning, using a graph convolution layer, a two-dimensional layout of at least a portion of the at least one set of unstructured documents, and learning, using the graph convolution layer, information pertaining to how one or more words within at least a portion of the at least one set of unstructured documents relate to one or more other words within at least a portion of the at least one set of unstructured documents.

16 . An apparatus comprising:

at least one processing device comprising a processor coupled to a memory;

the at least one processing device being configured:

to identify, from at least one set of unstructured documents comprising digital documents with varied layouts, one or more words and corresponding document position information by processing at least a portion of the at least one set of unstructured documents using one or more artificial intelligence-based text extraction techniques;

to generate an intermediate output by implementing character embeddings with respect to the at least one set of unstructured documents by processing at least a portion of the one or more identified words and corresponding document position information using at least one artificial intelligence-based encoder, wherein implementing the character embeddings comprises:

implementing at least one textual embedding by processing the at least a portion of the one or more identified words and corresponding document position information using at least one integral transform in connection with multiple named-entity recognition prefix-based text labels; and

implementing at least one visual embedding by processing the at least a portion of the one or more identified words and the corresponding document position information using one or more convolution layers with at least one dilation rate;

to determine structure-related information for at least a portion of the at least one set of unstructured documents by processing the intermediate output using one or more artificial intelligence-based graph-related techniques;

to generate a character-based representation of at least a portion of the at least one set of unstructured documents by processing at least a portion of the intermediate output in connection with the determined structure-related information using at least one artificial intelligence-based decoder;

to classify one or more portions of the character-based representation using one or more artificial intelligence-based statistical modeling techniques; and

to perform one or more automated actions based at least in part on the classifying of the one or more portions of the character-based representation.

17 . The apparatus of claim 16 , wherein performing one or more automated actions comprises:

generating one or more inferences based at least in part on the classifying; and

extracting information from at least a portion of the at least one set of unstructured documents based at least in part on the one or more inferences.

18 . The apparatus of claim 16 , wherein processing the intermediate output using one or more artificial intelligence-based graph-related techniques comprises learning, using a graph convolution layer, a two-dimensional layout of at least a portion of the at least one set of unstructured documents, and learning, using the graph convolution layer, information pertaining to how one or more words within at least a portion of the at least one set of unstructured documents relate to one or more other words within at least a portion of the at least one set of unstructured documents.

19 . The apparatus of claim 16 , wherein processing the intermediate output using one or more artificial intelligence-based graph-related techniques comprises processing the intermediate output using one or more two-dimensional connection learning layers.

20 . The apparatus of claim 16 , wherein implementing the at least one textual embedding comprises processing the at least a portion of the one or more identified words and corresponding document position information using at least one Fourier transform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2022
From: JHA, SAURABH; KUMAR, ATUL
To: DELL PRODUCTS L.P.
Reel/Frame 059881/0125 →
Continuity (1)
Related Publication 20230368553A1 · Nov 16, 2023
References Cited (28)
US 9378200B1 · Cohen et al. · 2016 [cited by applicant]
US 9672279B1 · Cohen et al. · 2017 [cited by applicant]
US 10127304B1 · Cohen et al. · 2018 [cited by applicant]
US 10235452B1 · Savir et al. · 2019 [cited by applicant]
US 10803399B1 · Cohen et al. · 2020 [cited by applicant]
US 11481605B2 · Nguyen · 2022 [cited by examiner]
US 20200334416A1 · Vianu · 2020 [cited by examiner]
US 20210012102A1 · Cristescu · 2021 [cited by examiner]
US 20210248367A1 · Gal · 2021 [cited by examiner]
US 20210374190A1 · Shepherd · 2021 [cited by examiner]
US 20220319217A1 · Paliwal · 2022 [cited by applicant]
US 20230005286A1 · Yebes Torres · 2023 [cited by applicant]
US 20230075290A1 · Wåreus · 2023 [cited by examiner]
US 20230368556A1 · Jha · 2023 [cited by applicant]
CN 113642293A · 2021 [cited by examiner]
W. Xue, Q. Li and Q. Xue, “Text Detection and Recognition for Images of Medical Laboratory Reports With a Deep Learning Approach,” in IEEE Access, vol. 8, pp. 407-416, 2020, doi: 10.1109/ACCESS.2019.2961964 (Year: 2020). [cited by examiner]
R. Gal, S. Ardazi, R. Shilkrot, “Cardinal Graph Convolution Framework for Document Information Extraction,” in ACM Digital Library, 2020, doi: 10.1145/3395027.3419584 (Year: 2020). [cited by examiner]
Jang, B., Kim, M., Harerimana, G., Kang, S., & Kim, J. W. (2020). Bi-LSTM Model to Increase Accuracy in Text Classification: Combining Word2vec CNN and Attention Mechanism. Applied Sciences, 10(17), 5841—. https://doi.o… [cited by examiner]
Wikipedia, Robotic process automation, https://en.wikipedia.org/w/index.php?title=Robotic_process_automation&oldid=1084045031 , Apr. 22, 2022. [cited by applicant]
Wikipedia, ABBYY Fine Reader, https://en.wikipedia.org/w/index.php?title=ABBYY_FineReader&oldid=1082321823, Apr. 12, 2022. [cited by applicant]
Wikipedia, Alteryx, https://en.wikipedia.org/w/index.php?title=Alteryx&oldid=1078943432 , Mar. 24, 2022. [cited by applicant]
Wikipedia, Microsoft Azure, https://en.wikipedia.org/w/index.php?title=Microsoft_Azure&oldid=1085299553 , Apr. 29, 2022. [cited by applicant]
Entrinsik.com, https://entrinsik.com/informer/ , May 5, 2022. [cited by applicant]
Wikipedia, Long short-term memory, https://en.wikipedia.org/w/index.php?title=Long_short-term_memory&oldid=1085879235 , May 2, 2022. [cited by applicant]
Wikipedia, Autoregressive integrated moving average, https://en.wikipedia.org/w/index.php?title=Autoregressive_integrated_moving_average&oldid=1086361303 , May 5, 2022. [cited by applicant]
Nishida, K., Exploratory.io, An Introduction to Time Series Forecasting with Prophet in Exploratory, Apr. 12, 2017. [cited by applicant]
Rigby, J., TowardsDataScience.com, AddressNet: How to build a robust street address parser using a Recurrent Neural Network, Dec. 5, 2018. [cited by applicant]
Github.com, Libpostal, https://github.com/openvenues/libpostal , May 5, 2022. [cited by applicant]
Cited By (1)
US 12,639,966