IP Library › Granted Patent US 12,039,798
Granted Patent B2
US 12,039,798 · App. 17/453,070 · Granted Jul 16, 2024

Processing forms using artificial intelligence models

Inventors: Mingfei Gao (Sunnyvale, CA); Ran Xu (Mountain View, CA)
Assignee: Salesforce, Inc.
G06V30/412G06F40/174G06F40/205G06N20/00G06V30/19007
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,039,798
App. No.
17/453,070
Granted
Jul 16, 2024
Kind
B2
Abstract

An application server may receive an input document including a set of input text fields and an input key phrase querying a value for a key-value pair that corresponds to one or more of the set of input text fields. The application server may extract, using an optical character recognition model, a set of character strings and a set of two-dimensional locations of the set of character strings on a layout of the input document. After extraction, the application server may input the extracted set of character strings and the set of two-dimensional locations into a machine learned model that is trained to compute a probability that a character string corresponds to the value for the key-value pair. The application server may then identify the value for the key-value pair corresponding to the input key phrase and may out the identified value.

Claims (62)

1. A method for form processing, comprising:

receiving an input document including a plurality of input text fields;

receiving an input key phrase querying a value for a key-value pair that corresponds to one or more of the plurality of input text fields;

extracting, using an optical character recognition model, a set of character strings and a set of two-dimensional locations of the set of character strings on a layout of the input document;

inputting the extracted set of character strings, the set of two-dimensional locations, and the input key phrase into a machine learned model that is trained to compute a set of probabilities for the set of character strings corresponding to the value for the key-value pair corresponding to the input key phrase;

identifying that a character string of the set of character strings corresponds to the value for the key-value pair corresponding to the input key phrase based at least in part on the inputting and on a respective probability for the character string being the value corresponding to the input key phrase being greater than one or more other respective probabilities of one or more other character strings of the set of character strings; and

transmitting the identified value corresponding to the input key phrase.

2. The method of claim 1 , further comprising:

identifying the set of probabilities for the set of character strings being the value for the key-value pair corresponding to the input key phrase based at least in part on the input key phrase being inputted into the machine learned model.

3. The method of claim 2 , further comprising:

ranking the set of probabilities based at least in part on a value of each probability in the set of probabilities, wherein identifying the value for the key-value pair corresponding to the input key phrase is based at least in part on ranking the set of probabilities.

4. The method of claim 1 , further comprising:

grouping one or more character strings into a value phrase based at least in part on an output of the machine learned model and the set of two-dimensional locations of the set of character strings, wherein identifying the value for the key-value pair is based at least in part on grouping the one or more character strings.

5. The method of claim 1 , further comprising:

determining that the input key phrase does not match a key that corresponds to one or more of the plurality of input text fields; and

identifying a dummy key corresponding to the input key phrase based at least in part on determining that the input key phrase does not match the key, wherein identifying the value for the key-value pair is based at least in part on identifying the dummy key.

6. The method of claim 1 , further comprising:

determining that the input key phrase is associated with an empty value field; and

identifying a dummy value corresponding to the input key phrase based at least in part on determining that the input key phrase is associated with the empty value field, wherein the identified value corresponding to the input key phrase comprises the dummy value.

7. The method of claim 1 , further comprising:

generating a first set of feature representations for a set of keywords included in the plurality of input text fields and a second set of feature representations for the extracted set of character strings; and

generating an unified feature representation for the input key phrase based at least in part on the first set of feature representations and the second set of feature representations, wherein identifying the value for the key-value pair corresponding to the input key phrase is based at least in part on the unified feature representation.

8. The method of claim 7 , further comprising:

applying a dot product between the unified feature representation for the input key phrase and each feature representation of the second set of feature representations, wherein the respective probability that the character string of the set of character strings corresponds to the value for the key-value pair corresponding to the input key phrase is computed based at least in part on applying the dot product.

9. The method of claim 1 , further comprising:

training the machine learned model based at least in part on inputting a plurality of input file formats into the machine learned model.

10. The method of claim 1 , wherein the input document comprises a fixed form, a non-fixed form, or both.

11. The method of claim 1 , wherein the machine learned model comprises a transformer-based machine learned model.

12. An apparatus for form processing, comprising:

one or more processors;

one or more memories coupled with the one or more processors; and

instructions stored in the one or more memories and executable by the one or more processors to cause the apparatus to:

receive an input document including a plurality of input text fields;

receive an input key phrase querying a value for a key-value pair that corresponds to one or more of the plurality of input text fields;

extract, using an optical character recognition model, a set of character strings and a set of two-dimensional locations of the set of character strings on a layout of the input document;

input the extracted set of character strings, the set of two-dimensional locations, and the input key phrase into a machine learned model that is trained to compute a set of probabilities for the set of character strings corresponding to the value for the key-value pair corresponding to the input key phrase;

identify that a character string of the set of character strings corresponds to the value for the key-value pair corresponding to the input key phrase based at least in part on the inputting and on a respective probability for the character string being the value corresponding to the input key phrase being greater than one or more other respective probabilities of one or more other character strings of the set of character strings; and

transmit the identified value corresponding to the input key phrase.

13. The apparatus of claim 12 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

identify the set of probabilities for the set of character strings being the value for the key-value pair corresponding to the input key phrase based at least in part on the input key phrase being inputted into the machine learned model.

14. The apparatus of claim 13 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

rank the set of probabilities based at least in part on a value of each probability in the set of probabilities, wherein identifying the value for the key-value pair corresponding to the input key phrase is based at least in part on ranking the set of probabilities.

15. The apparatus of claim 12 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

group one or more character strings into a value phrase based at least in part on an output of the machine learned model and the set of two-dimensional locations of the set of character strings, wherein identifying the value for the key-value pair is based at least in part on grouping the one or more character strings.

16. The apparatus of claim 12 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

determine that the input key phrase does not match a key that corresponds to one or more of the plurality of input text fields; and

identify a dummy key corresponding to the input key phrase based at least in part on determining that the input key phrase does not match the key, wherein identifying the value for the key-value pair is based at least in part on identifying the dummy key.

17. The apparatus of claim 12 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

determine that the input key phrase is associated with an empty value field; and

identify a dummy value corresponding to the input key phrase based at least in part on determining that the input key phrase is associated with the empty value field, wherein the identified value corresponding to the input key phrase comprises the dummy value.

18. The apparatus of claim 12 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

generate a first set of feature representations for a set of keywords included in the plurality of input text fields and a second set of feature representations for the extracted set of character strings; and

generate an unified feature representation for the input key phrase based at least in part on the first set of feature representations and the second set of feature representations, wherein identifying the value for the key-value pair corresponding to the input key phrase is based at least in part on the unified feature representation.

19. The apparatus of claim 18 , wherein the instructions are further executable by the one or more processors to cause the apparatus to:

apply a dot product between the unified feature representation for the input key phrase and each feature representation of the second set of feature representations, wherein the respective probability that the character string of the set of character strings corresponds to the value for the key-value pair corresponding to the input key phrase is computed based at least in part on applying the dot product.

20. A non-transitory computer-readable medium storing code for form processing, the code comprising instructions executable by one or more processors to:

receive an input document including a plurality of input text fields;

receive an input key phrase querying a value for a key-value pair that corresponds to one or more of the plurality of input text fields;

extract, using an optical character recognition model, a set of character strings and a set of two-dimensional locations of the set of character strings on a layout of the input document;

input the extracted set of character strings, the set of two-dimensional locations, and the input key phrase into a machine learned model that is trained to compute a set of probabilities for the set of character strings corresponding to the value for the key-value pair corresponding to the input key phrase;

identify that a character string of the set of character strings corresponds to the value for the key-value pair corresponding to the input key phrase based at least in part on the inputting and on a respective probability for the character string being the value corresponding to the input key phrase being greater than one or more other respective probabilities of one or more other character strings of the set of character strings; and

transmit the identified value corresponding to the input key phrase.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE RECEIVING PARTY DATA (ASSIGNEE) TO SALESFORCE.COM, INC. PREVIOUSLY RECORDED AT REEL: 57985 FRAME: 732. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 26, 2024
From: GAO, MINGFEI; XU, RAN
To: SALESFORCE.COM, INC
Reel/Frame 066909/0677 →
CHANGE OF NAME Recorded Mar 26, 2024
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 066909/0851 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2021
From: GAO, MINGFEI; XU, RAN
To: QUALCOMM INCORPORATED
Reel/Frame 057985/0732 →
Continuity (1)
Related Publication 20230133690A1 · May 4, 2023