IP Library Granted Patent US 12,236,355
Granted Patent B2
US 12,236,355 · App. 18/416,379 · Granted Feb 25, 2025

Generating machine-learning model for document extraction

Inventors: Michal Gdak (Warsaw, PL); Ganeshan Ramachandran Iyer (Redmond, WA); Tomasz Malisz (Bialystok, PL); Mikolaj Niedbala (Poznan, PL); Pawel Pollak (Warsaw, PL); Saurin Shah (Kirkland, WA); Jan Tomasz Topinski (Izabelin, PL); Daria Wieteska (Warsaw, PL)
Assignee: Snowflake Inc.
G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,355
App. No.
18/416,379
Granted
Feb 25, 2025
Kind
B2
Abstract

Systems and methods for generating a machine-learning (ML) model for extracting information from one or more electronic documents, where the ML model can be used as a data object, which can be part of a database command or as part of a document information extraction process that is continuously running (e.g., document information extraction pipeline).

Claims (53)

1. A system comprising:

at least one hardware processor; and

at least one memory storing instructions that cause the at least one hardware processor to perform operations comprising:

causing presentation of a user interface for training a select machine-learning model, the select machine-learning model being configured to extract values for one or more data points from electronic documents;

adding, by the user interface, a set of data points and a set of questions, the set of questions corresponding to the set of data points;

using the select machine-learning model to extract, from an uploaded electronic document, at least a set of values for the set of data points based on the set of questions;

causing presentation of the set of data points, the set of values, and the set of questions in the user interface;

receiving, by the user interface, user feedback with respect to one or more of the set of values;

performing a training process on the select machine-learning model to generate a custom machine-learning model from the select machine-learning model based on the user feedback and the uploaded electronic document;

publishing the custom machine-learning model as a database object on a data platform for use in extracting values for at least the set of data points from one or more electronic documents;

receiving a database command that generates a document information extraction pipeline based on the database object; and

in response to the database command, generating the document information extraction pipeline on the data platform based on the database object, the document information extraction pipeline comprising a software service that is continuously running on the data platform and that is configured to perform operations comprising:

monitoring for a set of input electronic documents;

receiving the set of input electronic documents;

using the custom machine-learning model of the database object to extract a set of extracted values from each input electronic document in the set of input electronic documents; and

storing each set of extracted values in a target table on the data platform, the target table being specified by the database command.

2. The system of claim 1 , wherein the database command is a Structured Query Language (SQL) query.

3. The system of claim 1 , wherein the database command comprises a create command.

4. The system of claim 1 , wherein the target table comprises a column for file location for each individual input electronic document, and a separate column for each data point of the set of data points for which at least one value is extracted from at least one input electronic document.

5. A machine-storage medium comprising instructions that, when executed by one or more processors of a machine, configure the machine to perform operations comprising:

causing presentation of a user interface for training a select machine-learning model, the select machine-learning model being configured to extract values for one or more data points from electronic documents;

adding, by the user interface, a set of data points and a set of questions, the set of questions corresponding to the set of data points;

using the select machine-learning model to extract, from an uploaded electronic document, at least a set of values for the set of data points based on the set of questions;

causing presentation of the set of data points, the set of values, and the set of questions in the user interface;

receiving, by the user interface, user feedback with respect to one or more of the set of values;

performing a training process on the select machine-learning model to generate a custom machine-learning model from the select machine-learning model based on the user feedback and the uploaded electronic document;

publishing the custom machine-learning model as a database object on a data platform for use in extracting values for at least the set of data points from one or more electronic documents;

receiving a database command that generates a document information extraction pipeline based on the database object; and

in response to the database command, generating the document information extraction pipeline on the data platform based on the database object, the document information extraction pipeline comprising a software service that is continuously running on the data platform and that is configured to perform operations comprising:

monitoring for a set of input electronic documents;

receiving the set of input electronic documents;

using the custom machine-learning model of the database object to extract a set of extracted values from each input electronic document in the set of input electronic documents; and

storing each set of extracted values in a target table on the data platform, the target table being specified by the database command.

6. The machine-storage medium of claim 5 , wherein the database command is a Structured Query Language (SQL) query.

7. The machine-storage medium of claim 5 , wherein the database command comprises a create command.

8. The machine-storage medium of claim 5 , wherein the target table comprises a column for file location for each individual input electronic document, and a separate column for each data point of the set of data points for which at least one value is extracted from at least one input electronic document.

9. A method comprising:

causing presentation, by one or more hardware processors, of a user interface for training a select machine-learning model, the select machine-learning model being configured to extract values for one or more data points from electronic documents;

adding, by the one or more hardware processors and by the user interface, a set of data points and a set of questions, the set of questions corresponding to the set of data points;

using, by the one or more hardware processors, the select machine-learning model to extract, from an uploaded electronic document, at least a set of values for the set of data points based on the set of questions;

causing presentation, by the one or more hardware processors, of the set of data points, the set of values, and the set of questions in the user interface;

receiving, by the one or more hardware processors and by the user interface, user feedback with respect to one or more of the set of values;

performing, by the one or more hardware processors, a training process on the select machine-learning model to generate a custom machine-learning model from the select machine-learning model based on the user feedback and the uploaded electronic document;

publishing, by the one or more hardware processors, the custom machine-learning model as a database object on a data platform for use in extracting values for at least the set of data points from one or more electronic documents;

receiving, by the one or more hardware processors, a database command that generates a document information extraction pipeline based on the database object; and

in response to the database command, generating the document information extraction pipeline on the data platform based on the database object, the document information extraction pipeline comprising a software service that is continuously running on the data platform and that is configured to perform operations comprising:

monitoring for a set of input electronic documents;

receiving the set of input electronic documents;

using the custom machine-learning model of the database object to extract a set of extracted values from each input electronic document in the set of input electronic documents; and

storing each set of extracted values in a target table on the data platform, the target table being specified by the database command.

10. The method of claim 9 , wherein the database command is a Structured Query Language (SQL) query.

11. The method of claim 9 , wherein the database command comprises a create command.

12. The method of claim 9 , wherein the target table comprises a column for file location for each individual input electronic document, and a separate column for each data point of the set of data points for which at least one value is extracted from at least one input electronic document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2024
From: GDAK, MICHAL; IYER, GANESHAN RAMACHANDRAN; MALISZ, TOMASZ; NIEDBALA, MIKOLAJ; POLLAK, PAWEL; SHAH, SAURIN; TOPINSKI, JAN TOMASZ; WIETESKA, DARIA
To: SNOWFLAKE INC.
Reel/Frame 066170/0895 →
Continuity (3)
Continuation 18472883 · Sep 22, 2023
Provisional Application 63495174 · Apr 10, 2023
Related Publication 20240338577A1 · Oct 10, 2024
References Cited (39)
US 10031912B2 · Allen et al. · 2018 [cited by applicant]
US 10521464B2 · Juneja · 2019 [cited by examiner]
US 11178389B2 · Sinha et al. · 2021 [cited by applicant]
US 11194963B1 · Schafer · 2021 [cited by examiner]
US 11573936B2 · Stolze · 2023 [cited by examiner]
US 11768884B2 · Benincasa · 2023 [cited by examiner]
US 11922328B1 · Gdak et al. · 2024 [cited by applicant]
US 20060218132A1 · Mukhin · 2006 [cited by examiner]
US 20100235451A1 · Yu et al. · 2010 [cited by applicant]
US 20110125511A1 · Bakst · 2011 [cited by examiner]
US 20140278916A1 · Nukala et al. · 2014 [cited by applicant]
US 20150120634A1 · Tateno · 2015 [cited by applicant]
US 20200026913A1 · Northrup · 2020 [cited by examiner]
US 20200151591A1 · Li · 2020 [cited by applicant]
US 20200402013A1 · Yeung · 2020 [cited by examiner]
US 20210089563A1 · Grabau et al. · 2021 [cited by applicant]
US 20210182661A1 · Li · 2021 [cited by examiner]
US 20210209500A1 · Hu · 2021 [cited by examiner]
US 20210217209A1 · Elder · 2021 [cited by examiner]
US 20210350258A1 · Divakarmurthy et al. · 2021 [cited by applicant]
US 20210406815A1 · Mimassi · 2021 [cited by examiner]
US 20220309109A1 · Benincasa et al. · 2022 [cited by applicant]
US 20220327138A1 · Benincasa et al. · 2022 [cited by applicant]
US 20230186104A1 · Noble · 2023 [cited by examiner]
US 20230336228A1 · Chintalapudi · 2023 [cited by examiner]
US 20230385249A1 · Bensberg · 2023 [cited by examiner]
US 20240193142A1 · Walzer · 2024 [cited by examiner]
US 20240338521A1 · Gdak et al. · 2024 [cited by applicant]
CN 118779368A · 2024 [cited by applicant]
WO WO2024215671A1 · 2024 [cited by applicant]
“U.S. Appl. No. 18/472,883, Examiner Interview Summary mailed Dec. 12, 2023”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 18/472,883, Non Final Office Action mailed Nov. 7, 2023”, 41 pgs. [cited by applicant]
“U.S. Appl. No. 18/472,883, Notice of Allowance mailed Jan. 4, 2024”, 7 pgs. [cited by applicant]
“U.S. Appl. No. 18/472,883, Response filed Dec. 8, 2023 to Non Final Office Action mailed Nov. 7, 2023”, 11 pgs. [cited by applicant]
“European Application Serial No. 24169385.2, Extended European Search Report mailed Aug. 26, 2024”, 8 pgs. [cited by applicant]
“European Application Serial No. 24169385.2, Invitation to remedy deficiencies (R. 58 EPC) mailed Apr. 25, 2024”, 3 pgs. [cited by applicant]
“European Application Serial No. 24169385.2, Response filed Jun. 24, 2024 to Invitation to remedy deficiencies (R. 58 EPC) mailed Apr. 25, 2024”, 9 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/023730, International Search Report mailed Jul. 31, 2024”, 2 pgs. [cited by applicant]
“International Application Serial No. PCT/US2024/023730, Written Opinion mailed Jul. 31, 2024”, 6 pgs. [cited by applicant]