IP Library › Granted Patent US 10,884,828
Granted Patent B2
US 10,884,828 · App. 16/779,276 · Granted Jan 5, 2021

Synchronous ingestion pipeline for data processing

Inventors: Agostino Deligia (Dorval, CA); Cristian Viorel Suciu (Richmond Hill, CA)
Assignee: OPEN TEXT SA ULC
G06F9/544G06F16/164G06F2209/541
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,884,828
App. No.
16/779,276
Granted
Jan 5, 2021
Kind
B2
Abstract

A method for synchronous ingestion of input content may include determining, from an ingestion request, applicable ingestion pipeline components and an order by which the ingestion pipeline components are to be applied to input content; applying the ingestion pipeline components to the input content in the order determined from the ingestion request; updating a metadata file as the input content is processed by the ingestion pipeline components; and returning processed content, the metadata file, or both to a client device. The method may further include determining whether the ingestion request specifies a computing facility such as an indexer or a database downstream from the ingestion pipeline. If so, a processing result may be communicated to the computing facility for further processing. A server system may implement synchronous ingestion, asynchronous ingestion, or both.

Claims (53)

1. A method, comprising:

receiving, from a client device by a data ingestion pipeline application programming interface (DIP API) running on a server machine, an ingestion request for ingestion of data;

determining, by the DIP API, whether the ingestion request includes a query parameter;

responsive to the query parameter not being found in the ingestion request, directing the ingestion request to an asynchronous ingestion pipeline in which the data is processed through a predetermined workflow;

responsive to the query parameter being found in the ingestion request, extracting the query parameter and determining, by the DIP API using the query parameter, what ingestion pipeline components to call and in what order in a synchronous ingestion pipeline in which the data is processed through a dynamic workflow; and

starting, by the DIP API, the synchronous ingestion pipeline consisting of the ingestion pipeline components determined by the DIP API using the query parameter from the ingestion request, wherein the data is ingested through the synchronous ingestion pipeline in which the ingestion pipeline components apply different functions to the data in the order determined by the DIP API based on the ingestion request.

2. The method according to claim 1 , wherein the synchronous ingestion pipeline is one of a plurality of synchronous ingestion pipeline instances, the plurality of synchronous ingestion pipeline instances corresponding to multiple ingestion requests by multiple client devices communicatively connected to the DIP API.

3. The method according to claim 1 , wherein the data accompanying the ingestion request comprises one or more files and wherein each of the one or more files is associated with a metadata file.

4. The method according to claim 3 , further comprising:

updating the metadata file to include metadata provided by the ingestion pipeline components of the synchronous ingestion pipeline.

5. The method according to claim 4 , further comprising:

providing the metadata to an indexing engine, wherein the indexing engine uses the metadata to index the data;

storing the metadata in a relational database; or

returning the metadata file thus updated to the client device.

6. The method according to claim 1 , further comprising:

returning a text file to the client device, wherein the text file contains text extracted from the data as the data is processed through the synchronous ingestion pipeline.

7. The method according to claim 1 , wherein the ingestion pipeline components comprise at least one of a text extractor, an image analyzer, a text analyzer, a document converter, a conversion service, a language detector, or an optical character recognition processor.

8. A system, comprising:

a processor;

a non-transitory computer-readable medium; and

stored instructions translatable by the processor for:

receiving, from a client device, an ingestion request for ingestion of data;

determining whether the ingestion request includes a query parameter;

responsive to the query parameter not being found in the ingestion request, directing the ingestion request to an asynchronous ingestion pipeline in which the data is processed through a predetermined workflow;

responsive to the query parameter being found in the ingestion request, extracting the query parameter and determining, using the query parameter, what ingestion pipeline components to call and in what order in a synchronous ingestion pipeline in which the data is processed through a dynamic workflow; and

starting the synchronous ingestion pipeline consisting of the ingestion pipeline components determined using the query parameter from the ingestion request, wherein the data is ingested through the synchronous ingestion pipeline in which the ingestion pipeline components apply different functions to the data in the order determined based on the ingestion request.

9. The system of claim 8 , wherein the synchronous ingestion pipeline is one of a plurality of synchronous ingestion pipeline instances, the plurality of synchronous ingestion pipeline instances corresponding to multiple ingestion requests by multiple client devices communicatively connected to the system.

10. The system of claim 8 , wherein the data accompanying the ingestion request comprises one or more files and wherein each of the one or more files is associated with a metadata file.

11. The system of claim 10 , wherein the stored instructions are further translatable by the processor for:

updating the metadata file to include metadata provided by the ingestion pipeline components of the synchronous ingestion pipeline.

12. The system of claim 11 , wherein the stored instructions are further translatable by the processor for:

providing the metadata to an indexing engine, wherein the indexing engine uses the metadata to index the data;

storing the metadata in a relational database; or

returning the metadata file thus updated to the client device.

13. The system of claim 8 , wherein the stored instructions are further translatable by the processor for:

returning a text file to the client device, wherein the text file contains text extracted from the data as the data is processed through the synchronous ingestion pipeline.

14. The system of claim 8 , wherein the ingestion pipeline components comprise at least one of a text extractor, an image analyzer, a text analyzer, a document converter, a conversion service, a language detector, or an optical character recognition processor.

15. A computer program product comprising a non-transitory computer-readable medium storing instructions translatable by a processor of a server system for:

receiving, from a client device, an ingestion request for ingestion of data;

determining whether the ingestion request includes a query parameter;

responsive to the query parameter not being found in the ingestion request, directing the ingestion request to an asynchronous ingestion pipeline in which the data is processed through a predetermined workflow;

responsive to the query parameter being found in the ingestion request, extracting the query parameter and determining, using the query parameter, what ingestion pipeline components to call and in what order in a synchronous ingestion pipeline in which the data is processed through a dynamic workflow; and

starting the synchronous ingestion pipeline consisting of the ingestion pipeline components determined using the query parameter from the ingestion request, wherein the data is ingested through the synchronous ingestion pipeline in which the ingestion pipeline components apply different functions to the data in the order determined based on the ingestion request.

16. The computer program product of claim 15 , wherein the synchronous ingestion pipeline is one of a plurality of synchronous ingestion pipeline instances, the plurality of synchronous ingestion pipeline instances corresponding to multiple ingestion requests by multiple client devices communicatively connected to the server system.

17. The computer program product of claim 15 , wherein the data accompanying the ingestion request comprises one or more files and wherein each of the one or more files is associated with a metadata file.

18. The computer program product of claim 17 , wherein the instructions are further translatable by the processor for:

updating the metadata file to include metadata provided by the ingestion pipeline components of the synchronous ingestion pipeline.

19. The computer program product of claim 18 , wherein the instructions are further translatable by the processor for:

providing the metadata to an indexing engine, wherein the indexing engine uses the metadata to index the data;

storing the metadata in a relational database; or

returning the metadata file thus updated to the client device.

20. The computer program product of claim 15 , wherein the instructions are further translatable by the processor for:

returning a text file to the client device, wherein the text file contains text extracted from the data as the data is processed through the synchronous ingestion pipeline.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 12, 2020
From: DELIGIA, AGOSTINO; SUCIU, CRISTIAN VIOREL
To: OPEN TEXT SA ULC
Reel/Frame 051805/0485 →
Continuity (3)
Continuation 15906904 · Feb 27, 2018
Provisional Application 62465411 · Mar 1, 2017
Related Publication 20200167213A1 · May 28, 2020
Cited By (7)
US 12,242,892 US 12,423,309 US 12,566,758 US 12,645,704 US 12,651,001 US 12,695,681 US 12,717,852