IP Library › Granted Patent US 11,775,864
Granted Patent B2
US 11,775,864 · App. 16/887,731 · Granted Oct 3, 2023

Feature management platform

Inventors: Frank Wisniewski (San Francisco, CA); Abhishek Jain (Mountain View, CA); Caio Vinicius Soares (Redwood City, CA); Tristan Cooper Baker (San Diego, CA); Joseph Brian Cessna (San Diego, CA)
Assignee: INTUIT, INC.
G06N20/00G06F16/254
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,864
App. No.
16/887,731
Granted
Oct 3, 2023
Kind
B2
Abstract

Certain aspects of the present disclosure provide techniques for operation of a feature management platform. A feature management platform is an end-to-end platform developed to manage the full lifecycle of data features. For example, to create a feature the feature management platform can receive a processing artifact (e.g., a configuration file and code fragment) from a computing device. The processing artifact defines the feature, including the data source to retrieve event data from, when to retrieve the event data, the type of transform to apply, etc. Based on the processing artifact, the feature management system generates a processing job, which when initiated generates a vector that encapsulates the feature data. The vector is transmitted to the computing device that locally hosts a model, which generates a prediction. The prediction is transmitted to the feature management platform and can be transmitted to other computing devices, upon request.

Claims (76)

1. A method, comprising:

receiving, at a feature management platform from a first computing device, a processing artifact that defines a feature associated with a data source and a transform;

generating, based on the processing artifact, a processing job configured to retrieve event data from the data source;

initiating the processing job that includes:

retrieving the event data from the data source;

applying the transform to the event data to generate a set of feature values; and

encapsulating the set of feature values within a feature vector for the feature to store in a feature store of the feature management platform;

storing the feature vector in the feature store for a certain amount of time, wherein the feature store is a fast-retrieval database, and wherein the feature management platform is associated with a dual storage system comprising the fast-retrieval database for storing features for the certain amount of time and a separate training data database for storing respective features for longer than the certain amount of time;

storing feature metadata associated with the feature vector in a feature registry;

transmitting the feature vector representing the set of feature values to a model hosted on the first computing device;

receiving, at the feature management platform, from the first computing device, a prediction generated by the model;

transmitting, by the feature management platform, the prediction to the feature store of the feature management platform; and

upon receiving a request from a second computing device for the feature within the certain amount of time after the storing of the feature vector in the feature store:

determining that the feature vector for the feature is stored in the fast-retrieval database of the dual storage system based on locating the feature metadata in the feature registry;

retrieving the feature vector from the fast-retrieval database based on the locating of the feature metadata in the feature registry; and

transmitting the feature vector to the second computing device from the fast-retrieval database without repeating, in response to the request, the retrieving of the event data, the applying of the transform, or the encapsulating of the set of feature values that were earlier performed to create the feature vector.

2. The method of claim 1 , further comprising:

receiving a request from the first computing device for the feature; and

determining the feature is not present in the feature store.

3. The method of claim 1 , further comprising:

storing the prediction in the feature store;

receiving a request from another computing device for the prediction; and

transmitting the prediction to the other computing device.

4. The method of claim 1 , further comprising retrieving, based on the processing artifact, the event data from a batch data source.

5. The method of claim 1 , further comprising retrieving, based on the processing artifact, the event data from a streaming data source.

6. The method of claim 1 , further comprising transmitting via an API of the feature management platform to the first computing device, an interface configured to receive an indication of a type of transform that is to be applied to the event data.

7. The method of claim 1 , wherein initiating the processing job includes:

implementing an aggregation operation to generate a set of aggregated feature values;

generating a stateful feature based on the set of aggregated feature values; and

encapsulating the stateful feature in a feature vector.

8. The method of claim 1 , further comprising transmitting the feature vector to a feature queue, wherein the feature queue is a write path to a multi-storage persistent layer in the feature management platform.

9. The method of claim 1 , wherein the processing artifact includes a configuration file and a code fragment.

10. A system, comprising:

a processor; and

a memory storing instructions, which when executed by the processor perform a method comprising:

receiving, at a feature management platform from a first computing device, a processing artifact that defines a feature associated with a data source and a transform;

generating, based on the processing artifact, a processing job configured to retrieve event data from the data source;

initiating the processing job that includes:

retrieving the event data from the data source;

applying the transform to the event data to generate a set of feature values; and

encapsulating the set of feature values within a feature vector for the feature to store in a feature store of the feature management platform;

storing the feature vector in the feature store, for a certain amount of time, wherein the feature store is a fast-retrieval database, and wherein the feature management platform is associated with a dual storage system comprising the fast-retrieval database for storing features for the certain amount of time and a separate training data database for storing respective features for longer than the certain amount of time;

storing feature metadata associated with the feature vector in a feature registry;

transmitting the feature vector representing the set of feature values to a model hosted on the first computing device;

receiving, at the feature management platform, from the first computing device, a prediction generated by the model;

transmitting, by the feature management platform, the prediction to the feature store of the feature management platform; and

upon receiving a request from a second computing device for the feature within the certain amount of time after the storing of the feature vector in the feature store:

determining that the feature vector for the feature is stored in the fast-retrieval database of the dual storage system based on locating the feature metadata in the feature registry;

retrieving the feature vector from the fast-retrieval database based on the locating of the feature metadata in the feature registry; and

transmitting the feature vector to the second computing device from the fast-retrieval database without repeating, in response to the request, the retrieving of the event data, the applying of the transform, or the encapsulating of the set of feature values that were earlier performed to create the feature vector.

11. The system of claim 10 , wherein the method further comprises:

receiving a request from the first computing device for the feature; and

determining the feature is not present in the feature store.

12. The system of claim 10 , wherein the method further comprises:

storing the prediction in the feature store;

receiving a request from another computing device for the prediction; and

transmitting the prediction to the other computing device.

13. The system of claim 10 , wherein the method further comprises retrieving, based on the processing artifact, the event data from a batch data source.

14. The system of claim 10 , wherein initiating the processing job includes:

implementing an aggregation operation to generate a set of aggregated feature values;

generating a stateful feature based on the set of aggregated feature values; and

encapsulating the stateful feature in a feature vector.

15. The system of claim 10 , wherein the method further comprises transmitting via an API of the feature management platform to the first computing device, an interface configured to receive an indication of a type of transform that is to be applied for application of the transform to the event data.

16. The system of claim 10 , wherein the feature defined in the processing artifact includes hierarchical feature data.

17. The system of claim 10 , wherein the method further comprises transmitting the feature vector to a feature queue, wherein the feature queue is a write path to a multi-storage persistent layer.

18. The system of claim 10 , wherein the processing artifact includes a configuration file and a code fragment.

19. A method, comprising:

receiving, at a feature management platform from a first computing device, a request for a feature;

determining that a feature store of the feature management platform includes the feature based on locating feature metadata that is associated with the feature in a feature registry, wherein the feature store is a fast-retrieval database, wherein the feature management platform is associated with a dual storage system comprising the fast-retrieval database for storing features for a certain amount of time and a separate training data database for storing respective features for longer than the certain amount of time, and wherein the request was received within the certain amount of time after the feature was stored in the feature store;

retrieving a feature vector of the feature from the fast-retrieval database on the locating of the feature metadata in the feature registry, and

transmitting the feature vector to a second computing device that hosts a model without repeating, in response to the request, a plurality of processing tasks that were earlier performed to create the feature vector;

receiving, the feature management platform, from the second computing device, a prediction generated by the model, based on the feature vector;

updating, by the feature management platform, the feature store based on the prediction;

transmitting the prediction to the feature store of the feature management platform; and

upon receiving a request from a third computing device for the prediction, transmitting the prediction to the third computing device from the feature store without regenerating the prediction based on the request from the third computing device.

20. The method of claim 19 , wherein the feature metadata in the feature registry indicates that the feature vector is stored in the feature store.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2020
From: WISNIEWSKI, FRANK; JAIN, ABHISHEK; SOARES, CAIO VINICIUS; BAKER, TRISTAN COOPER; CESSNA, JOSEPH BRIAN
To: INTUIT INC.
Reel/Frame 052791/0165 →
Continuity (1)
Related Publication 20210374600A1 · Dec 2, 2021