IP Library Granted Patent US 12705511
Granted Patent B2
US 12705511 · App. 17/006,673 · Granted Aug 11, 2026

Automated taxonomy classification system

Inventors: Melania Calinescu (San Francisco, CA); Xuexin Ren (Millbrae, CA); Han Liu (South San Francisco, CA)
Assignee: Data.ai Inc.
G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705511
App. No.
17/006,673
Granted
Aug 11, 2026
Kind
B2
Abstract

A taxonomy classification system assigns taxonomy labels to content items of an online system. To assign the taxonomy labels, the taxonomy classification system applies one or more taxonomy model to the content items to determine scores or probabilities that a particular label applies to the content item. Each taxonomy model includes multiple sub-models. Each sub-model corresponds to a different type of information for the content item. For example, a first sub-model corresponds to a description of the content item, a second sub-model corresponds to metrics of the content item in one or more content item publishers, a third sub-model corresponds to similar content items to the content item being evaluated. The taxonomy classification system combines the output from every sub-model to determine a label score for one or more labels in a label class and a taxonomy label from the label class is selected based on the determined label score.

Claims (51)

1 . A method for classifying a mobile device application, the method comprising:

assigning one or more taxonomy labels to the mobile device application, comprising:

selecting a taxonomy label using a first trained model, comprising:

applying a first sub-model of the first trained model based on a description of the mobile device application;

applying a second sub-model of the first trained model based on metrics information for the mobile device application from one or more mobile device application publishers, wherein applying the second sub-model based on metrics information comprises:

identifying categories of metrics information that are available for the mobile device application;

selecting a single version of the second sub-model based on the categories of metrics information that are available for the mobile device application, the single version of the second sub-model selected from a plurality of versions of the second sub-model, each version of the plurality of versions of the second sub-model trained using metrics information corresponding to different combinations of categories of metrics information, wherein selecting the version of the second sub-model comprises: determining a set of feature categories on which each of the plurality of versions of the second sub-model were trained, and selecting the single version of the plurality of versions based on a priority of a match of the available categories of metrics information of the mobile device application to the sets of feature categories used to train each of the versions of the second sub-model and based on whether all categories for the version of the second sub-model are in the set of available categories for the mobile device application, the priority indicating a relative quantity of matching categories between the version of the second sub-model and the mobile device application relative to quantities for the plurality of versions of the second sub-model; and

applying the selected single version of the second sub-model to generate metrics scores for one or more labels in a label class;

applying a third sub-model of the first trained model based on a list of content items having one or more characteristics that are also had by the mobile device application;

combining an output of the first sub-model, an output of the second sub-model, and an output of the third sub-model to generate label scores for one or more labels in a label class; and

selecting the taxonomy label from the label class based on the generated label scores.

2 . The method of claim 1 , further comprising:

assigning one or more tags to the mobile device application, comprising for each tag of a plurality of tags:

applying a corresponding model to determine a likelihood that a tag applies to the mobile device application; and

responsive to determining that the likelihood that the tag applies to the mobile device application is above a threshold value, assigning the tag to the mobile device application.

3 . The method of claim 1 , wherein a taxonomy of the one or more taxonomy labels has multiple levels, wherein the selected taxonomy label is a first-level taxonomy label corresponding to a first level of the taxonomy, and wherein assigning the one or more taxonomy labels to the mobile device application further comprises:

selecting a second-level model from a plurality of second-level models, the second-level model selected based on the selected first-level taxonomy label; and

selecting a second-level taxonomy label corresponding to a second level of the taxonomy, the second-level taxonomy label from a set of second-level taxonomy labels associated with the selected first-level taxonomy label, the second-level taxonomy label selected using the selected second-level model.

4 . The method of claim 3 , wherein selecting the second-level model from the plurality of second-level models comprises:

selecting the second-level model corresponding to the selected first-level taxonomy label, each second level-model from the plurality of second-level models corresponding to a different taxonomy label in the first level of the taxonomy.

5 . The method of claim 1 , wherein selecting the version of the second sub-model comprises:

filtering the plurality of versions of the second sub-model based on the categories of metrics information that are available for the mobile device application; and

selecting the version of the second sub-model with a highest priority.

6 . The method of claim 1 , wherein assigning the one or more taxonomy labels to the mobile device application further comprises:

determining a confidence score for the taxonomy label; and

assigning the selected taxonomy label to the mobile device application responsive to the confidence score being above a threshold value.

7 . The method of claim 6 , wherein determining the confidence score for the taxonomy label comprises:

generating a selection score for each taxonomy label in a label class;

identifying a highest selection score from the generated selection scores of each taxonomy label in the label class;

identifying a second highest selection score from the generated selection scores of each taxonomy label in the label class;

determining a pre-confidence score based on a difference between the highest selection score and the second highest selection score; and

determining the confidence score based on the pre-confidence score, the confidence score determined using estimated parameters calculated by fitting a probability curve to a training dataset.

8 . The method of claim 7 , wherein the estimated parameters are calculated by fitting an exponential curve to the training dataset.

9 . The method of claim 7 , wherein assigning the selected taxonomy label to the mobile device application responsive to the confidence score being above the threshold value comprises:

responsive to the confidence score being above the threshold value, assigning a taxonomy label with the highest selection score to the mobile device application.

10 . The method of claim 6 , further comprising:

responsive to the confidence score being below the threshold value, sending the mobile device application for manual classification.

11 . A method for applying a trained model to a mobile device application, comprising:

identifying a set of available feature categories for the mobile device application;

selecting a single version of the trained model based on the identified set of available feature categories for the mobile device application, the single version of the trained model selected from a plurality of versions of the trained model, each version in the plurality of versions of the trained model trained using a different set of feature categories of content items in a training dataset, wherein selecting the version of the trained model comprises:

determining a set of feature categories on which each of the plurality of versions of the trained model were trained; and

selecting the single version of the plurality of versions based on a priority of a match of the available feature categories of the mobile device application to the sets of feature categories used to train each of the versions of the trained model and based on whether all feature categories for the version of the trained model are in the set of available feature categories for the mobile device application, the priority indicating a relative quantity of matching feature categories between the version of the trained model and the mobile device application relative to quantities for the plurality of versions of the trained model; and

applying the selected single version of the trained model.

12 . The method of claim 11 , wherein selecting the version of the trained model further comprises:

responsive to determining that the available features for the mobile device application includes every feature category used to train a first version of the trained model, selecting the first version of the trained model.

13 . The method of claim 11 , wherein the selected version of the trained model is trained using more feature categories than a second version of the trained model, and wherein the selected version of the trained model is more accurate than the second version of the trained model.

14 . The method of claim 11 , wherein selecting the single version of the plurality of versions based on the priority of the match of the available features of the mobile device application to the sets of feature categories used to train each of the versions of the trained model includes:

determining whether the set of available features for the mobile device application includes every feature category used to train a first version of the trained model; and

responsive to determining that the available features for the mobile device application do not include every feature category used to train the first version of the trained model:

determining that the set of available features for the mobile device application includes every feature category used to train one or more other versions of the trained model, the one or more other versions of the trained model having lower priority levels than the first version of the trained model; and

selecting a second version of the trained model, as opposed to the first version of the trained model, to be the selected version of the trained model, the second version being selected from among the one or more other versions based on the second version having a highest priority from among the one or more other versions.