IP Library › Granted Patent US 9,355,414
Granted Patent B2
US 9,355,414 · App. 12/790,844 · Granted May 31, 2016

Collaborative filtering model having improved predictive performance

Inventors: Martin B. Scholz (San Francisco, CA); George Forman (Port Orchard, WA); Rong Pan (Guangdong, CN)
Assignee: Hewlett Packard Enterprise Development LP
G06Q30/0282G06Q30/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,355,414
App. No.
12/790,844
Filed
May 30, 2010
Granted
May 31, 2016
Kind
B2
Art Unit
2166
USPC
707/748
Abstract

For each first entity of a subset of a number of first entities, an expected improvement of a predictive performance of a collaborative filtering model if additional ratings of the first entity in relation to a plurality of second entities were obtained is estimated. Particular first entities from the subset of the first entities of which to obtain the additional ratings in relation to the second entities are selected based at least on the expected improvements that have been determined. The additional ratings of the particular first entities in relation to the second entities are obtained.

Claims (67)

1. A method comprising:

for each first entity of a subset of a plurality of first entities, estimating by a computing device an expected improvement of a predictive performance of a collaborative filtering model constructed from training data including a plurality of ratings of the first entity in relation to the second entities and that has a predictive accuracy resulting from being tested against testing data, if additional ratings of the first entity in relation to a plurality of second entities were obtained, by:

removing a number of the ratings of the first entity in relation to the second entities from the training data;

constructing a reduced collaborative filtering model from the training data from which the number of the ratings have been removed;

testing the reduced collaborative filtering model against the testing data to determine a predictive accuracy thereof;

determining the expected improvement based on a degradation of the predictive accuracy of the reduced collaborative filtering model compared to the predictive accuracy of the collaborating filtering model;

selecting by a computing device particular first entities from the subset of the first entities of which to obtain the additional ratings in relation to the second entities, based at least on the expected improvements that have been determined; and

obtaining a number of the additional ratings of the particular first entities in relation to the second entities based on the number of the ratings that were removed from the training data.

2. The method of claim 1 , wherein the first entities are items and the second entities are users,

such that the expected improvement of the predictive performance of the collaborative filtering model if additional ratings for each item were obtained is estimated,

such that particular items are selected for which to obtain the additional ratings, based at least on the expected improvements that have been determined,

and such that the additional ratings for the particular items are obtained.

3. The method of claim 1 , wherein the first entities are users and the second entities are items,

such that the expected improvement of the predictive performance of the collaborative filtering model if additional ratings were obtained from each user is estimated,

such that particular users are selected from whom to obtain the additional ratings, based at least on the expected improvements that have been determined,

and such that the additional ratings are obtained from the particular users.

4. The method of claim 1 , wherein the additional ratings of just the particular first entities in relation to the second entities are obtained, such that the additional ratings of the first entities that are not the particular first entities are not obtained.

5. The method of claim 1 , wherein, for each given first entity, estimating the expected improvement of the predictive performance of the collaborative filtering model comprises:

dividing data having a number of ratings of the first entities in relation to the second entities into training data and test data, the test data including just ratings of the given first entity in relation to the second entities;

repeating a plurality of times,

determining a predictive performance of the collaborative filtering model for the given first entity using the training data and the test data, to yield a data point of the predictive performance of the collaborative filtering model for the given first entity by the number of ratings of the given first entity;

removing some of the ratings of the given first entity from the training data;

fitting a regression model that approximates the predictive performance of the collaborative filtering model for the given first entity as a function of the number of ratings of the given first entity in relation to the second entities, using the data points; and,

determining the expected improvement of the predictive performance of the collaborative filtering model for the given first entity based on the regression model.

6. The method of claim 5 , wherein, for each given first entity, estimating the expected improvement of the predictive performance of the collaborative filtering model further comprises:

modifying the expected improvement of the predictive performance of the collaborative filtering model for the given first entity based on an importance of the given first entity, to yield the expected improvement of the predictive performance of the collaborative filtering model as a whole.

7. The method of claim 6 , wherein modifying the expected improvement of the predictive performance of the collaborative filtering model based on the importance of the given first entity comprises multiplying the expected improvement by the importance of the given first entity.

8. The method of claim 1 , wherein, for each given first entity, estimating the expected improvement of the predictive performance of the collaborative filtering model comprises:

dividing data having a number of ratings of multiple first entities in relation to multiple second entities into training data and test data;

repeating a plurality of times,

determining a predictive performance of the collaborative filtering model as a whole using the training data and the test data, to yield a data point of the predictive performance of the collaborative filtering model by the number of ratings of the given first entity;

removing some of the ratings of the given first entity from the training data;

fitting a model that approximates the predictive performance of the collaborative filtering model as a whole as a function of the number of ratings of the given first entity in relation to the second entities, using the data points; and,

determining the expected improvement of the predictive performance of the collaborative filtering model as a whole based on the model.

9. The method of claim 8 , wherein determining the predictive performance of the collaborative filtering model as a whole using the training data and the test data comprises:

training the collaborative filtering model using the training data;

determining a predictive performance of the collaborative filtering model for each first entity; and,

combining the predictive performance of the collaborative filtering model for each first entity to yield the predictive performance of the collaborative filtering model as a whole.

10. The method of claim 9 , wherein determining the predictive performance of the collaborative filtering model as a whole using the training data and the test data further comprises, prior to combining the predictive performance of the collaborative filtering model for each first entity:

modifying the predictive performance of the collaborative filtering model for each first entity based on an importance of the first entity.

11. The method of claim 1 , wherein the first entities are items and the second entities are users, such that the particular first entities are particular items,

and wherein obtaining the additional ratings of the particular entities in relation to the second entities comprises querying the users for the additional ratings for the particular items using a round robin-based approach.

12. The method of claim 1 , wherein the first entities are items and the second entities are users, such that the particular first entities are particular items,

and wherein obtaining the additional ratings of the particular entities in relation to the second entities comprises, for each user:

determining specific items of the particular items for which the user is likely to be able or willing to provide the additional ratings; and,

querying the user for the additional ratings for just the specific items.

13. The method of claim 1 , further comprising:

training the collaborative filtering model using a plurality of ratings of the particular first entities in relation to the second entities, including the additional ratings that have been obtained; and,

using the collaborative filtering model, as trained, to make predictions.

14. A system comprising:

a processor;

a computer-readable data storage medium to store one or more computer programs that are executable by the processor;

a first component implemented by the computer programs to, for each first entity of a subset of a plurality of first entities, estimate an expected improvement of a predictive performance of a collaborative filtering model constructed from training data including a plurality of ratings of the first entity in relation to the second entities and that has a predictive accuracy resulting from being tested against testing data, if additional ratings of the first entity in relation to a plurality of second entities were obtained, by:

removing a number of the ratings of the first entity in relation to the second entities from the training data;

constructing a reduced collaborative filtering model from the training data from which the number of the ratings of the first entity in relation to the second entities have been removed;

testing the reduced collaborative filtering model against the testing data to determine a predictive accuracy of the reduced collaborative filtering model;

determining the expected improvement based on a degradation of the predictive accuracy of the reduced collaborative filtering model compared to the predictive accuracy of the collaborative filtering model;

a second component implemented by the computer programs to select particular first entities from the subset of the first entities of which to obtain the additional ratings in relation to the second entities, based at least on the expected improvements that have been determined; and,

a third component implemented by the computer programs to obtain a number of the additional ratings of the particular first entities in relation to the second entities based on the number of the ratings that were removed from the training data.

15. A non-transitory computer-readable data storage medium having one or more computer programs stored thereon, wherein execution of the computer programs by a processor causes a method to be performed, the method comprising:

for each first entity of a subset of a plurality of first entities, estimating by a computing device an expected improvement of a predictive performance of a collaborative filtering model constructed from training data including a plurality of ratings of the first entity in relation to the second entities and that has a predictive accuracy resulting from being tested against testing data, if additional ratings of the first entity in relation to a plurality of second entities were obtained, by:

removing a number of the ratings of the first entity in relation to the second entities from the training data;

constructing a reduced collaborative filtering model from the training data from which the number of the ratings have been removed;

testing the reduced collaborative filtering model against the testing data to determine a predictive accuracy thereof;

determining the expected improvement based on a degradation of the predictive accuracy of the reduced collaborative filtering model compared to the predictive accuracy of the collaborating filtering model;

selecting by a computing device particular first entities from the subset of the first entities of which to obtain the additional ratings in relation to the second entities, based at least on the expected improvements that have been determined; and

obtaining a number of the additional ratings of the particular first entities in relation to the second entities based on the number of the ratings that were removed from the training data.

Assignments (8)
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0577 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC)
Reel/Frame 063560/0001 →
RELEASE OF SECURITY INTEREST REEL/FRAME 044183/0718 Recorded Feb 2, 2023
From: JPMORGAN CHASE BANK, N.A.
To: MICRO FOCUS LLC (F/K/A ENTIT SOFTWARE LLC); BORLAND SOFTWARE CORPORATION; MICRO FOCUS (US), INC.; SERENA SOFTWARE, INC; ATTACHMATE CORPORATION; MICRO FOCUS SOFTWARE INC. (F/K/A NOVELL, INC.); NETIQ CORPORATION
Reel/Frame 062746/0399 →
CHANGE OF NAME Recorded Aug 8, 2019
From: ENTIT SOFTWARE LLC
To: MICRO FOCUS LLC
Reel/Frame 050004/0001 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ATTACHMATE CORPORATION; BORLAND SOFTWARE CORPORATION; NETIQ CORPORATION; MICRO FOCUS (US), INC.; MICRO FOCUS SOFTWARE, INC.; ENTIT SOFTWARE LLC; ARCSIGHT, LLC; SERENA SOFTWARE, INC.
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0718 →
SECURITY INTEREST Recorded Oct 11, 2017
From: ENTIT SOFTWARE LLC; ARCSIGHT, LLC
To: JPMORGAN CHASE BANK, N.A.
Reel/Frame 044183/0577 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2017
From: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
To: ENTIT SOFTWARE LLC
Reel/Frame 042746/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2015
From: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
To: HEWLETT PACKARD ENTERPRISE DEVELOPMENT LP
Reel/Frame 037079/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2010
From: SCHOLZ, MARTIN B.; FORMAN, GEORGE; PAN, RONG
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 024806/0800 →
Continuity (1)
Related Publication 20110295762A1 · Dec 1, 2011