IP Library › Granted Patent US 10,437,894
Granted Patent B2
US 10,437,894 · App. 14/706,631 · Granted Oct 8, 2019

Method and system for app search engine leveraging user reviews

Inventors: Dae Hoon Park (San Jose, CA); Mengwen Liu (San Jose, CA); Lifan Guo (San Jose, CA)
Assignee: TCL RESEARCH AMERICA INC.
G06F16/951G06F16/3331G06F11/3409
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,437,894
App. No.
14/706,631
Granted
Oct 8, 2019
Kind
B2
Abstract

A method for an app search engine leveraging user reviews is provided. The method includes receiving an app search query from a user, determining a plurality of relevant apps based on the received app search query, and extracting app descriptions and user reviews associated with the plurality of relevant apps from an app database. The method also includes preprocessing the extracted app descriptions and user reviews of each of the plurality of relevant apps to generate a text corpus and creating a topic-based language model for each of the plurality of relevant apps based on the generated text corpus. Further, the method includes ranking a list of relevant apps using the topic-based language model and providing the ranked app list for the user.

Claims (731)

1. A method for an app search engine leveraging user

reviews, comprising:

receiving an app search query from a user;

based on the received app search query, determining a plurality of relevant apps;

extracting app descriptions and user reviews associated with the plurality of relevant apps from an app database;

preprocessing the extracted app descriptions and user reviews of each of the plurality of relevant apps to generate a text corpus;

based on the generated text corpus, creating a topic-based language model for each of the plurality of relevant apps, wherein the topic-based language model combines the app descriptions and the user reviews;

ranking a list of relevant apps using the topic-based language model, comprising:

removing noises in the user reviews;

scoring each given query word g in each topic associated with each of the plurality of relevant apps a;

calculating an app score for the app search query from scores of the query words, wherein a score function for q and a is defined as:

score

⁡

(

q

,

a

)

=

⁢

∏

w

∈

q

⁢

p

⁡

(

w

|

a

)

=

⁢

∏

w

∈

q

⁢

[

(

1

-

μ

)

⁢

p

⁡

(

w

|

d

)

+

μ

⁢

⁢

p

⁡

(

w

|

r

)

]

=

⁢

∏

w

∈

q

⁢

[

(

1

-

μ

)

⁢

(

(

1

-

κ

d

)

⁢

p

m

⁢

⁢

l

⁡

(

w

|

d

)

+

κ

d

⁢

p

⁡

(

w

|

D

)

)

+

⁢

μ

⁡

(

(

1

-

κ

r

)

⁢

p

m

⁢

⁢

l

⁡

(

w

|

r

)

+

κ

r

⁢

p

⁡

(

w

|

R

)

)

]

where D is a document corpus, and pml(w|d) and p(w|D) are estimated by a Maximum Likelihood Estimator (MLE), which makes

p

m

⁢

⁢

l

⁡

(

w

|

d

)

=

c

⁡

(

w

,

d

)

∑

w

′

⁢

c

⁡

(

w

′

,

d

)

⁢

⁢

and

⁢

⁢

p

⁡

(

w

|

D

)

=

c

⁡

(

w

,

D

)

∑

w

′

⁢

c

⁡

(

w

′

,

D

)

,

 p(w|R) is a background language model in all reviews R, and K d and K r are smoothing parameters for app descriptions and user reviews; and

ranking the list of relevant apps that are scored according to relevance to the received app search query, wherein a language model for an app a is created by linearly interpolating a unigram language model for the app descriptions d and the user reviews r, which is defined by:

p ( w|a )=(1−μ) p ( w|d )+μ p ( w|r ),

wherein w is a certain word in the app search query; and μ is a parameter to determine a proportion of a review language model in p(w|a);

retrieving, based on the unigram language model, the list of relevant apps by

the certain word not contained in the app descriptions or the user reviews; and providing the ranked app list for the user.

2. The method according to claim 1 , wherein providing the ranked app list for the user further includes:

formatting the ranked app list to be viewable by a mobile device used by the user.

3. The method according to claim 1 , wherein preprocessing the extracted app descriptions and user reviews of each of the plurality of relevant apps to generate a text corpus further includes:

normalizing content of the text corpus to a canonical form.

4. The method according to claim 1 , wherein:

the app score indicates strength of association between the query words and the app.

5. The method according to claim 1 , wherein:

provided that each document d contains N d words and a whole document collections build a word vocabulary V, the topic-based language model for each app a in each topic z is defined by:

p

lda

⁡

(

w

|

a

)

=

∑

z

=

1

K

⁢

⁢

p

⁡

(

w

|

z

,

W

d

,

Z

^

d

,

β

)

⁢

p

⁡

(

z

|

a

,

Z

^

d

,

Z

^

r

,

α

d

,

α

r

)

∝

∑

z

=

1

K

⁢

⁢

N

^

w

|

z

+

β

N

^

z

+

V

⁢

⁢

β

⁢

N

^

z

|

d

+

K

⁢

⁢

α

d

+

N

^

z

|

r

+

K

⁢

⁢

α

r

⁢

N

^

z

|

d

+

α

d

N

d

+

K

⁢

⁢

α

d

N

d

+

K

⁢

⁢

α

d

+

Σ

z

⁢

N

^

z

|

r

+

K

⁢

⁢

α

r

wherein α and β are symmetric prior vectors; w is a certain word in the app search query; W d is all the words in descriptions of all apps; K is a total number of all shared topics; {circumflex over (N)}with subscription is an estimated number of words satisfying subscription condition; and {circumflex over (Z)} d and {circumflex over (Z)} r are topics for the app descriptions and the user reviews estimated from app latent dirichlet allocation (AppLDA), respectively.

6. A system for an app search engine leveraging user reviews, comprising:

a memory;

a processor coupled to the memory;

a plurality of program units stored in the memory to be executed by the processor, the plurality of program units comprising:

a receiving module configured to receive an app search query from a user;

an extraction module configured to determine a plurality of relevant apps based on the received app search query and extract app descriptions and user reviews associated with the plurality of relevant apps from an app database;

a preprocessing module configured to perform preliminary processing for the extracted app descriptions and user reviews of each of the plurality of relevant apps to generate a text corpus;

a language model creating module configured to, based on the generated text corpus, create a topic-based language model for each of the plurality of relevant apps, wherein the topic-based language model combines the app descriptions and the user reviews;

a result ranking module configured to rank a list of relevant apps using the topic-based language model and provide the ranked app list to the user; and

an app scorer configured to score the given query word g in each topic associated with each app, and calculate each app score for the given query from the scores of the query words, wherein a language model for an app a is created by linearly interpolating a unigram language model for the app descriptions d and the user reviews r, which is defined by:

p ( w|a )=(1−μ) p ( w|d )+μ p ( w|r )

wherein w is a certain word in the app search query; and μ is a parameter to determine a proportion of a review language model in p(w|a), and the list of relevant apps is retrieved by the certain word not contained in the app descriptions or the user reviews based on the unigram language model; and

a score function for g and a is defined as:

score

⁡

(

q

,

a

)

=

⁢

∏

w

∈

q

⁢

p

⁡

(

w

|

a

)

=

⁢

∏

w

∈

q

⁢

[

(

1

-

μ

)

⁢

p

⁡

(

w

|

d

)

+

μ

⁢

⁢

p

⁡

(

w

|

r

)

]

=

⁢

∏

w

∈

q

⁢

[

(

1

-

μ

)

⁢

(

(

1

-

κ

d

)

⁢

p

m

⁢

⁢

l

⁡

(

w

|

d

)

+

κ

d

⁢

p

⁡

(

w

|

D

)

)

+

⁢

μ

⁡

(

(

1

-

κ

r

)

⁢

p

m

⁢

⁢

l

⁡

(

w

|

r

)

+

κ

r

⁢

p

⁡

(

w

|

R

)

)

]

wherein D is a document corpus, and pml(w|d) and p(w|D) are estimated by a Maximum Likelihood Estimator (MLE), which makes

p

m

⁢

⁢

l

⁡

(

w

|

d

)

=

c

⁡

(

w

,

d

)

∑

w

′

⁢

c

⁡

(

w

′

,

d

)

⁢

⁢

and

⁢

⁢

p

⁡

(

w

|

D

)

=

c

⁡

(

w

,

D

)

∑

w

′

⁢

c

⁡

(

w

′

,

D

)

,

 p(w|R) is a background language model in all reviews R, and K d and K r are smoothing parameters for app descriptions and user reviews.

7. The system according to claim 6 , wherein:

the ranked app list is formatted to be viewable by a mobile device used by the user.

8. The system according to claim 6 , wherein:

the preprocessing module is further configured to normalize content of the text corpus to a canonical form.

9. The system according to claim 6 , wherein:

the app score indicates strength of association between the query words and the app.

10. The system according to claim 6 , wherein:

provided that each document d contains Nwords and a whole document collections build a word vocabulary V, the topic-based language model for each app a in each topic z is defined by:

p

lda

⁡

(

w

|

a

)

=

∑

z

=

1

K

⁢

⁢

p

⁡

(

w

|

z

,

W

d

,

Z

^

d

,

β

)

⁢

p

⁡

(

z

|

a

,

Z

^

d

,

Z

^

r

,

α

d

,

α

r

)

∝

∑

z

=

1

K

⁢

⁢

N

^

w

|

z

+

β

N

^

z

+

V

⁢

⁢

β

⁢

N

^

z

|

d

+

K

⁢

⁢

α

d

+

N

^

z

|

r

+

K

⁢

⁢

α

r

⁢

N

^

z

|

d

+

α

d

N

d

+

K

⁢

⁢

α

d

N

d

+

K

⁢

⁢

α

d

+

Σ

z

⁢

N

^

z

|

r

+

K

⁢

⁢

α

r

wherein α and β are symmetric prior vectors; w is a certain word in the app search query; W d is all the words in descriptions of all apps; K is a total number of all shared topics; {circumflex over (N)}with subscription is an estimated number of words satisfying subscription condition; and {circumflex over (Z)} d and {circumflex over (Z)} r are topics for the app descriptions and the user reviews estimated from app latent dirichlet allocation (AppLDA), respectively.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2026
From: TCL RESEARCH AMERICA INC.
To: HONGFA GLOBAL LIMITED
Reel/Frame 075814/0200 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2015
From: PARK, DAE HOON; LIU, MENGWEN; GUO, LIFAN
To: TCL RESEARCH AMERICA INC.
Reel/Frame 035589/0642 →
Continuity (1)
Related Publication 20160328403A1 · Nov 10, 2016