IP Library › Granted Patent US 10,635,678
Granted Patent B2
US 10,635,678 · App. 15/538,727 · Granted Apr 28, 2020

Method and apparatus for processing search data

Inventors: Pengjun Xie (Hangzhou, CN); Xin Zhou (Hangzhou, CN); Jun Lang (Hangzhou, CN)
Assignee: ALIBABA GROUP HOLDING LIMITED
G06F16/24578G06F16/9535
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,635,678
App. No.
15/538,727
Granted
Apr 28, 2020
Kind
B2
Abstract

The disclosure provides a method and apparatus for processing search data. For a historical search query that includes a knowledge requirement, the disclosure mines entity information for the historical search query and uses that as an answer recommended to users. Thus, the accuracy of entity information recommended to users is improved, and the current problem of poor search results for a historical search query that includes a knowledge requirement is solved.

Claims (36)

1. A method comprising:

acquiring, by a processor, search result information associated with a historical search query, the historical search query including a knowledge requirement, the knowledge requirement including a shopping query for information and the search result information including text content, a website identifier, the text content comprising a number of supporters and a number of opponents of an answer to the shopping query;

extracting, by the processor, candidate entity information from the search result information based on a type of the shopping query, wherein the candidate entity information corresponds to the historical search query associated with the search result information;

determining, by the processor, that a subset of the candidate entity information is entity information associated with the historical search query based on the search result information;

receiving, by the processor over a network, a current search query from a user, the current search query including the knowledge requirement;

identifying, by the processor, the historical search query as corresponding to the current search query; and

transmitting, by the processor over the network, the entity information corresponding to the historical search query to the user in a response to the current search query.

2. The method of claim 1 wherein acquiring search result information associated with a historical search query comprises:

identifying a type of the historical search query based on text content included in the historical search query;

identifying a method for extracting candidate entity information based on the type of the historical search query; and

extracting candidate entity information using the method for extracting candidate entity information.

3. The method of claim 2 wherein identifying a type of the historical search query based on text content included in the historical search query comprises identifying a presence of one or more pre-defined n-grams or patterns.

4. The method of claim 1 wherein extracting candidate entity information from the search result information comprises extracting candidate entity information from the answer included within the text content.

5. The method of claim 1 wherein extracting candidate entity information from the search result information further comprises screening the candidate entity information and selecting a subset of the candidate entity information.

6. The method of claim 1 wherein extracting candidate entity information from the search result information further comprises scoring the candidate entity information and selecting, as the entity information, a highest scoring subset of the candidate entity information.

7. The method of claim 6 wherein scoring the candidate entity information comprises scoring the candidate entity information based on a presence of an entity word appearing within an answer within text content of a website, a weight associated with the website, and a weight associated with the answer.

8. The method of claim 7 wherein the weight associated with the answer is determined based on a number of supporters of the answer and a number of opponents of the answer.

9. An apparatus comprising:

a processor; and

a non-transitory memory storing computer-executable instructions therein that, when executed by the processor, cause the apparatus to perform the operations of:

acquiring search result information associated with a historical search query, the historical search query including a knowledge requirement, the knowledge requirement including a shopping query for information and the search result information including text content, a website identifier, the text content comprising a number of supporters and a number of opponents of an answer to the shopping query;

extracting candidate entity information from the search result information based on a type of the shopping query, wherein the candidate entity information corresponds to the historical search query associated with the search result information;

determining that a subset of the candidate entity information is entity information associated with the historical search query based on the search result information;

receiving a current search query from a user over a network, the current search query including the knowledge requirement;

identifying the historical search query as corresponding to the current search query; and

transmitting the entity information corresponding to the historical search query to the user over the network in a response to the current search query.

10. The apparatus of claim 9 wherein acquiring search result information associated with a historical search query comprises:

identifying a type of the historical search query based on text content included in the historical search query;

identifying a method for extracting candidate entity information based on the type of the historical search query; and

extracting candidate entity information using the method for extracting candidate entity information.

11. The apparatus of claim 10 wherein identifying a type of the historical search query based on text content included in the historical search query comprises identifying a presence of one or more pre-defined n-grams or patterns.

12. The apparatus of claim 9 wherein extracting candidate entity information from the search result information comprises extracting candidate entity information from the answer included within the text content.

13. The apparatus of claim 9 wherein extracting candidate entity information from the search result information further comprises screening the candidate entity information and selecting a subset of the candidate entity information.

14. The apparatus of claim 9 wherein extracting candidate entity information from the search result information further comprises scoring the candidate entity information and selecting, as the entity information, a highest scoring subset of the candidate entity information.

15. The apparatus of claim 14 wherein scoring the candidate entity information comprises scoring the candidate entity information based on a presence of an entity word appearing within an answer within text content of a website, a weight associated with the website, and a weight associated with the answer.

16. The apparatus of claim 15 wherein the weight associated with the answer is determined based on a number of supporters of the answer and a number of opponents of the answer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2018
From: XIE, PENGJUN; ZHOU, XIN; LANG, JUN
To: ALIBABA GROUP HOLDING LIMITED
Reel/Frame 047272/0456 →
Priority Claims (1)
CN 2014 1 0836116 · Dec 23, 2014 · national
Continuity (1)
Related Publication 20180011857A1 · Jan 11, 2018
Cited By (1)
US 12,517,899