IP Library Granted Patent US 11,714,840
Granted Patent B2
US 11,714,840 · App. 17/375,429 · Granted Aug 1, 2023

Method and apparatus for information query and storage medium

Inventor: Wanshun Chen (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
G06F16/3344G06F16/319G06F16/3329G10L15/02G10L15/04G10L15/18
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,714,840
App. No.
17/375,429
Granted
Aug 1, 2023
Kind
B2
Abstract

The present application discloses a method and an apparatus for information query, and an electronic device, which relates to a field of deep learning (DL), natural language processing (NLP) and artificial intelligence (AI) technology. The method includes: receiving a query sentence, segmenting the query sentence to obtain word segments, and obtaining a dependency relationship between two word segments and part of speech of the word segments; obtaining a coding sequence of the query sentence according to the dependency relationship and the part of speech of the word segments; matching the coding sequence with a generalized template to obtain a core corpus of the query sentence, wherein the generalized template comprises part of speech to be extracted and a dependency relationship to be extracted; and obtaining a query result corresponding to the query sentence based on the core corpus. The application no longer relies on the accumulation of massive business scenario data to enhance a generalization ability, which ensures accurate and efficient information query, and improves the efficiency and reliability of the information query process. At the same time, it may support information query in different business scenarios, with strong expansion capability and high universality.

Claims (56)

1. A method for information query, comprising:

receiving a query sentence, segmenting the query sentence to obtain word segments, and obtaining a dependency relationship between two word segments and part of speech of the word segments;

obtaining a coding sequence of the query sentence according to the dependency relationship and the part of speech of the word segments;

matching the coding sequence with a generalized template to obtain a core corpus of the query sentence, wherein the generalized template comprises part of speech to be extracted and a dependency relationship to be extracted; and

obtaining a query result corresponding to the query sentence based on the core corpus,

wherein, obtaining the query result corresponding to the query sentence based on the core corpus comprises:

retrieving, based on the core corpus, in a preset corpus database to obtain a target seed corpus corresponding to the core corpus;

determining label data corresponding to the target seed corpus as the query result corresponding to the query sentence,

wherein, retrieving, based on the core corpuses, in the preset corpus database to obtain the target seed corpus corresponding to the core corpus comprises:

retrieving, based on the core corpus, in the preset corpus database to obtain at least one candidate seed corpus corresponding to the core corpus; wherein, the preset corpus database stores a plurality of seed corpuses, and the core corpus is a generalized corpus of at least one seed corpus; and

determining the target seed corpus from the at least one candidate seed corpus.

2. The method of claim 1 , wherein, matching the coding sequence with the generalized template to obtain the core corpus of the query sentence comprises:

extracting coding fragments consistent with the part of speech to be extracted from the coding sequence according to the part of speech to be extracted;

determining a generalized boundary among the coding fragments according to the dependence relationship to be extracted, and extracting the core corpus according to the generalized boundary.

3. The method of claim 1 , wherein, retrieving, based on the core corpus, in the corpus database to obtain the at least one candidate seed corpus corresponding to the core corpus comprises:

retrieving the core corpus from the corpus database based on an inverted index and semantic similarity computation, to obtain at least one candidate seed corpus.

4. The method of claim 1 , wherein, determining the target seed corpus from the at least one candidate seed corpus comprises:

obtaining a similarity between the core corpus and each candidate seed corpus, and selecting a candidate seed corpus with the highest similarity as the target seed corpus.

5. The method of claim 1 , further comprising:

obtaining the number of the at least one candidate seed corpus; and

updating the generalized template in response to the number of the at least one candidate seed corpus greater than a first preset number or less than a second preset number.

6. An apparatus for information query, comprising:

one or more processors;

a memory storing instructions executable by the one or more processors;

wherein the one or more processors are configured to:

receive a query sentence, segment the query sentence to obtain word segments, and obtain a dependency relationship between two word segments and part of speech of the word segments;

obtain a coding sequence of the query sentence according to the dependency relationship and the part of speech of the word segments;

match the coding sequence with a generalized template to obtain a core corpus of the query sentence; wherein, the generalized template comprises part of speech to be extracted and a dependency relationship to be extracted;

obtain a query result corresponding to the query sentence based on the core corpus,

wherein the one or more processors are configured to:

retrieve, based on the core corpus, in a preset corpus database to obtain a target seed corpus corresponding to the core corpus;

determine label data corresponding to the target seed corpus as the query result corresponding to the query sentence,

wherein the one or more processors are configured to:

retrieve, based on the core corpuses, in the preset corpus database to obtain at least one candidate seed corpus corresponding to the core corpus; wherein, the preset corpus database stores a plurality of seed corpuses, and the core corpus is a generalized corpus of at least one seed corpus; and

determine the target seed corpus from the at least one candidate seed corpus.

7. The apparatus of claim 6 , wherein the one or more processors are configured to:

extract coding fragments consistent with the part of speech to be extracted from the coding information according to the part of speech to be extracted;

determine a generalized boundary among the coding fragments according to the dependence relationship to be extracted, and extract the core corpus according to the generalized boundary.

8. The apparatus of claim 6 , wherein the one or more processors are configured to:

retrieve the core corpus from the corpus database based on an inverted index and semantic similarity computation, to obtain at least one candidate seed corpus.

9. The apparatus of claim 6 , wherein the one or more processors are configured to:

obtain a similarity between the core corpus and each candidate seed corpus, and select a candidate seed corpus with the highest similarity as the target seed corpus.

10. The apparatus of claim 6 , wherein the one or more processors are configured to:

obtain the number of the at least one candidate seed corpus; and

update the generalized template in response to the number of the at least one candidate seed corpus greater than a first preset number or less than a second preset number.

11. A non-transitory computer-readable storage medium storing computer instructions, wherein when the computer instructions are executed by a computer, the computer is caused to perform a method for information query, and the method comprises:

receiving a query sentence, segmenting the query sentence to obtain word segments, and obtaining a dependency relationship between two word segments and part of speech of the word segments;

obtaining a coding sequence of the query sentence according to the dependency relationship and the part of speech of the word segments;

matching the coding sequence with a generalized template to obtain a core corpus of the query sentence, wherein the generalized template comprises part of speech to be extracted and a dependency relationship to be extracted; and

obtaining a query result corresponding to the query sentence based on the core corpus,

wherein, obtaining the query result corresponding to the query sentence based on the core corpus comprises:

retrieving, based on the core corpus, in a preset corpus database to obtain a target seed corpus corresponding to the core corpus;

determining label data corresponding to the target seed corpus as the query result corresponding to the query sentence,

wherein, retrieving, based on the core corpuses, in the preset corpus database to obtain the target seed corpus corresponding to the core corpus comprises:

retrieving, based on the core corpus, in the preset corpus database to obtain at least one candidate seed corpus corresponding to the core corpus; wherein, the preset corpus database stores a plurality of seed corpuses, and the core corpus is a generalized corpus of at least one seed corpus; and

determining the target seed corpus from the at least one candidate seed corpus.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2022
From: CHEN, WANSHUN
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061990/0247 →
Priority Claims (1)
CN 202011538686.1 · Dec 23, 2020 · national
Continuity (1)
Related Publication 20210342376A1 · Nov 4, 2021
Cited By (1)
US 12,651,593