IP Library Granted Patent US 10,235,358
Granted Patent B2
US 10,235,358 · App. 13/773,269 · Granted Mar 19, 2019

Exploiting structured content for unsupervised natural language semantic parsing

Inventors: Gokhan Tur (Los Altos, CA); Dilek Hakkani-Tur (Los Altos, CA); Larry Heck (Los Altos, CA); Minwoo Jeong (Bellevue, WA); Ye-Yi Wang (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F17/2785
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,235,358
App. No.
13/773,269
Filed
Feb 21, 2013
Granted
Mar 19, 2019
Kind
B2
Art Unit
2677
USPC
704/9
Abstract

Structured web pages are accessed and parsed to obtain implicit annotation for natural language understanding tasks. Search queries that hit these structured web pages are automatically mined for information that is used to semantically annotate the queries. The automatically annotated queries may be used for automatically building statistical unsupervised slot filling models without using a semantic annotation guideline. For example, tags that are located on a structured web page that are associated with the search query may be used to annotate the query. The mined search queries may be filtered to create a set of queries that is in a form of a natural language query and/or remove queries that are difficult to parse. A natural language model may be trained using the resulting mined queries. Some queries may be set aside for testing and the model may be adapted using in-domain sentences that are not annotated. The models may be tested using these implicitly annotated natural-language-like queries in an unsupervised fashion.

Claims (37)

1. A method for natural language semantic parsing, comprising:

accessing structured content including structured web pages;

parsing a semantic structure identified in the structured content to identify entities linked by a relationship, wherein each entity has a respective tag;

mining a plurality of natural language search queries that can access the structured content to identify, from the plurality of natural language search queries, at least one natural language search query that includes at least one of the entities; and

automatically annotating the at least one natural language search query using the respective tag.

2. The method of claim 1 , further comprising building an unsupervised slot filling model using the at least one natural language search query annotated in the automatically annotating.

3. The method of claim 2 , further comprising adapting the unsupervised slot filling model using in-domain unannotated sentences.

4. The method of claim 3 , further comprising testing a performance of the model based on the at least one natural language search query.

5. The method of claim 1 , wherein the structured content is defined by a triple that consists of two entities linked by a relation.

6. The method of claim 1 , further comprising filtering the natural language search queries by removing at least a portion of the natural language search queries that have un-annotated stopwords.

7. The method of claim 1 , wherein the respective tag of each entity is included in at least one of the structured web pages.

8. A computer-readable storage device storing computer-executable instructions that perform a method when executed, the method comprising:

accessing structured content including structured web pages;

parsing the structured content to identify two entities linked by a relationship, wherein each entity has a respective tag;

mining a plurality of natural language search queries that can access the structured content to identify, from the plurality of natural language search queries, at least one natural language search query that includes at least one of the two entities;

automatically annotating the at least one natural language search query to form at least one annotated natural language search query; and

creating an understanding model including slots using the at least one natural language search query annotated in the automatically annotating.

9. The computer-readable storage device of claim 8 , wherein the understanding model is created in an unsupervised manner.

10. The computer-readable storage device of claim 8 , wherein the method further comprises testing a performance of the model based on the at least one natural language search query.

11. The computer-readable storage device of claim 8 , wherein the structured content is defined by a triple that consists of two entities linked by a relation.

12. The computer-readable storage device of claim 8 , further comprising filtering the natural language search queries by removing natural language search queries that have un-annotated non-stopwords.

13. A system for natural language semantic parsing, comprising:

a processor and memory;

an operating environment executing using the processor; and

a knowledge manager that is configured to perform actions comprising:

accessing structured content including structured web pages;

parsing the structured content to identify two entities linked by a relationship,

wherein each of the entities has a respective tag;

mining a plurality of natural language search queries that can access the structured content to identify, from the plurality of natural language search queries, at least one natural language search query that includes at least one of the two entities;

automatically annotating the at least one natural language search query using the respective tags; and

creating an understanding model including slots using the at least one natural language search query annotated in the automatically annotating.

14. The system of claim 13 , wherein the understanding model is created in an unsupervised manner.

15. The system of claim 13 , further comprising testing a performance of the model based on the at least one natural language query.

16. The system of claim 13 , wherein the structured content includes multiple triples that consist of two entities linked by a relation.

17. The system of claim 13 , further comprising filtering the natural language search queries by removing some of the natural language search queries that have an un-annotated non-stopword.

18. The system of claim 13 , wherein the structured web pages include a semantic web.

19. The system of claim 13 , wherein the structured content is defined by a triple that consists of two entities linked by a relation.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 039025/0454 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2013
From: TUR, GOKHAN; HAKKANI-TUR, DILEK; HECK, LARRY; JEONG, MINWOO; WANG, YE-YI
To: MICROSOFT CORPORATION
Reel/Frame 030063/0872 →
Continuity (1)
Related Publication 20140236575A1 · Aug 21, 2014
Cited By (4)
US 12,217,003 US 12,306,874 US 12,341,751 US 12,602,209