IP Library Granted Patent US 7,890,539
Granted Patent B2
US 7,890,539 · App. 11/974,022 · Granted Feb 15, 2011

Semantic matching using predicate-argument structure

Assignee: Raytheon BBN Technologies Corp.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,890,539
App. No.
11/974,022
Granted
Feb 15, 2011
Kind
B2
Abstract

The invention relates to topic classification systems in which text intervals are represented as proposition trees. Free-text queries and candidate responses are transformed into proposition trees, and a particular candidate response can be matched to a free-text query by transforming the proposition trees of the free-text query into the proposition trees of the candidate responses. Because proposition trees are able to capture semantic information of text intervals, the topic classification system accounts for the relative importance of topic words, for paraphrases and re-wordings, and for omissions and additions. Redundancy of two text intervals can also be identified.

Claims (59)

1. A system that processes text intervals, comprising:

a memory;

a processor configured to execute a plurality of modules stored in the memory;

the modules including:

a preprocessing module configured to:

extract a first proposition from a first text interval;

a generation module configured to:

generate a first proposition tree from the first proposition, wherein the first proposition tree comprises a set of nodes and a set of edges, and wherein each edge includes a semantic relationship between nodes of the first proposition tree; and

a matching module configured to:

determine a first similarity value between the first text interval and a second text interval based on a comparison of the first proposition tree and a second proposition tree corresponding to the second text interval, wherein the second text interval is different from the first text interval,

determine a second similarity value between the second text interval and the first text interval,

find the second text interval redundant to the first text interval in response to the first and second similarity values each exceeding respective thresholds, and

selectively output the second text interval based on at least one of the first similarity value and the redundancy finding.

2. The system of claim 1 , wherein the first text interval is a query and the second text interval is a candidate response and the matching module is further configured to output the second text interval if the first similarity value exceeds a threshold.

3. The system of claim 1 , wherein the matching module is further configured to:

refrain from outputting redundant text intervals.

4. The system of claim 1 , wherein determining the first similarity value and the second similarity value comprises performing a two-way comparison between the first and second intervals.

5. The system of claim 1 , further comprising:

an augmentation module configured to, for at least one node in the first proposition tree, associate, with the at least one node, a word having a relationship to the at least one node to form a first augmented proposition tree.

6. The system of claim 5 , wherein the relationship is a co-reference relationship.

7. The system of claim 6 , wherein associating the related word based on a co-reference relationship comprises:

identifying a real-world object, concept, or event included in the at least one node,

identifying alternative words in a document in which the first text interval is included that correspond to the same real-world object, concept, or event, and

augmenting the at least one node in the proposition tree with the identified alternative words to create the first augmented proposition tree.

8. The system of claim 5 , wherein the relationship is a synonym, hypernym, hyponym, or substitutable label relationship.

9. The system of claim 5 , wherein the matching module is configured to determine the first similarity value by determining a number of nodes and a number of edges that match between the first augmented proposition tree and the second proposition tree.

10. The system of claim 9 , wherein a first node in the first augmented proposition tree matches a second node in the second proposition tree in response to at least one word in, or associated with, the first node matching a word in, or associated with, the second node.

11. The system of claim 9 , wherein in response to the first node matching the second node as a result of a word associated with the first node matching the second node, decreasing the first similarity score based on the relationship between the associated word and the first node.

12. The system of claim 9 , wherein a first edge in the first augmented proposition tree matches a second edge in the second proposition tree in response to a semantic relationship associated with the first edge being substitutable for the semantic relationship associated with the second edge.

13. The system of claim 9 , wherein in response to the first edge matching the second edge as a result based on a substitute semantic relationship, decreasing the first similarity score based on the substitution.

14. The system of claim 5 , wherein the matching module is configured to determine the first similarity value by calculating a transformation score based on costs associated with transforming the first augmented proposition tree to the second proposition tree.

15. The system of claim 5 , wherein the second proposition tree comprises an augmented proposition tree.

16. The system of claim 1 , wherein the generation system is configured to:

generate a first plurality of proposition subtrees and a second plurality of proposition subtrees; and

the matching system is configured to:

determine the second similarity value by matching the first plurality of proposition subtrees to the second plurality of proposition subtrees.

17. The system of claim 1 , wherein the generation system is configured to:

generate a first bag of nodes from the first proposition tree and a second bag of nodes from the second proposition tree; and

the matching system is configured to:

determine the second similarity value by matching the first bag of nodes to the second bag of nodes.

18. A method of processing text intervals, comprising:

extracting a first proposition from a first text interval;

generating a first proposition tree from the first proposition, wherein the first proposition tree comprises a set of nodes and a set of edges, and wherein each edge includes a semantic relationship between nodes of the first proposition tree;

determining a first similarity value between the first text interval and a second text interval based on a comparison of the first proposition tree and a second proposition tree corresponding to the second text interval, wherein the second text interval is different from the first text interval;

determining a second similarity value between the second text interval and the first text interval;

finding the second text interval redundant to the first text interval in response to the first and second similarity values each exceeding respective thresholds; and

selectively outputting, using a processor, the second text interval based on at least one of the first similarity value and the redundancy finding.

19. The method of claim 18 , comprising outputting the second text interval if the first similarity value exceeds a threshold, wherein the first text interval is a query and the second text interval is a candidate response.

20. The method of claim 18 , comprising:

refraining from outputting redundant text intervals.

21. The method of claim 18 , comprising associating, for at least one node in the first proposition tree, a word having a relationship to the at least one node to form a first augmented proposition tree.

22. The method of claim 21 , wherein the relationship is a co-reference relationship.

23. The method of claim 21 , wherein the relationship is a synonym, hypernym, hyponym, or substitutable label relationship.

24. The method of claim 18 , comprising:

generating a first plurality of proposition subtrees and a second plurality of proposition subtrees; and

determining the second similarity value by matching the first plurality of proposition subtrees to the second plurality of proposition subtrees.

25. The method of claim 18 , comprising:

generating a first bag of nodes from the first proposition tree and a second bag of nodes from the second proposition tree; and

determining the second similarity value by matching the first bag of nodes to the second bag of nodes.

Assignments (6)
CHANGE OF NAME Recorded Aug 22, 2024
From: RAYTHEON BBN TECHNOLOGIES CORP.
To: RTX BBN TECHNOLOGIES, INC.
Reel/Frame 068748/0419 →
CHANGE OF NAME Recorded May 28, 2010
From: BBN TECHNOLOGIES CORP.
To: RAYTHEON BBN TECHNOLOGIES CORP.
Reel/Frame 024456/0537 →
RELEASE OF SECURITY INTEREST Recorded Oct 27, 2009
From: BANK OF AMERICA, N.A. (SUCCESSOR BY MERGER TO FLEET NATIONAL BANK)
To: BBN TECHNOLOGIES CORP. (AS SUCCESSOR BY MERGER TO BBNT SOLUTIONS LLC)
Reel/Frame 023427/0436 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT Recorded Dec 4, 2008
From: BBN TECHNOLOGIES CORP.
To: BANK OF AMERICA, N.A.
Reel/Frame 021926/0017 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2008
From: BOSCHEE, ELIZABETH MEGAN; LEVIT, MICHAEL; FREEDMAN, MARJORIE RUTH
To: BBN TECHNOLOGIES CORP.
Reel/Frame 021843/0151 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT Recorded Sep 19, 2008
From: BBN TECHNOLOGIES CORP.
To: BANK OF AMERICA, N.A., AS AGENT
Reel/Frame 021701/0085 →
Continuity (1)
Related Publication 20090100053A1 · Apr 16, 2009