IP Library Granted Patent US 7,860,707
Granted Patent B2
US 7,860,707 · App. 11/638,071 · Granted Dec 28, 2010

Compound word splitting for directory assistance services

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,860,707
App. No.
11/638,071
Granted
Dec 28, 2010
Kind
B2
Abstract

A computer-implemented method is disclosed for improving the accuracy of a directory assistance system. The method includes constructing a prefix tree based on a collection of alphabetically organized words. The prefix tree is utilized as a basis for generating splitting rules for a compound word included in an index associated with the directory assistance system. A language model check and a pronunciation check are conducted in order to determine which of the generated splitting rules are mostly likely correct. The compound word is split into word components based on the most likely correct rule or rules. The word components are incorporated into a data set associated with the directory assistance system, such as into a recognition grammar and/or the index.

Claims (42)

1. A computer-implemented method of improving the accuracy of an automatic directory assistance system, comprising:

obtaining a compound word;

identifying a rule that specifies a way to split the compound word;

applying the rule so as to split the compound word into a plurality of word components;

utilizing a computer processor that is a component of a computing device to perform a pronunciation check by comparing an obtained pronunciation of the compound word with an obtained pronunciation of the plurality of word components;

determining, based on the pronunciation check, that the obtained pronunciation of the compound word is different than the obtained pronunciation of the plurality of word components; and

as a response to said determination, effectuating an exclusion of the plurality of word components from a data set that is a listing of identification candidates utilized by the directory assistance system as a reference for identifying a user-initiated input.

2. The method of claim 1 , wherein the data set is a speech recognition grammar.

3. The method of claim 1 , wherein the data set is a query processing index.

4. The method of claim 1 , further comprising:

identifying a second rule that specifies a second way to split the compound word;

applying the second rule so as to split the compound word into a second plurality of word components, the first plurality of word components being different than the second plurality of word components;

utilizing the computer processor to perform a second pronunciation check by comparing the obtained pronunciation of the compound word with an obtained pronunciation of the second plurality of word components;

determining, based on the second pronunciation check, that the obtained pronunciation of the compound word is the same as the obtained pronunciation of the second plurality of word components; and

as a response to the determination that the pronunciations are the same, effectuating an inclusion of the second plurality of individual word components within said data set.

5. The method of claim 1 , wherein identifying a rule comprises constructing a prefix tree.

6. The method of claim 5 , wherein constructing a prefix tree comprises constructing a prefix tree based on an alphabetically organized set of words that includes the compound word and the plurality of word components.

7. The method of claim 6 , wherein the alphabetically organized set of words is a collection of dictionary words in the same language as the user-initiated input.

8. The method of claim 6 , wherein the alphabetically organized set of words is a collection of words contained in a speech engine lexicon, the collection of words being in the same language as the user-initiated input.

9. The method of claim 4 , further comprising performing a language model check to determine which of the first and second plurality of word components is more commonly utilized in a language system that is the same language as the user-initiated input.

10. A computer-implemented method of improving the accuracy of an automatic directory assistance system, comprising:

obtaining a compound word;

identifying a rule that specifies a way to split the compound word;

applying the rule so as to split the compound word into a plurality of word components;

utilizing a computer processor that is a component of a computing device to perform a pronunciation check by comparing an obtained pronunciation of the compound word with an obtained pronunciation of the plurality of word components;

determining, based on the pronunciation check, that the obtained pronunciation of the compound word is the same as the obtained pronunciation of the plurality of word components; and

as a response to the determination that the pronunciations are the same, effectuating an inclusion of the plurality of individual word components within a data set that is a listing of identification candidates utilized by the directory assistance system as a reference for identifying a user-initiated input.

11. The method of claim 10 , wherein the data set is a speech recognition grammar.

12. The method of claim 10 , wherein the data set is a query processing index.

13. The method of claim 10 , wherein identifying a rule comprises constructing a prefix tree.

14. The method of claim 13 , wherein constructing a prefix tree comprises constructing a prefix tree based on an alphabetically organized set of words that includes the compound word and the plurality of word components.

15. The method of claim 14 , wherein the alphabetically organized set of words is a collection of dictionary words in the same language as the user-initiated input.

16. A computer-implemented method of improving the accuracy of an automatic directory assistance system, comprising:

obtaining a compound word;

identifying a rule that specifies a way to split the compound word;

applying the rule so as to split the compound word into a plurality of individual word components;

utilizing a computer processor that is a component of a computing device to compare an obtained pronunciation of the compound word with an obtained pronunciation of the plurality of individual word components;

if the comparison indicates that the pronunciations are the same, then accepting the rule and thereby allowing the plurality of individual word components to be included in a data set that is a listing of identification candidates utilized by the directory assistance system as a reference for identifying a user-initiated input; and

if the comparison indicates that the pronunciations are different, rejecting the rule and thereby establishing a prohibition against including the plurality of individual word components in the data set.

17. The method of claim 16 , identifying a rule comprises constructing a prefix tree based on an alphabetically organized set of words that includes the compound word and the plurality of word components.

18. The method of claim 17 , wherein the alphabetically organized set of words is a collection of dictionary words in the same language as the user-initiated input.

19. The method of claim 18 , wherein the alphabetically organized set of words is a collection of words contained in a speech engine lexicon, the collection of words being in the same language as the user-initiated input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034542/0001 →