IP Library Granted Patent US 10,068,008
Granted Patent B2
US 10,068,008 · App. 14/471,886 · Granted Sep 4, 2018

Spelling correction of email queries

Inventors: Raghavendra Uppinakuduru Udupa (Bangalore, IN); Abhijit Narendra Bhole (Gujarat, IN)
Assignee: Microsoft Technologies Licensing, LLC
G06F17/30675G06F17/277G06F17/3064G06Q10/107H04L51/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,068,008
App. No.
14/471,886
Granted
Sep 4, 2018
Kind
B2
Abstract

Techniques and constructs to facilitate spelling correction of email queries can leverage features of email data to obtain candidate corrections particular to the email data being queried. The constructs may enable accurate spelling correction of email queries across languages and domains based on, for example, one or more of a language model such as a bigram language model and/or a normalized token IDF based language model, a translation model such as an edit distance translation model and/or a fuzzy match translation model, content-based features, and/or contextual features. Content-based features can include features associated with the subject line of emails, content including identified phrases, contacts, and/or the number of candidate emails returned. Contextual features can include a time window of subject match and/or contact match, a frequency of emails received from a contact, and/or device characteristics.

Claims (67)

1. A computer-implemented method for automatically correcting words in an email query, the method comprising:

receiving an email query for mining at least one email repository, wherein the email query comprises at least one misspelled word;

determining two or more features for personalized candidate spelling correction, wherein at least one feature comprises contextual feature of emails in at least one email repository, and wherein the context feature comprises one or more of:

subject lines,

frequent contacts,

recent contacts, and

input device types;

generating at least one personalized candidate spelling correction for the email query based on the two or more determined features;

ranking the at least one personalized candidate spelling corrections based at least in part on a weighted score of two or more determined features;

determining at least one personalized spelling correction for the at least one email query from the ranked personalized candidate spelling corrections; and

applying the at least one personalized spelling correction to the email query to produce a second email query with automatically corrected words.

2. A method as claim 1 recites, wherein at least a part of the email repository is associated with an entity from which the email query is received.

3. A method as claim 1 recites, wherein ranking the personalized candidate spelling corrections is performed based on at least one statistical ranking.

4. A method as claim 1 recites, wherein the contextual feature of emails in the at least one email repository further comprises at least one of a time range of recent emails having a subject and contact field associated with the determined features.

5. A method as claim 1 recites, wherein the at least one personalized candidate spelling correction comprises a misspelling of a known word based at least in part on ranking associated with the personalized candidate spelling corrections.

6. A method as claim 1 recites, further comprising:

matching at least one of the personalized candidate spelling corrections to data from the email repository; and

removing one or more personalized candidate spelling corrections from the ranking personalized candidate spelling corrections based on the matched data.

7. A method as claim 1 recites, further comprising identifying at least one email from the email repository, wherein the at least one email contains at least one of the personalized candidate spelling corrections without matching words in the email query.

8. A method as claim 1 recites, further comprising identifying at least one email from the email repository, wherein the at least one email contains at least one of the personalized candidate spelling corrections with matching words in the email query.

9. A method as claim 1 recites, wherein the email repository includes at least two subdivisions associated with an entity from which a same email query is received; and further comprising:

determine at least a first email from the email repository associated with a first subdivision, wherein the first email contains at least a first match of words with at least one of the personalized candidate spelling corrections, and wherein the first match of words comprise at least one of the misspelled words failing to match emails based on the email query;

determine at least a second email from the email repository associated with a second subdivision, wherein the second email contains at least a second match of words with at least one of the personalized candidate spelling corrections, and wherein the first match of words comprise at least one of the misspelled words failing to match emails based on the email query; and

wherein the first match differs from the second match.

10. A method as claim 1 recites, wherein the email repository includes at least one subdivision associated with an entity from which an email query is received; and further comprising:

identifying at least a first email from the email repository associated with a first subdivision, the first email containing at least a first match to at least one of the personalized candidate spelling corrections matching the email query;

identifying at least a second email from the email repository associated with a second subdivision, the second email containing at least a second match to at least one of the personalized candidate spelling corrections matching the email query; and the first match differing from the second match.

11. A method as claim 1 recites, wherein the email repository is associated with at least two entities from which a same email query is received; and further comprising:

identifying at least a first email from the email repository associated with a first of the at least two entities, the first email containing at least a first match to at least one of the personalized candidate spelling corrections without matching the email query;

identifying at least a second email from the email repository associated with a second of the at least two entities, the second email containing at least a second match to at least one of the personalized candidate spelling corrections without matching the email query; and the first match differing from the second match.

12. A method as claim 1 recites, wherein the email repository is associated with at least two entities from which a same email query is received; and further comprising:

identifying at least a first email from the email repository associated with a first of the at least two entities, the first email containing at least a first match to at least one of the personalized candidate spelling corrections matching the email query;

identifying at least a second email from the email repository associated with a second of the at least two entities, the second email containing at least a second match to at least one of the personalized candidate spelling corrections matching the email query; and the first match differing from the second match.

13. A computing device, comprising:

at least one processing unit; and

at least one memory storing computer executable instructions for correcting a plurality of misspelled words in emails, the instructions when executed by the at least processing unit causing the computing device to:

identify, by a feature function module, features from email data in the email repository, wherein the feature function module comprises context-based features module configured to score personalized candidate spelling corrections based at least on where personalized candidate spelling corrections occur in context of the email data, and wherein the context-based features module is configured to treat output of a translation model module as a feature;

generate, by a candidate generation module, the personalized candidate spelling corrections corresponding to a first email query, wherein the personalized candidate spelling corrections being based at least in part on the email data; and

rank, by a ranking module, the personalized candidate spelling corrections based at least in part on a weighted score of the features from the email data;

applying the ranked personalized candidate spelling corrections to the email query to produce a second email query with automatically corrected words.

14. A computing device as claim 13 recites, wherein the context-based features module is further configured to treat output of a language model module as a feature.

15. A computing device as claim 13 recites, further comprising a content-based features module, wherein the content-based features module is configured to score content-based features based on at least one of:

subject cover, wherein the subject cover comprises at least a part of email data containing at least one candidate token of the personalized candidate spelling corrections in the subject line;

subject match, wherein the subject match comprises at least one candidate token associated with the personalized candidate spelling corrections appearing in at least one subject line in the email data;

key phrase cover, wherein the key phrase cover comprises a score based on a number of candidate tokens associated with the personalized candidate spelling corrections contained in phrases in the email data;

contact match, wherein the contact match comprises at least one candidate token associated with at least one contact of at least one email account in the email data; and

excessive emails indicator, wherein the excessive emails indicator indicates when a number of retrieved emails based on executing the second email query using the ranked personalized candidates corrections is greater than a threshold.

16. A computing device as claim 13 recites, wherein the feature function module further comprises at least one of:

the translation model module configured to score lexical similarity of personalized candidate spelling corrections; and

a content features module configured to score personalized candidate spelling corrections according to whether respective of the personalized candidate spelling corrections occur in content of the email data.

17. A computing device as claim 16 recites, wherein the translation model module is configured to employ at least one translation model of edit distance or fuzzy match.

18. A computing device as claim 16 recites, wherein the context-based features module is configured to score context-based features according to at least one of:

recent subject match, wherein the recent subject match comprises at least a part of candidate tokens appearing in a subject line in the email data in the email repository for a time window of recent emails;

recent contact match, wherein the recent contact match comprises at least a part of candidate tokens appearing in a contact in the email data in the email repository for a time window of recent contact interactions;

frequent contact match, wherein the frequent contact match comprises at least a part of candidate tokens that are present more than a certain number of times in the email data in the email repository for a time window of recent contact interactions, and wherein the frequent contact match further comprises at least a part of candidate tokens appearing in the email data in the email repository more frequently than a threshold frequency for a time window of the email data; and

device match, wherein the device match comprises at least a part candidate tokens present with an indication of input from a particular input type.

19. A computer storage medium having thereon computer executable instructions, the computer-executable instructions upon execution configuring a computer to perform operations comprising:

extracting a plurality of tokens in an email query;

determine at least one feature associated with respective of the tokens, wherein the feature comprises at least one context of email;

generating a vector of the at least one feature associated with the tokens;

generating one or more personalized candidate spelling corrections associated with at least one of the tokens;

ranking the one or more personalized candidate spelling corrections based at least in part on the vector of the feature associated with the tokens;

selecting at least one word from the personalized candidate spelling corrections based on a number of emails associated with the personalized candidate spelling correction;

replacing at least one word in the email query with the selected at least one word from the personalized candidate spelling corrections;

executing the email query; and

retrieving at least one email from an email repository based on a result of the executed email query.

20. A computer storage medium as claim 19 recites, wherein the contextual aspects include at least one of an indication of input from a particular input type, an indication of input from a device type, and/or an indication of input from a particular device.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 9, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 039025/0454 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 28, 2014
From: UDUPA, RAGHAVENDRA UPPINAKUDURU; BHOLE, ABHIJIT NARENDRA
To: MICROSOFT CORPORATION
Reel/Frame 033633/0321 →
Continuity (1)
Related Publication 20160063094A1 · Mar 3, 2016
Cited By (3)
US 12,282,941 US 12,412,039 US 12,664,568