IP Library › Granted Patent US 11,429,785
Granted Patent B2
US 11,429,785 · App. 17/327,072 · Granted Aug 30, 2022

System and method for generating test document for context sensitive spelling error correction

Inventors: Hyukchul Kwon (Busan, KR); Jung-Hun Lee (Busan, KR)
Assignee: Pusan National University Industry-University Cooperation Foundation
G06F40/232G06F40/166G06F40/242G06F40/279G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,429,785
App. No.
17/327,072
Filed
May 21, 2021
Granted
Aug 30, 2022
Kind
B2
Art Unit
2144
USPC
715/257
Abstract

Disclosed is a system for generating test documents for context-sensitive spelling error correction. The system includes: an input unit inputting an error-free document for generating an error document; an error target word segment test unit checking possibility of an error in a word segment by sequentially examining word segments of the entire sentences in the document input through the input unit and searching for a candidate word appearing at the corresponding position together with surrounding context; an error word candidate selection unit selecting error word candidates among candidate words found at the corresponding position by considering edit distances to a correct word and keyboard typographical errors; and an error word determination and presentation unit calculating probabilities of an error word candidate and its surrounding context and determining an error word of the highest priority as a final error word.

Claims (130)

1. A system for generating test documents for context-sensitive spelling error correction, comprising:

a processor configured to:

input an error-free document for generating an error document;

check possibility of an error in a word segment by sequentially examining word segments of entire sentences in the document and search for a candidate word appearing at a corresponding position together with surrounding context;

select error word candidates through filtering based on an edit distance between a word found by checking the possibility of the error in the word segment and a correct word and keyboard typographical error categories for characters; and

calculate probabilities of an error word candidate and its surrounding context and determine an error word of a highest priority as a final error word,

recognize a relationship between context probabilities of a correct word and an error candidate word using a noisy channel model and select a candidate word not exceeding the probability of the correct word as an error word by calculating

I

^

⁢

=

argmax

I

⁢

p

⁡

(

I

❘

O

)

=

argmax

I

⁢

p

⁡

(

I

)

⁢

p

⁡

(

O

❘

I

)

p

⁡

(

O

)

=

argmax

I

⁢

p

⁡

(

I

)

⁢

p

⁡

(

O

❘

I

)

,

where probability p(O) of output data is a constant, p(I) represents a language model, p(O|I) represents a channel probability distribution, the language model p(I) is defined as a probability distribution of a character string that a user tries to input, the channel probability p(O|I) is defined as an occurrence rate of typing error, and the approximate value Î is obtained by calculating a probability formed by a candidate word and its context.

2. The system of claim 1 , wherein the processor is further configured to find all words co-occurring in the surrounding context of a key word for generating error words using information on N-grams.

3. The system of claim 1 , wherein the processor is further configured to find candidate words through a pre-built N-gram dictionary using

candidate words=< w i−2 ,w i−1 ,*>∪<w i−1 ,*,w i+1 >∪<*,w i+1 ,w i+2 >,*=w i

and search N-grams spanning word segments in both sides (adjacent word segments: w i−2 , w i−1 , w i+1 , w i+2 ) around the position “*” of a key word (*=w i ), wherein the search is conducted to find all statistical candidate words appearing simultaneously with context words adjacent to the key word position “*”.

4. The system of claim 1 , wherein the processor is further configured to select an error candidate word using an error word filter utilizing a set of all candidate words co-occurring in the surrounding context.

5. The system of claim 4 , wherein the error word filter operates based on a distance between adjacent keys corresponding to a keyboard input, alphabet letters, and an edit distance between a key word and a candidate word.

6. The system of claim 1 , wherein the keyboard typographical error categories include omission of a letter, addition of a letter, repeated typing of a letter, omission of a repeated letter, typing of a wrong letter, interchange of two adjacent letters, and a combination of all preceding input errors, and

an error input is determined among neighboring keyboard letters adjacent to a target input letter on the keyboard.

7. A method for generating test documents for context-sensitive spelling error correction, comprising:

inputting, by a processor, an error-free document for generating an error document and sequentially examining word segments of entire sentences in the document;

determining, by the processor, whether an error candidate word appears in a word segment and determining a corresponding word segment as an error generating word segment when the error candidate word exists in the word segment;

filtering, by the processor, candidate words into error candidate words and selecting an error candidate word by considering a keyboard input process and an edit distance;

calculating, by the processor, probabilities of filtered error candidate words and surrounding context of a key word segment and reflecting an error in the document using an error word having the highest priority through comparison of probabilities or through random selection among error candidate words; and

recognizing, by the processor, a relationship between context probabilities of a correct word and an error candidate word using a noisy channel model and selecting a candidate word not exceeding the probability of the correct word as an error word by calculating

I

^

=

argmax

I

⁢

p

⁡

(

I

|

O

)

=

argmax

I

⁢

p

⁡

(

I

)

⁢

p

⁡

(

O

|

I

)

p

⁡

(

O

)

=

argmax

I

⁢

p

⁡

(

I

)

⁢

p

⁡

(

O

|

I

)

,

where probability p(O) of output data is a constant, p(I) represents a language model, p(O|I) represents a channel probability distribution, the language model p(I) is defined as a probability distribution of a character string that a user tries to input, the channel probability p(O|I) is defined as an occurrence rate of typing error, and the approximate value Î is obtained by calculating a probability formed by a candidate word and its context.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2021
From: KWON, HYUKCHUL; LEE, JUNG-HUN
To: PUSAN NATIONAL UNIVERSITY INDUSTRY-UNIVERSITY COOPERATION FOUNDATION
Reel/Frame 056315/0876 →
Priority Claims (1)
KR 10-2020-0158196 · Nov 23, 2020 · national
Continuity (1)
Related Publication 20220164530A1 · May 26, 2022
Cited By (1)
US 12,210,848