IP Library › Granted Patent US 11,481,547
Granted Patent B2
US 11,481,547 · App. 17/142,718 · Granted Oct 25, 2022

Framework for chinese text error identification and correction

Inventors: Tao Yang (Mountain View, CA); Zeyu You (San Jose, CA); Min Tu (Cupertino, CA); Shangqing Zhang (San Jose, CA); Xu Wang (Palo Alto, CA); Lianyi Han (Palo Alto, CA); Wei Fan (New York, NY)
Assignee: TENCENT AMERICA LLC
G06F40/232G06F40/129G06F40/274G06F40/53
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,481,547
App. No.
17/142,718
Granted
Oct 25, 2022
Kind
B2
Abstract

A method, computer program, and computer system is provided for text error identification and correction. A text input having a phonetic component and a glyphic component is received. Information corresponding to the phonetic component and the glyphic component is coded as a fixed-length sequence. One or more candidate replacement words corresponding to the fixed-length sequence are identified. At least a portion of the text input is replaced with a candidate replacement word from among the one or more candidate replacement words.

Claims (265)

1. A method of text error identification and correction, executable by a processor, the method comprising:

receiving a text input having a phonetic component and a glyphic component;

coding a fixed-length sequence comprising a plurality of positions, a first portion of the plurality of positions corresponding to the phonetic component and a second portion of the plurality of positions corresponding to the glyphic component of the received text;

wherein the plurality of positions corresponding to the phonetic component comprises:

first, second, and third positions corresponding to initial, final, and auxiliary final components corresponding to a character associated with the text input;

a fourth position corresponding to a tone component corresponding to the character associated with the text input;

wherein the plurality of positions corresponding to the glyphic component comprises:

fifth, sixth, seventh, eighth, and ninth positions corresponding to a five-digit Four-Corner code corresponding to the character associated with the text input;

a tenth position corresponding to a structure corresponding to the character associated with the text input;

an eleventh position corresponding to a number of strokes corresponding to the character associated with the text input;

identifying a plurality of candidate replacement words according to a calculated similarity between the fixed-length sequence of the received text and a fixed length sequence of each of the plurality of candidate replacement words, where each of the plurality of positions of the fixed-length sequence are individually weighted by a respective weighting function;

ranking the plurality of candidate replacement words, based on the calculated similarity; and

outputting a correction text comprising a candidate replacement word having a highest ranking among the plurality of candidate replacement words.

2. The method of claim 1 , further comprising replacing at least a portion of the text input with the correction text.

3. The method of claim 1 , wherein at least a portion of the plurality of candidate replacement words correspond to a domain-specific application associated with the text input.

4. The method of claim 1 , wherein the text input comprises one or more from among traditional Chinese characters, simplified Chinese characters, and Pinyin input.

5. The method of claim 1 , wherein the calculated similarity is calculated as:

s

⁡

(

C

1

,

C

2

)

=

∑

i

=

0

5

w

gi

⁢

f

⁡

(

g

1

⁢

i

+

g

2

⁢

i

)

+

w

g

⁢

6

⁢

h

⁡

(

g

16

,

g

2

⁢

6

)

+

∑

i

=

0

3

w

pi

⁢

f

⁡

(

P

1

⁢

i

,

P

2

⁢

i

)

where C x denotes the character, s(C 1 ,C 2 ) denotes a similarity between two characters C i and C 2 , g x denotes a set of the glyphic components of the fixed-length sequence, p x denotes a set of the phonetic components of the fixed-length sequence, w denotes a set of tunable weighting variables, f(a,b) denotes a binary measurement function for two codes a and b, and h(a,b) denotes a continuous measurement function for the two codes a and b.

6. A computer system for text error identification and correction, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access said computer program code and operate as instructed by said computer program code, said computer program code including:

receiving code configured to cause the one or more computer processors to receive a text input having a phonetic component and a glyphic component;

coding code configured to cause the one or more computer processors to code a fixed-length sequence comprising a plurality of positions, a first portion of the plurality of positions corresponding to the phonetic component and a second portion of the plurality of positions corresponding to the glyphic component of the received text;

wherein the plurality of positions corresponding to the phonetic component comprises:

first, second, and third positions corresponding to initial, final, and auxiliary final components corresponding to a character associated with the text input;

a fourth position corresponding to a tone component corresponding to the character associated with the text input;

wherein the plurality of positions corresponding to the glyphic component comprises:

fifth, sixth, seventh, eighth, and ninth positions corresponding to a five-digit Four-Corner code corresponding to the character associated with the text input;

a tenth position corresponding to a structure corresponding to the character associated with the text input;

an eleventh position corresponding to a number of strokes corresponding to the character associated with the text input;

identifying code configured to cause the one or more computer processors to identify a plurality of candidate replacement words according to a calculated similarity between the fixed-length sequence of the received text and a fixed length sequence of each of the plurality of candidate replacement words, where each of the plurality of positions of the fixed-length sequence are individually weighted by a respective weighting function;

ranking code configured to cause the one or more computer processors to rank the plurality of candidate replacement words, based on the calculated similarity; and

outputting code configured to cause the one or more computer processors to output a correction text comprising a candidate replacement word having a highest ranking among the plurality of candidate replacement words.

7. The computer system of claim 6 , further comprising replacing code configured to cause the one or more computer processors to replace at least a portion of the text input with the correction text.

8. The computer system of claim 6 , wherein at least a portion of the plurality of candidate replacement words correspond to a domain-specific application associated with the text input.

9. The computer system of claim 6 , wherein the text input comprises one or more from among traditional Chinese characters, simplified Chinese characters, and Pinyin input.

10. The computer system of claim 6 , wherein the calculated similarity is calculated as:

s

⁡

(

C

1

,

C

2

)

=

∑

i

=

0

5

w

gi

⁢

f

⁡

(

g

1

⁢

i

+

g

2

⁢

i

)

+

w

g

⁢

6

⁢

h

⁡

(

g

16

,

g

2

⁢

6

)

+

∑

i

=

0

3

w

pi

⁢

f

⁡

(

P

1

⁢

i

,

P

2

⁢

i

)

where C x denotes the character, s(C 1 ,C 2 ) denotes a similarity between two characters C i and C 2 , g x denotes a set of the glyphic components of the fixed-length sequence, p x denotes a set of the phonetic components of the fixed-length sequence, w denotes a set of tunable weighting variables, f(a,b) denotes a binary measurement function for two codes a and b, and h(a,b) denotes a continuous measurement function for the two codes a and b.

11. A non-transitory computer readable medium having stored thereon a computer program for text error identification and correction, the computer program configured to cause one or more computer processors to:

receive a text input having a phonetic component and a glyphic component;

code a fixed-length sequence comprising a plurality of positions, a first portion of the plurality of positions corresponding to the phonetic component and a second portion of the plurality of positions corresponding to the glyphic component of the received text;

wherein the plurality of positions corresponding to the phonetic component comprises:

first, second, and third positions corresponding to initial, final, and auxiliary final components corresponding to a character associated with the text input;

a fourth position corresponding to a tone component corresponding to the character associated with the text input;

wherein the plurality of positions corresponding to the glyphic component comprises:

fifth, sixth, seventh, eighth, and ninth positions corresponding to a five-digit Four-Corner code corresponding to the character associated with the text input;

a tenth position corresponding to a structure corresponding to the character associated with the text input;

an eleventh position corresponding to a number of strokes corresponding to the character associated with the text input;

identify a plurality of candidate replacement words according to a calculated similarity between the fixed-length sequence of the received text and a fixed length sequence of each of the plurality of candidate replacement words, where each of the plurality of positions of the fixed-length sequence are individually weighted by a respective weighting function;

rank the plurality of candidate replacement words, based on the calculated similarity; and

output a correction text comprising a candidate replacement word having a highest ranking among the plurality of candidate replacement words.

12. The non-transitory computer readable medium of claim 11 , wherein the computer program is further configured to cause the one or more computer processors to replace at least a portion of the text input with the correction text.

13. The non-transitory computer readable medium of claim 11 , wherein the plurality of candidate replacement words correspond to a domain-specific application associated with the text input.

14. The non-transitory computer readable medium of claim 11 , wherein the calculated similarity is calculated as:

s

⁡

(

C

1

,

C

2

)

=

∑

i

=

0

5

w

gi

⁢

f

⁡

(

g

1

⁢

i

+

g

2

⁢

i

)

+

w

g

⁢

6

⁢

h

⁡

(

g

16

,

g

2

⁢

6

)

+

∑

i

=

0

3

w

pi

⁢

f

⁡

(

P

1

⁢

i

,

P

2

⁢

i

)

where C x denotes the character, s(C 1 ,C 2 ) denotes a similarity between two characters C i and C 2 , g x denotes a set of the glyphic components of the fixed-length sequence, p x denotes a set of the phonetic components of the fixed-length sequence, w denotes a set of tunable weighting variables, f(a,b) denotes a binary measurement function for two codes a and b, and h(a,b) denotes a continuous measurement function for the two codes a and b.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2021
From: YANG, TAO; YOU, ZEYU; TU, MIN; ZHANG, SHANGQING; WANG, XU; HAN, LIANYI; FAN, WEI
To: TENCENT AMERICA LLC
Reel/Frame 054830/0878 →
Continuity (1)
Related Publication 20220215170A1 · Jul 7, 2022