IP Library › Granted Patent US 11,256,869
Granted Patent B2
US 11,256,869 · App. 16/553,014 · Granted Feb 22, 2022

Word vector correction method

Inventor: Hwiyeol Jo (Seoul, KR)
Assignee: LG ELECTRONICS INC.
G06F40/30G06F16/3347G06F40/247G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,256,869
App. No.
16/553,014
Granted
Feb 22, 2022
Kind
B2
Abstract

The present disclosure provides a word vector correction method using artificial intelligence technology. A word vector correction method using a word vector with n dimensions includes generating a first (n+1)-dimensional word vector using an average of elements included in a first n-dimensional word vector; generating a second (n+1)-dimensional word vector using an average of elements included in a second n-dimensional word vector; and determining whether a first word corresponding to the first word vector and a second word corresponding to the second word vector are similar to each other on the basis of specified synonym information.

Claims (36)

1. A word vector correction method comprising:

generating a first (n+1)-dimensional word vector corresponding to a first n-dimensional word vector for a first word, wherein first to n th values of the first (n+1)-dimensional word vector are set to be equal to first to n th values of the first n-dimensional word vector and a (n+1)th value of the first (n+1)-dimensional word vector is set to an average of the first to n th values of the first n-dimensional word vector;

generating a second (n+1)-dimensional word vector corresponding to a second n-dimensional word vector for a second word, wherein first to n th values of the second (n+1)-dimensional word vector are set to be equal to first to n th values of the second n-dimensional word vector and a (n+1)th value of the second (n+1)-dimensional word vector is set to an average of the first to n th values of the second n-dimensional word vector;

determining whether the first word and the second word are similar to each other on the basis of specified synonym information;

updating the (n+1)th value of the first (n+1)-dimensional word vector and the (n+1)th value of the second (n+1)-dimensional word vector based on the first word and the second word being determined to be similar to each other; and

correcting the first and second n-dimensional word vectors by applying the (n+1)th values of the first and second (n+1)-dimensional word vectors to each of the n values of the first and second n-dimensional word vectors.

2. The word vector correction method of claim 1 , wherein the updating comprises:

calculating an average of the (n+1)th value of the first (n+1)-dimensional word vector and the (n+1)th value of the second (n+1)-dimensional word vector; and

changing the (n+1)th values of each of the first and second (n+1)-dimensional word vectors to the calculated average when the first word and the second word are determined to be similar to each other.

3. The word vector correction method of claim 2 , wherein when the first word and the second word are determined not to be similar to each other, the (n+1)th values of the first and second (n+1)-dimensional word vectors are not changed.

4. The word vector correction method of claim 1 , wherein,

0 to nth values of the first (n+1)-dimensional word vector are the same as those of the first n-dimensional word vector, and

0 to nth values of the second (n+1)-dimensional word vector are the same as those of the second n-dimensional word vector.

5. The word vector correction method of claim 1 , wherein the correcting comprises applying a linear discriminant analysis (LDA) algorithm to apply the (n+1)th values of the first and second (n+1)-dimensional word vectors to the values of the first and second n-dimensional word vectors, respectively.

6. The word vector correction method of claim 5 , wherein a distance between the first and second n-dimensional word vectors in a vector space after the correction is different from a distance between the first and second n-dimensional word vectors before the correction.

7. The word vector correction method of claim 6 , wherein the distance between the first and second n-dimensional word vectors in the vector space after the correction is less than the distance between the first and second n-dimensional word vectors before the correction when the first word and the second word are determined to be similar to each other.

8. The word vector correction method of claim 7 , wherein the distance between the first and second n-dimensional word vectors in the vector space after the correction is greater than the distance between the first and second n-dimensional word vectors before the correction when the first word and the second word are determined to not be similar to each other.

9. The word vector correction method of claim 1 , wherein the specified synonym information comprises a learning model for distributed word representation.

10. A word vector correction method comprising:

based on a first n-dimensional word vector for a first word, determining a first average value of the n values in the first n-dimensional word vector;

based on a second n-dimensional word vector for a second word, determining a second average value of the n values in the second n-dimensional word vector;

determining whether the first word and the second word are similar to each other; and

correcting the first and second n-dimensional word vectors by applying a first calculated value to the values of the first n-dimensional vector and a second calculated value to the values of the second n-dimensional vector,

wherein based on the first word and the second word being determined to be similar to each other, the first calculated value and the second calculated value both equal an average of the first average value and the second average value, and

wherein based on the first word and the second word being determined not to be similar to each other, the first calculated value is equal to the first average value and the second calculated value is equal to the second average value.

11. The word vector correction method of claim 10 , wherein the correcting comprises using a linear discriminant analysis (LDA) algorithm to apply the first calculated value to the values of the first n-dimensional word vector and to apply the second calculated value to the values of the second n-dimensional word vector.

12. The word vector correction method of claim 11 , wherein the similarity between the first word and the second word is determined using a learning model for distributed word representation.

13. A machine-readable non-transitory medium having stored thereon machine-executable instructions for word vector correction, the instructions comprising:

based on a first n-dimensional word vector for a first word, determining a first average value of the n values in the first n-dimensional word vector;

based on a second n-dimensional word vector for a second word, determining a second average value of the n values in the second n-dimensional word vector;

determining whether the first word and the second word are similar to each other; and

correcting the first and second n-dimensional word vectors by applying a first calculated value to the values of the first n-dimensional vector and a second calculated value to the values of the second n-dimensional vector,

wherein based on the first word and the second word being determined to be similar to each other, the first calculated value and the second calculated value both equal an average of the first average value and the second average value, and

wherein based on the first word and the second word being determined not to be similar to each other, the first calculated value is equal to the first average value and the second calculated value is equal to the second average value.

14. The machine-readable non-transitory medium of claim 13 , wherein the correcting comprises using a linear discriminant analysis (LDA) algorithm to apply the first calculated value to the values of the first n-dimensional word vector and to apply the second calculated value to the values of the second n-dimensional word vector.

15. The machine-readable non-transitory medium of claim 14 , wherein the similarity between the first word and the second word is determined using a learning model for distributed word representation.

Continuity (3)
Continuation PCTKR2019095025 · May 31, 2019
Provisional Application 62728063 · Sep 6, 2018
Related Publication 20190384817A1 · Dec 19, 2019