IP Library › Granted Patent US 11,630,957
Granted Patent B2
US 11,630,957 · App. 16/807,997 · Granted Apr 18, 2023

Natural language processing method and apparatus

Inventors: Yasheng Wang (Shenzhen, CN); Jiansheng Wei (Shenzhen, CN); Yang Zhang (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F40/30G06F40/242G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,630,957
App. No.
16/807,997
Granted
Apr 18, 2023
Kind
B2
Abstract

A natural language processing method includes obtaining a to-be-processed phrase, where the to-be-processed phrase includes M words, determining polarity characteristic information of m to-be-processed words in the M words, where polarity characteristic information of an i th word in the m to-be-processed words includes n polarity characteristic values, and each polarity characteristic value corresponds to one sentiment polarity, determining a polarity characteristic vector of the to-be-processed phrase based on the polarity characteristic information of the m to-be-processed words, where the polarity characteristic vector includes n groups of components in a one-to-one correspondence with n sentiment polarities, and determining a sentiment polarity of the to-be-processed phrase based on the polarity characteristic vector of the to-be-processed phrase using a preset classifier, and outputting the sentiment polarity.

Claims (75)

1. A natural language processing method, comprising:

obtaining a single to-be-processed phrase that comprises M words;

determining first respective polarity characteristic information for each respective word of first to-be-processed words in the M words, wherein, for each respective word of the first to-be-processed words, the first respective polarity characteristic information comprises a first respective plurality of n polarity characteristic values, wherein, for each respective word of the first to-be-processed words, each respective polarity characteristic value of the first respective plurality of n polarity characteristic values corresponds to a respective sentiment polarity of n sentiment polarities and is determined at least in part by:

determining, from all of one or more second phrases that are in a preset dictionary and that comprise the respective word, a first quantity of one or more first phrases that correspond to the respective sentiment polarity; and

determining, using the first quantity, a first respective percentage of the one or more second phrases that correspond to the one or more first phrases;

determining, based on the first respective polarity characteristic information of the first to-be-processed words, a first polarity characteristic vector of the single to-be-processed phrase, wherein the first polarity characteristic vector comprises first n respective groups, wherein each respective group of the first n respective groups comprises a first respective plurality of m components, wherein the first n respective groups are in a one-to-one correspondence with the n sentiment polarities, and wherein each respective group of the first n respective groups corresponds to a respective sentiment polarity of the n sentiment polarities and is determined based on a first respective subset of polarity characteristic values of the first respective plurality of n polarity characteristic values of the first to-be-processed words that correspond to the respective sentiment polarity;

determining, using a preset classifier and based on the first polarity characteristic vector, a first sentiment polarity of the single to-be-processed phrase; and

outputting the first sentiment polarity of the single to-be-processed phrase, wherein M, m, and n are positive integers, and wherein M≥m.

2. The natural language processing method of claim 1 , wherein when m>1, determining the first polarity characteristic vector comprises combining the first respective plurality of n polarity characteristic values in the first respective polarity characteristic information of the first to-be-processed words into the first polarity characteristic vector, and wherein a first group of components in the n groups comprises m polarity characteristic values that are obtained by combining first polarity characteristic values of the first respective plurality of n polarity characteristic values.

3. The natural language processing method of claim 2 , wherein when M>m, a greatest value of each of the first respective plurality of n polarity characteristic values is greater than any polarity characteristic value of any one of remaining (M−m) words.

4. The natural language processing method of claim 2 , wherein combining the first respective plurality of n polarity characteristic values comprises combining the first respective plurality of n polarity characteristic values into the first polarity characteristic vector according to an arrangement order of the first to-be-processed words.

5. The natural language processing method of claim 1 , wherein when m>1, determining the first polarity characteristic vector comprises:

traversing the range [1,n] for x, wherein x is a positive integer; and

determining an x th group of components in the first n respective groups in any one of the following manners:

finding an average of x th polarity characteristic values of all of the first to-be-processed words;

finding a sum of the x th polarity characteristic values; or

finding a greatest value of the x th polarity characteristic values.

6. The natural language processing method of claim 1 , further comprising:

treating the single to-be-processed phrase as a first processed phrase; and

adding the first processed phrase into the preset dictionary.

7. The natural language processing method of claim 1 , further comprising:

obtaining a training sample from the preset dictionary, wherein the training sample comprises Y phrases with the n sentiment polarities, and wherein each phrase of the Y phrases comprises second to-be-processed words;

training the preset classifier using the training sample;

determining second respective polarity characteristic information of each word of third to-be-processed words, of the second to-be-processed words, comprised in a second phrase of the Y phrases, wherein corresponding polarity characteristic information of respective words in the second to-be-processed words comprises a second plurality of n polarity characteristic values, wherein, for each respective word in the second to-be-processed words, the second plurality of n polarity characteristic values corresponds to a respective sentiment priority of the n sentiment priorities and is determined at least in part by:

determining, from all of one or more fourth phrases that are in the preset dictionary and that comprise the respective word, a second quantity of one or more third phrases that correspond to the respective sentiment polarity; and

determining, using the second quantity, a second respective percentage of the one or more fourth phrases that correspond to the one or more third phrases;

determining a second polarity characteristic vector of the second phrase based on the second respective polarity characteristic information of the third to-be-processed words, wherein the second polarity characteristic vector comprises second n groups of a second respective plurality of m components in a one-to-one correspondence with the n sentiment polarities, and wherein each respective group of the second n respective groups corresponds to a respective sentiment polarity of the n sentiment polarities and is determined based on a second respective subset of polarity characteristic values of the second respective plurality of n polarity characteristic values of the second to-be-processed words that correspond to the respective sentiment polarity; and

training the preset classifier using a second sentiment polarity of the second phrase and the second polarity characteristic vector, wherein Y is a positive integer.

8. The natural language processing method of claim 7 , further comprising:

treating the second to-be-processed phrase as a second processed phrase, and

training the preset classifier with the second processed phrase as the training sample.

9. A natural language processing apparatus, comprising:

a processor; and

a memory coupled to the processor and storing instructions that, when executed by the processor, cause the natural language processing apparatus to be configured to:

obtain a single to-be-processed phrase that comprises M words;

determine first respective polarity characteristic information for each respective word of first to-be-processed words in the M words, wherein, for each respective word of the first to-be-processed words, the first respective polarity characteristic information comprises a first respective plurality of n polarity characteristic values, wherein, for each respective word of the first to-be-processed words, each respective polarity characteristic value of the first respective plurality of n polarity characteristic values corresponds to a respective sentiment polarity of n sentiment polarities and is determined at least in part by:

determining, from all of one or more second phrases that are in a preset dictionary and that comprise the respective word, a first quantity of one or more first phrases that correspond to the respective sentiment polarity; and

determining, using the first quantity, a first respective percentage of the one or more second phrases that correspond to the one or more first phrases;

determine, based on the first respective polarity characteristic information of the first to-be-processed words, a first polarity characteristic vector of the single to-be-processed phrase, wherein the first polarity characteristic vector comprises first n respective groups, wherein each respective group of the first n respective groups comprises a first respective plurality of m components, wherein the first n respective groups are in a one-to-one correspondence with the n sentiment polarities, and wherein each respective group of the first n respective groups corresponds to a respective sentiment polarity of the n sentiment polarities and is determined based on a first respective subset of polarity characteristic values of the first respective plurality of n polarity characteristic values of the first to-be-processed words that correspond to the respective sentiment polarity;

determine, using a preset classifier and based on the first polarity characteristic vector, a first sentiment polarity of the single to-be-processed phrase; and

output the first sentiment polarity of the single to-be-processed phrase, wherein M, m, and n are positive integers.

10. The natural language processing apparatus of claim 9 , wherein when m>1, the instructions further cause the natural language processing apparatus to be configured to combine the first respective plurality of n polarity characteristic values in the first respective polarity characteristic information of the first to-be-processed words into the first polarity characteristic vector, and wherein a first group of components in the n groups comprises m polarity characteristic values that are obtained by combining first polarity characteristic values of the first respective plurality of n polarity characteristic values.

11. The natural language processing apparatus of claim 10 , wherein when M>m, a greatest value of each of the first respective plurality of n polarity characteristic values is greater than any polarity characteristic value of any one of remaining (M−m) words.

12. The natural language processing apparatus of claim 10 , wherein the instructions further cause the natural language processing apparatus to be configured to combine the first respective plurality of n polarity characteristic values into the first polarity characteristic vector according to an arrangement order of the first to-be-processed words.

13. The natural language processing apparatus of claim 9 , wherein when M>1, the instructions further cause the natural language processing apparatus to be configured to:

traverse the range [1,n] for x, wherein x is a positive integer; and

determine an x th group of components in the first n respective groups of components by:

finding an average of x th polarity characteristic values of all of the first to-be-processed words;

finding a sum of the x th polarity characteristic values; or

finding a greatest value of the x th polarity characteristic values.

14. The natural language processing apparatus of claim 9 , wherein the instructions further cause the natural language processing apparatus to be configured to:

treat the first to-be-processed phrase as a first processed phrase; and

add the first processed phrase into the preset dictionary.

15. The natural language processing apparatus of claim 9 , wherein the instructions further cause the natural language processing apparatus to be configured to:

obtain a training sample from the preset dictionary, wherein the training sample comprises Y phrases with the n sentiment polarities, and wherein each phrase of the Y phrases comprises second to-be-processed words;

train the preset classifier using the training sample;

determine second respective polarity characteristic information of each word of third to-be-processed words, of the second to-be-processed words, comprised in a second phrase of the Y phrases, wherein corresponding polarity characteristic information of respective words in the second to-be-processed words comprises a second plurality of n polarity characteristic values, wherein, for each respective word in the second to-be-processed words, the second plurality of n polarity characteristic values corresponds to a respective sentiment priority of the n sentiment priorities and is determined at least in part by:

determine from all of one or more fourth phrases that are in the preset dictionary and that comprise the respective word, a second quantity of one or more third phrases that correspond to the respective sentiment polarity; and

determine, using the second quantity, a second respective percentage of the one or more fourth phrases that correspond to the one or more third phrases;

determine a second polarity characteristic vector of the second phrase based on the second respective polarity characteristic information of the third to-be-processed words, wherein the second polarity characteristic vector comprises second n groups of a second respective plurality of m components in a one-to-one correspondence with the n sentiment polarities, and wherein each respective group of the second n respective groups corresponds to a respective sentiment polarity of the n sentiment polarities and is determined based on a second respective subset of polarity characteristic values of the second respective plurality of n polarity characteristic values of the second to-be-processed words that correspond to the respective sentiment polarity; and

train the preset classifier using a second sentiment polarity of the second phrase and the second polarity characteristic vector, wherein Y is a positive integer.

16. The natural language processing apparatus of claim 15 , wherein the instructions further cause the natural language processing apparatus to be configured to:

treat the second to-be-processed phrase as a second processed phrase; and

train the preset classifier using the second processed phrase as the training sample.

17. A computer program product comprising computer-executable instructions that are stored on a non-transitory computer-readable medium and that, when executed by a processor, cause a computer to:

obtain a single to-be-processed phrase that comprises M words;

determine first respective polarity characteristic information for each respective word of first to-be-processed words in the M words, wherein, for each respective word of the first to-be-processed words, the first respective polarity characteristic information comprises a first respective plurality of n polarity characteristic values, wherein, for each respective word of the first to-be-processed words, each first respective polarity characteristic value of the first respective plurality of n polarity characteristic values corresponds to a first respective sentiment polarity of n sentiment polarities and is determined at least in part by:

determining, from all of one or more second phrases that are in a preset dictionary and that comprise the respective word, a first quantity of one or more first phrases that correspond to the respective sentiment polarity; and

determining, using the first quantity, a respective percentage of the one or more second phrases that correspond to the one or more first phrases;

determine, based on the first respective polarity characteristic information of the first to-be-processed words, a first polarity characteristic vector of the single to-be-processed phrase, wherein the first polarity characteristic vector comprises first n respective groups, wherein each respective group of the first n respective groups comprises a first respective plurality of m components, wherein the first n respective groups are in a one-to-one correspondence with the n sentiment polarities, and wherein each respective group of the first n respective groups corresponds to a respective sentiment polarity of the first plurality of n sentiment polarities and is determined based on a first respective subset of polarity characteristic values of the first respective plurality of n polarity characteristic values of the first to-be-processed words that correspond to the respective sentiment polarity;

determine, using a preset classifier and based on the first polarity characteristic vector, a first sentiment polarity of the single to-be-processed phrase; and

output the first sentiment polarity of the single to-be-processed phrase, wherein M, m, and n are positive integers.

18. The computer program product of claim 17 , wherein when m>1, the instructions further cause the computer to be configured to combine the first respective plurality of n polarity characteristic values in the first respective polarity characteristic information of the first to-be-processed words into the first polarity characteristic vector, and wherein a first group of components in the n groups comprises m polarity characteristic values that are obtained by combining first polarity characteristic values of the first respective plurality of n polarity characteristic values.

19. The computer program product of claim 18 , wherein when M>m, a greatest value of each of the first respective plurality of n polarity characteristic values is greater than any polarity characteristic value of any one of remaining (M−m) words.

20. The computer program product of claim 18 , wherein the instructions further cause the computer to combine the first respective plurality of n polarity characteristic values into the first polarity characteristic vector according to an arrangement order of the first to-be-processed words.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2020
From: WANG, YASHENG; WEI, JIANSHENG; ZHANG, YANG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 052141/0975 →
Priority Claims (1)
CN 201710786457.3 · Sep 4, 2017 · national
Continuity (2)
Continuation PCTCN2018103837 · Sep 3, 2018
Related Publication 20200202075A1 · Jun 25, 2020
Cited By (1)
US 12,394,406