IP Library › Granted Patent US 11,010,561
Granted Patent B2
US 11,010,561 · App. 16/234,080 · Granted May 18, 2021

Sentiment prediction from textual data

Inventor: Jerome R. Bellegarda (Saratoga, CA)
Assignee: Apple Inc.
G06F40/30G06F40/10G06N3/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,010,561
App. No.
16/234,080
Filed
Dec 27, 2018
Granted
May 18, 2021
Kind
B2
Art Unit
2659
USPC
704/9
Abstract

Techniques for predicting sentiment from textual data are described herein. In some examples, the described techniques utilize a sentiment prediction model having bidirectional long short-term memory (LSTM) networks with one or more convolution-and-pooling stages. The bidirectional LSTM networks process vector representations of words in a textual word sequence to determine forward and backward word-level context feature vectors. Forward and backward phrase-level feature vectors are determined based on the forward and backward word-level context feature vectors. The one or more convolution-and-pooling stages pool the forward and backward phrase-level feature vectors to determine pooled phrase-level feature vectors. A sentiment representing the textual word sequence is determined based on the pooled phrase-level feature vectors.

Claims (177)

1. A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for:

receiving a word sequence in textual form, wherein a first word, a second word, and a third word are consecutively arranged in the word sequence;

determining a plurality of word-level context feature vectors for the word sequence;

determining, based on the plurality of word-level context feature vectors, a first phrase-level feature vector for a first phrase in the word sequence and a second phrase-level feature vector for a second phrase in the word sequence;

determining a pooled phrase-level feature vector by pooling the first phrase-level feature vector and the second phrase-level feature vector;

determining, based on the pooled phrase-level feature vector, a sentiment conveyed by the word sequence, wherein the sentiment is selected from a plurality of predefined sentiment states based on the pooled phrase-level feature vector; and

outputting the result, wherein the result is generated by performing an action in accordance with the determined sentiment.

2. The computer-readable storage medium of claim 1 , wherein the first phrase and the second phrase are adjacent phrases in the word sequence.

3. The computer-readable storage medium of claim 2 , wherein the one or more programs further include instructions for:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence and a fourth phrase-level feature vector for a fourth phrase in the word sequence, wherein the third phrase and the fourth phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the third phrase-level feature vector and the fourth phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

4. The computer-readable storage medium of claim 2 , wherein the one or more programs further include instructions for:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence, wherein the second phrase and the third phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the second phrase-level feature vector and the third phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

5. The computer-readable storage medium of claim 3 , wherein the one or more programs further include instructions for:

determining a plurality of pooled phrase-level feature vectors that includes the first pooled phrase-level feature vector for a first portion of the word sequence, the second pooled phrase-level feature vector for a second portion of the word sequence, and a third pooled phrase-level feature vector for a third portion of the word sequence, wherein the first, second, and third portions of the word sequence are consecutive portions of the word sequence;

determining, based on the plurality of pooled phrase-level feature vectors, a plurality of phrase-level context feature vectors that includes a first phrase-level context feature vector for the first portion of the word sequence and a second phrase-level context feature vector for the second portion of the word sequence, wherein the second phrase-level context feature vector is determined based on the second pooled phrase-level feature vector and the first phrase-level context feature vector;

determining, based on the plurality of phrase-level context feature vectors, a first sentence-level feature vector for a first sentence in the word sequence and a second sentence-level feature vector for a second sentence in the word sequence;

and

determining a pooled sentence-level feature vector by pooling the first sentence-level feature vector and the second sentence-level feature vector, wherein the sentiment is determined further based on the pooled sentence-level feature vector.

6. The computer-readable storage medium of claim 1 , wherein the first phrase-level feature vector and the second phrase-level feature vector are pooled using a max pooling function.

7. The computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

determining, based on the plurality of word-level context feature vectors, a plurality of phrase-level feature vectors for a plurality of phrases in the word sequence, the plurality of phrase-level feature vectors including the first phrase-level feature vector and the second phrase-level feature vector, and wherein the plurality of phrases includes the first phrase and the second phrase;

and

determining a plurality of pooled phrase-level feature vectors by pooling the plurality of word-level context feature vectors, the plurality of pooled phrase-level feature vectors including the pooled phrase-level feature vector, wherein the pooling is such that a total number of pooled phrase-level feature vectors in the plurality of pooled phrase-level feature vectors is less than a total number of phrases in the plurality of phrases by a predetermined factor.

8. The computer-readable storage medium of claim 1 , wherein:

the plurality of word-level context feature vectors includes a fourth word-level context feature vector for a fourth word in the word sequence, the fourth word is a starting word of the first phrase;

the first phrase-level feature vector comprises the fourth word-level context feature vector;

the plurality of word-level context feature vectors includes a fifth word-level context feature vector for a fifth word in the word sequence, the fifth word is an ending word of the first phrase; and

the first phrase-level feature vector comprises the fifth word-level context feature vector.

9. The computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

determining one or more phrase boundaries in the word sequence; and

identifying the first phrase and the second phrase in the word sequence based on the one or more phrase boundaries.

10. The computer-readable storage medium of claim 1 , wherein the plurality of word-level context feature vectors is determined using a long short-term memory (LSTM) network.

11. The computer-readable storage medium of claim 1 , wherein the one or more programs further include instructions for:

determining, based on the pooled phrase-level feature vector, a probability distribution across the plurality of predefined sentiment states, wherein the sentiment is selected based on the probability distribution.

12. A method for predicting sentiment from text to generate a result, the method comprising:

at an electronic device having a processor and memory:

receiving a word sequence in textual form, wherein a first word, a second word, and a third word are consecutively arranged in the word sequence;

determining a plurality of word-level context feature vectors for the word sequence;

determining, based on the plurality of word-level context feature vectors, a first phrase-level feature vector for a first phrase in the word sequence and a second phrase-level feature vector for a second phrase in the word sequence;

determining a pooled phrase-level feature vector by pooling the first phrase-level feature vector and the second phrase-level feature vector;

determining, based on the pooled phrase-level feature vector, a sentiment conveyed by the word sequence, wherein the sentiment is selected from a plurality of predefined sentiment states based on the pooled phrase-level feature vector; and

outputting the result, wherein the result is generated by performing an action in accordance with the determined sentiment.

13. An electronic device, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving a word sequence in textual form, wherein a first word, a second word, and a third word are consecutively arranged in the word sequence;

determining a plurality of word-level context feature vectors for the word sequence;

determining, based on the plurality of word-level context feature vectors, a first phrase-level feature vector for a first phrase in the word sequence and a second phrase-level feature vector for a second phrase in the word sequence;

determining a pooled phrase-level feature vector by pooling the first phrase-level feature vector and the second phrase-level feature vector;

determining, based on the pooled phrase-level feature vector, a sentiment conveyed by the word sequence, wherein the sentiment is selected from a plurality of predefined sentiment states based on the pooled phrase-level feature vector; and

outputting the result, wherein the result is generated by performing an action in accordance with the determined sentiment.

14. The computer-readable storage medium of claim 1 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first forward word-level context feature vector.

15. The computer-readable storage medium of claim 14 , wherein:

the first phrase includes the first word, the second word, and the third word;

the plurality of word-level context feature vectors includes a third word-level context feature vector for the third word; and

the first phrase-level feature vector is determined based on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

16. The computer-readable storage medium of claim 15 , wherein:

the first phrase-level feature vector is determined by applying an attention mechanism on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

17. The computer-readable storage medium of claim 1 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first word-level context feature vector.

18. The computer-readable storage medium of claim 1 , wherein the plurality of word-level context feature vectors for the word sequence are a plurality of forward word-level context feature vectors for the word sequence, the first phrase-level feature vector is a first forward phrase-level feature vector, the second phrase-level feature vector is a second forward phrase-level feature vector, and wherein the one or more programs further include instructions for:

determining a plurality of backward word-level context feature vectors for the word sequence;

determining, based on the plurality of backward word-level context feature vectors, a first backward phrase-level feature vector for the first phrase and a second backward phrase-level feature vector for the second phrase; and

determining the pooled phrase-level feature vector by pooling the first forward phrase-level feature vector, the second forward phrase-level feature vector, the first backward phrase-level feature vector, and the second backward phrase-level feature vector.

19. The computer-readable storage medium of claim 18 , wherein determining the pooled phrase-level feature vector further comprises:

determining a first combined phrase-level feature vector for the first phrase by pooling the first forward phrase-level feature vector and the first backward phrase-level feature vector;

determining a second combined phrase-level feature vector for the second phrase by pooling the second forward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the first combined phrase-level feature vector and the second combined phrase-level feature vector to determine the pooled phrase-level feature vector.

20. The computer-readable storage medium of claim 18 , wherein determining the pooled phrase-level feature vector further comprises:

determining a combined forward phrase-level feature vector by pooling the first forward phrase-level feature vector and the second forward phrase-level feature vector;

determining a combined backward phrase-level feature vector by pooling the first backward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the combined forward phrase-level feature vector and the combined backward phrase-level feature vector to determine the pooled phrase-level feature vector.

21. The method of claim 12 , wherein the first phrase and the second phrase are adjacent phrases in the word sequence.

22. The method of claim 21 , further comprising:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence and a fourth phrase-level feature vector for a fourth phrase in the word sequence, wherein the third phrase and the fourth phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the third phrase-level feature vector and the fourth phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

23. The method of claim 21 , further comprising:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence, wherein the second phrase and the third phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the second phrase-level feature vector and the third phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

24. The method of claim 22 , further comprising:

determining a plurality of pooled phrase-level feature vectors that includes the first pooled phrase-level feature vector for a first portion of the word sequence, the second pooled phrase-level feature vector for a second portion of the word sequence, and a third pooled phrase-level feature vector for a third portion of the word sequence, wherein the first, second, and third portions of the word sequence are consecutive portions of the word sequence;

determining, based on the plurality of pooled phrase-level feature vectors, a plurality of phrase-level context feature vectors that includes a first phrase-level context feature vector for the first portion of the word sequence and a second phrase-level context feature vector for the second portion of the word sequence, wherein the second phrase-level context feature vector is determined based on the second pooled phrase-level feature vector and the first phrase-level context feature vector;

determining, based on the plurality of phrase-level context feature vectors, a first sentence-level feature vector for a first sentence in the word sequence and a second sentence-level feature vector for a second sentence in the word sequence;

and

determining a pooled sentence-level feature vector by pooling the first sentence-level feature vector and the second sentence-level feature vector, wherein the sentiment is determined further based on the pooled sentence-level feature vector.

25. The method of claim 12 , wherein the first phrase-level feature vector and the second phrase-level feature vector are pooled using a max pooling function.

26. The method of claim 12 , further comprising:

determining, based on the plurality of word-level context feature vectors, a plurality of phrase-level feature vectors for a plurality of phrases in the word sequence, the plurality of phrase-level feature vectors including the first phrase-level feature vector and the second phrase-level feature vector, and wherein the plurality of phrases includes the first phrase and the second phrase;

and

determining a plurality of pooled phrase-level feature vectors by pooling the plurality of word-level context feature vectors, the plurality of pooled phrase-level feature vectors including the pooled phrase-level feature vector, wherein the pooling is such that a total number of pooled phrase-level feature vectors in the plurality of pooled phrase-level feature vectors is less than a total number of phrases in the plurality of phrases by a predetermined factor.

27. The method of claim 12 , wherein:

the plurality of word-level context feature vectors includes a fourth word-level context feature vector for a fourth word in the word sequence, the fourth word is a starting word of the first phrase;

the first phrase-level feature vector comprises the fourth word-level context feature vector;

the plurality of word-level context feature vectors includes a fifth word-level context feature vector for a fifth word in the word sequence, the fifth word is an ending word of the first phrase; and

the first phrase-level feature vector comprises the fifth word-level context feature vector.

28. The method of claim 12 , further comprising:

determining one or more phrase boundaries in the word sequence; and

identifying the first phrase and the second phrase in the word sequence based on the one or more phrase boundaries.

29. The method of claim 12 , wherein the plurality of word-level context feature vectors is determined using a long short-term memory (LSTM) network.

30. The method of claim 12 , further comprising:

determining, based on the pooled phrase-level feature vector, a probability distribution across the plurality of predefined sentiment states, wherein the sentiment is selected based on the probability distribution.

31. The method of claim 12 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first forward word-level context feature vector.

32. The method of claim 31 , wherein:

the first phrase includes the first word, the second word, and the third word;

the plurality of word-level context feature vectors includes a third word-level context feature vector for the third word; and

the first phrase-level feature vector is determined based on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

33. The method of claim 32 , wherein:

the first phrase-level feature vector is determined by applying an attention mechanism on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

34. The method of claim 12 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first word-level context feature vector.

35. The method of claim 12 , wherein the plurality of word-level context feature vectors for the word sequence are a plurality of forward word-level context feature vectors for the word sequence, the first phrase-level feature vector is a first forward phrase-level feature vector, the second phrase-level feature vector is a second forward phrase-level feature vector, and wherein the method further comprises:

determining a plurality of backward word-level context feature vectors for the word sequence;

determining, based on the plurality of backward word-level context feature vectors, a first backward phrase-level feature vector for the first phrase and a second backward phrase-level feature vector for the second phrase; and

determining the pooled phrase-level feature vector by pooling the first forward phrase-level feature vector, the second forward phrase-level feature vector, the first backward phrase-level feature vector, and the second backward phrase-level feature vector.

36. The method of claim 35 , wherein determining the pooled phrase-level feature vector further comprises:

determining a first combined phrase-level feature vector for the first phrase by pooling the first forward phrase-level feature vector and the first backward phrase-level feature vector;

determining a second combined phrase-level feature vector for the second phrase by pooling the second forward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the first combined phrase-level feature vector and the second combined phrase-level feature vector to determine the pooled phrase-level feature vector.

37. The method of claim 35 , wherein determining the pooled phrase-level feature vector further comprises:

determining a combined forward phrase-level feature vector by pooling the first forward phrase-level feature vector and the second forward phrase-level feature vector;

determining a combined backward phrase-level feature vector by pooling the first backward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the combined forward phrase-level feature vector and the combined backward phrase-level feature vector to determine the pooled phrase-level feature vector.

38. The electronic device of claim 13 , wherein the first phrase and the second phrase are adjacent phrases in the word sequence.

39. The electronic device of claim 38 , the one or more programs further including instructions for:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence and a fourth phrase-level feature vector for a fourth phrase in the word sequence, wherein the third phrase and the fourth phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the third phrase-level feature vector and the fourth phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

40. The electronic device of claim 38 , the one or more programs further including instructions for:

determining, based on the plurality of word-level context feature vectors, a third phrase-level feature vector for a third phrase in the word sequence, wherein the second phrase and the third phrase are adjacent phrases in the word sequence;

and

determining a second pooled phrase-level feature vector by pooling the second phrase-level feature vector and the third phrase-level feature vector, wherein the sentiment is determined further based on the second pooled phrase-level feature vector.

41. The electronic device of claim 39 , the one or more programs further including instructions for:

determining a plurality of pooled phrase-level feature vectors that includes the first pooled phrase-level feature vector for a first portion of the word sequence, the second pooled phrase-level feature vector for a second portion of the word sequence, and a third pooled phrase-level feature vector for a third portion of the word sequence, wherein the first, second, and third portions of the word sequence are consecutive portions of the word sequence;

determining, based on the plurality of pooled phrase-level feature vectors, a plurality of phrase-level context feature vectors that includes a first phrase-level context feature vector for the first portion of the word sequence and a second phrase-level context feature vector for the second portion of the word sequence, wherein the second phrase-level context feature vector is determined based on the second pooled phrase-level feature vector and the first phrase-level context feature vector;

determining, based on the plurality of phrase-level context feature vectors, a first sentence-level feature vector for a first sentence in the word sequence and a second sentence-level feature vector for a second sentence in the word sequence;

and

determining a pooled sentence-level feature vector by pooling the first sentence-level feature vector and the second sentence-level feature vector, wherein the sentiment is determined further based on the pooled sentence-level feature vector.

42. The electronic device of claim 13 , wherein the first phrase-level feature vector and the second phrase-level feature vector are pooled using a max pooling function.

43. The electronic device of claim 13 , the one or more programs further including instructions for:

determining, based on the plurality of word-level context feature vectors, a plurality of phrase-level feature vectors for a plurality of phrases in the word sequence, the plurality of phrase-level feature vectors including the first phrase-level feature vector and the second phrase-level feature vector, and wherein the plurality of phrases includes the first phrase and the second phrase;

and

determining a plurality of pooled phrase-level feature vectors by pooling the plurality of word-level context feature vectors, the plurality of pooled phrase-level feature vectors including the pooled phrase-level feature vector, wherein the pooling is such that a total number of pooled phrase-level feature vectors in the plurality of pooled phrase-level feature vectors is less than a total number of phrases in the plurality of phrases by a predetermined factor.

44. The electronic device of claim 13 , wherein:

the plurality of word-level context feature vectors includes a fourth word-level context feature vector for a fourth word in the word sequence, the fourth word is a starting word of the first phrase;

the first phrase-level feature vector comprises the fourth word-level context feature vector;

the plurality of word-level context feature vectors includes a fifth word-level context feature vector for a fifth word in the word sequence, the fifth word is an ending word of the first phrase; and

the first phrase-level feature vector comprises the fifth word-level context feature vector.

45. The electronic device of claim 13 , the one or more programs further including instructions for:

determining one or more phrase boundaries in the word sequence; and

identifying the first phrase and the second phrase in the word sequence based on the one or more phrase boundaries.

46. The electronic device of claim 13 , wherein the plurality of word-level context feature vectors is determined using a long short-term memory (LSTM) network.

47. The electronic device of claim 13 , the one or more programs further including instructions for:

determining, based on the pooled phrase-level feature vector, a probability distribution across the plurality of predefined sentiment states, wherein the sentiment is selected based on the probability distribution.

48. The electronic device of claim 13 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first forward word-level context feature vector.

49. The electronic device of claim 48 , wherein:

the first phrase includes the first word, the second word, and the third word;

the plurality of word-level context feature vectors includes a third word-level context feature vector for the third word; and

the first phrase-level feature vector is determined based on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

50. The electronic device of claim 49 , wherein:

the first phrase-level feature vector is determined by applying an attention mechanism on the first word-level context feature vector, the second word-level context feature vector, and the third word-level context feature vector.

51. The electronic device of claim 13 , wherein the plurality of word-level context feature vectors include a first word-level context feature vector for the first word and a second word-level context feature vector for the second word after the first word, wherein the second word-level context feature vector is determined based on a vector representation of the second word and the first word-level context feature vector.

52. The electronic device of claim 13 , wherein the plurality of word-level context feature vectors for the word sequence are a plurality of forward word-level context feature vectors for the word sequence, the first phrase-level feature vector is a first forward phrase-level feature vector, the second phrase-level feature vector is a second forward phrase-level feature vector, and wherein the method further comprises:

determining a plurality of backward word-level context feature vectors for the word sequence;

determining, based on the plurality of backward word-level context feature vectors, a first backward phrase-level feature vector for the first phrase and a second backward phrase-level feature vector for the second phrase; and

determining the pooled phrase-level feature vector by pooling the first forward phrase-level feature vector, the second forward phrase-level feature vector, the first backward phrase-level feature vector, and the second backward phrase-level feature vector.

53. The electronic device of claim 52 , wherein determining the pooled phrase-level feature vector further comprises:

determining a first combined phrase-level feature vector for the first phrase by pooling the first forward phrase-level feature vector and the first backward phrase-level feature vector;

determining a second combined phrase-level feature vector for the second phrase by pooling the second forward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the first combined phrase-level feature vector and the second combined phrase-level feature vector to determine the pooled phrase-level feature vector.

54. The electronic device of claim 52 , wherein determining the pooled phrase-level feature vector further comprises:

determining a combined forward phrase-level feature vector by pooling the first forward phrase-level feature vector and the second forward phrase-level feature vector;

determining a combined backward phrase-level feature vector by pooling the first backward phrase-level feature vector and the second backward phrase-level feature vector; and

pooling the combined forward phrase-level feature vector and the combined backward phrase-level feature vector to determine the pooled phrase-level feature vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2019
From: BELLEGARDA, JEROME R.
To: APPLE INC.
Reel/Frame 048578/0929 →
Continuity (2)
Provisional Application 62737848 · Sep 27, 2018
Related Publication 20200104369A1 · Apr 2, 2020
Cited By (2)
US 12,260,178 US 12,430,512