Multimodal sentiment classification
Sentiment classification can be implemented by an entity-level multimodal sentiment classification neural network. The neural network can include left, right, and target entity subnetworks. The neural network can further include an image network that generates representation data that is combined and weighted with data output by the left, right, and target entity subnetworks to output a sentiment classification for an entity included in a network post.
1. A method comprising:
identifying a multimodal message comprising an image and a plurality of terms, the plurality of terms comprising one or more entity terms and non-entity terms;
generating, using a recurrent neural network, a target entity representation from the one or more entity terms;
generating, using one or more additional recurrent neural networks, a textual representation from the non-entity terms;
generating, using a convolutional neural network (CNN), an image representation from the image in the multimodal message; and
generating a sentiment classification from the target entity representation, the textual representation, and the image representation,
wherein the one or more additional recurrent neural networks comprise a left network and a right network, the left network configured to generate a left representation from terms to the left of the one or more entity terms, the right network configured to generate a right representation from terms to the right of the one or more entity terms,
wherein the target entity representation is linearly combined with the left representation and linearly combined with the right representation.
2. The method of claim 1 , further comprising:
generating a merged representation by merging the target entity representation, the textual representation, and the image representation, wherein the sentiment classification is generated by applying a classification layer to the merged representation.
3. The method of claim 2 , wherein the sentiment classification is one or more of the following: a positive sentiment, a negative sentiment, or neutral sentiment.
4. The method of claim 2 , wherein the merging is concatenation.
5. The method of claim 2 , wherein the classification layer is a SoftMax layer.
6. The method of claim 1 , further comprising:
generating the image using an image sensor of a user device.
7. The method of claim 6 , further comprising:
identifying content data using the generated sentiment classification; and
displaying the content data on the user device.
8. The method of claim 7 , further comprising:
generating a modified multimodal message from the content data and the multimodal message; and
publishing, to a network site, the modified multimodal message as an ephemeral message.
9. The method of claim 1 , wherein the image representation is linearly combined with the left representation and linearly combined with the right representation.
10. The method of claim 9 , wherein the linearly combined left and right representations are weighted to emphasize significant characters.
11. The method of claim 1 , wherein the image representation is weighted to emphasize regions within the image that are more related to the one or more entity terms.
12. The method of claim 1 , wherein the recurrent neural network, the one or more additional recurrent neural networks, and the convolutional neural network are subnetworks of a network trained end-to-end on a training dataset.
13. The method of claim 12 , wherein the training dataset comprises a plurality of multimodal messages, each of the multimodal messages comprising an image and one or more terms corresponding to an entity.
14. The method of claim 1 , further comprising:
identifying, using an entity recognition machine learning scheme, the one or more entity terms in the plurality of terms.
15. A system comprising:
one or more processors of a machine; and
a memory storing instructions that, when executed by at least one processor among the one or more processors, causes the machine to perform operations comprising:
identifying a multimodal message comprising an image and a plurality of terms, the plurality of terms comprising one or more entity terms and non-entity terms;
generating, using a recurrent neural network, a target entity representation from the one or more entity terms;
generating, using one or more additional recurrent neural networks, a textual representation from the non-entity terms;
generating, using a convolutional neural network (CNN), an image representation from the image in the multimodal message; and
generating a sentiment classification from the target entity representation, the textual representation, and the image representation,
wherein the one or more additional recurrent neural networks comprise a left network and a right network, the left network configured to generate a left representation from terms to the left of the one or more entity terms, the right network configured to generate a right representation from terms to the right of the one or more entity terms,
wherein the target entity representation is linearly combined with the left representation and linearly combined with the right representation.
16. The system of claim 15 , the operations further comprising:
generating a merged representation by merging the target entity representation, the textual representation, and the image representation, wherein the sentiment classification is generated by applying a classification layer to the merged representation.
17. The system of claim 15 , wherein the sentiment classification is one or more of the following: a positive sentiment, a negative sentiment, or neutral sentiment.
18. A machine-readable storage device embodying instructions that, when executed by a device, cause the device to perform operations comprising:
identifying a multimodal message comprising an image and a plurality of terms, the plurality of terms comprising one or more entity terms and non-entity terms;
generating, using a recurrent neural network, a target entity representation from the one or more entity terms;
generating, using one or more additional recurrent neural networks, a textual representation from the non-entity terms;
generating, using a convolutional neural network (CNN), an image representation from the image in the multimodal message; and
generating a sentiment classification from the target entity representation, the textual representation, and the image representation,
wherein the one or more additional recurrent neural networks comprise a left network and a right network, the left network configured to generate a left representation from terms to the left of the one or more entity terms, the right network configured to generate a right representation from terms to the right of the one or more entity terms,
wherein the target entity representation is linearly combined with the left representation and linearly combined with the right representation.