IP Library › Granted Patent US 12,609,114
Granted Patent B2
US 12,609,114 · App. 18/676,744 · Granted Apr 21, 2026

Multi-modal cross attention sentiment analysis of textual and audio embeddings

Inventors: Priyanka Pathak (Hyderabad, IN); Saurabh Jha (Austin, TX)
Assignee: Dell Products L.P.
G10L15/1815G10L15/16G10L15/183G10L15/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,609,114
App. No.
18/676,744
Granted
Apr 21, 2026
Kind
B2
Abstract

A method for managing sentiment analysis includes obtaining, by a data classification system, raw data, associated with a meeting, wherein the raw data the comprises a textual transcript and an audio file of the meeting, applying a generative adversarial network (GAN) to the audio file to obtain audio embeddings and extracted text, applying the textual transcripts and the extracted text to a trained language model to obtain textual embeddings, applying the audio embeddings and the textual embeddings to a multi-modal cross attention module to obtain a fused embedding, performing a sentiment classification on the fused embedding to obtain a sentiment prediction, and implementing a remediation on the data classification system using the sentiment prediction.

Claims (52)

1 . A method for managing sentiment analysis, the method comprising:

obtaining, by a data classification system, raw data, associated with a meeting, wherein the raw data the comprises a textual transcript and an audio file of the meeting;

applying a generative adversarial network (GAN) to the audio file to obtain audio embeddings and extracted text;

applying the textual transcripts and the extracted text to a trained language model to obtain textual embeddings;

applying the audio embeddings and the textual embeddings to a multi-modal cross attention module to obtain a fused embedding;

performing a sentiment classification on the fused embedding to obtain a sentiment prediction; and

implementing a remediation on the data classification system using the sentiment prediction.

2 . The method of claim 1 , wherein the GAN comprises a sample generator and a sample discriminator.

3 . The method of claim 2 , wherein applying the GAN to the audio file comprises:

using data distribution of the audio file to obtain a set of audio samples;

generating, by the sample generator and using the data distribution, a set of synthetic samples;

applying the set of synthetic samples and the set of audio samples to the sample discriminator and modifying the set of synthetic samples in an iterative fashion until a pre-defined percentage threshold of the set of synthetic samples are labeled as not synthetic; and

obtaining the audio embeddings as a final modified set of synthetic samples.

4 . The method of claim 2 , wherein the sample generator has a higher learning rate than that of the sample discriminator.

5 . The method of claim 1 , wherein the sentiment prediction comprises tagged textual embeddings of the textual transcript, wherein the tagged textual embeddings comprises at least one of: a positive tagged embedding, a negative tagged embedding, a very positive tagged embedding, a very negative tagged embedding, and a neutral tagged embedding.

6 . The method of claim 1 , wherein the trained language model is an auto-regressive language model trained using a set of earnings call transcripts.

7 . The method of claim 1 , wherein one of the fused embeddings is tagged with a pitch, frequency, and tone associated with a textual embedding based on a corresponding audio embedding.

8 . A non-transitory computer readable medium comprising computer readable program code, which when executed by a computer processor enables the computer processor to perform a method for managing sentiment analysis, the method comprising:

obtaining, by a data classification system, raw data, associated with a meeting, wherein the raw data the comprises a textual transcript and an audio file of the meeting;

applying a generative adversarial network (GAN) to the audio file to obtain audio embeddings and extracted text;

applying the textual transcripts and the extracted text to a trained language model to obtain textual embeddings;

applying the audio embeddings and the textual embeddings to a multi-modal cross attention module to obtain a fused embedding;

performing a sentiment classification on the fused embedding to obtain a sentiment prediction; and

implementing a remediation on the data classification system using the sentiment prediction.

9 . The non-transitory computer readable medium of claim 8 , wherein the GAN comprises a sample generator and a sample discriminator.

10 . The non-transitory computer readable medium of claim 9 , wherein applying the GAN to the audio file comprises:

using data distribution of the audio file to obtain a set of audio samples;

generating, by the sample generator and using the data distribution, a set of synthetic samples;

applying the set of synthetic samples and the set of audio samples to the sample discriminator and modifying the set of synthetic samples in an iterative fashion until a pre-defined percentage threshold of the set of synthetic samples are labeled as not synthetic; and

obtaining the audio embeddings as a final modified set of synthetic samples.

11 . The non-transitory computer readable medium of claim 9 , wherein the sample generator has a higher learning rate than that of the sample discriminator.

12 . The non-transitory computer readable medium of claim 8 , wherein the sentiment prediction comprises tagged textual embeddings of the textual transcript, wherein the tagged textual embeddings comprises at least one of: a positive tagged embedding, a negative tagged embedding, a very positive tagged embedding, a very negative tagged embedding, and a neutral tagged embedding.

13 . The non-transitory computer readable medium of claim 8 , wherein the trained language model is an auto-regressive language model trained using a set of earnings call transcripts.

14 . The non-transitory computer readable medium of claim 8 , wherein one of the fused embeddings is tagged with a pitch, frequency, and tone associated with a textual embedding based on a corresponding audio embedding.

15 . A system, comprising:

a processor; and

memory including instructions, which when executed by the processor, perform a method comprising:

obtaining, by a data classification system, raw data, associated with a meeting, wherein the raw data the comprises a textual transcript and an audio file of the meeting;

applying a generative adversarial network (GAN) to the audio file to obtain audio embeddings and extracted text;

applying the textual transcripts and the extracted text to a trained language model to obtain textual embeddings;

applying the audio embeddings and the textual embeddings to a multi-modal cross attention module to obtain a fused embedding;

performing a sentiment classification on the fused embedding to obtain a sentiment prediction; and

implementing a remediation on the data classification system using the sentiment prediction.

16 . The system of claim 15 , wherein the GAN comprises a sample generator and a sample discriminator, and wherein the sample generator has a higher learning rate than that of the sample discriminator.

17 . The system of claim 16 , wherein applying the GAN to the audio file comprises:

using data distribution of the audio file to obtain a set of audio samples;

generating, by the sample generator and using the data distribution, a set of synthetic samples;

applying the set of synthetic samples and the set of audio samples to the sample discriminator and modifying the set of synthetic samples in an iterative fashion until a pre-defined percentage threshold of the set of synthetic samples are labeled as not synthetic; and

obtaining the audio embeddings as a final modified set of synthetic samples.

18 . The system of claim 15 , wherein the sentiment prediction comprises tagged textual embeddings of the textual transcript, wherein the tagged textual embeddings comprises at least one of: a positive tagged embedding, a negative tagged embedding, a very positive tagged embedding, a very negative tagged embedding, and a neutral tagged embedding.

19 . The system of claim 15 , wherein the trained language model is an auto-regressive language model trained using a set of earnings call transcripts.

20 . The system of claim 15 , wherein one of the fused embeddings is tagged with a pitch, frequency, and tone associated with a textual embedding based on a corresponding audio embedding.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 29, 2024
From: PATHAK, PRIYANKA; JHA, SAURABH
To: DELL PRODUCTS L.P.
Reel/Frame 067549/0392 →
Continuity (1)
Related Publication 20250372085A1 · Dec 4, 2025
References Cited (12)
US 11521639B1 · Shon · 2022 [cited by examiner]
US 11854538B1 · Rozgic · 2023 [cited by examiner]
US 20200335092A1 · Georgiou · 2020 [cited by examiner]
US 20240127812A1 · Inbavaluthi · 2024 [cited by examiner]
US 20250201267A1 · Kim · 2025 [cited by examiner]
US 20250209802A1 · Singh · 2025 [cited by examiner]
US 20250232122A1 · Hajavi · 2025 [cited by examiner]
US 20250252949A1 · Saha · 2025 [cited by examiner]
Chatziagapi, Aggelina, et al. “Data augmentation using GANs for speech emotion recognition.” Interspeech. 2019. (Year: 2019). [cited by examiner]
Zadeh, Amir, et al. “Tensor fusion network for multimodal sentiment analysis.” arXiv preprint arXiv:1707.07250 (2017). (Year: 2017). [cited by examiner]
Sahu, Saurabh, Rahul Gupta, and Carol Espy-Wilson. “On enhancing speech emotion recognition using generative adversarial networks.” arXiv preprint arXiv:1806.06626 (2018). (Year: 2018). [cited by examiner]
Tharani, A. D., and J. Aravinth. “Multimodal sentimental analysis using hierarchical fusion technique.” 2023 IEEE 4th Annual Flagship India Council International Subsections Conference (INDISCON). IEEE, 2023. (Year: 202… [cited by examiner]