IP Library › Granted Patent US 12,718,539
Granted Patent B2
US 12,718,539 · App. 18/057,643 · Granted Aug 25, 2026

Multi-modal understanding of emotions in video content

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,718,539
App. No.
18/057,643
Filed
Nov 21, 2022
Granted
Aug 25, 2026
Kind
B2
Art Unit
2675
USPC
382/100
Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2022
From: CHOUDHARY, DIVYA; GOYAL, PALASH
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 061845/0614 →
Continuity (1)
Related Publication 20240169711A1 · May 23, 2024
References Cited (36)
US 7242810B2 · Chang · 2007 [cited by applicant]
US 11238289B1 · Tao et al. · 2022 [cited by applicant]
US 11282297B2 · Zeng et al. · 2022 [cited by applicant]
US 11375256B1 · Dorner · 2022 [cited by applicant]
US 11386712B2 · Yadav et al. · 2022 [cited by applicant]
US 20190341025A1 · Omote · 2019 [cited by examiner]
US 20200279156A1 · Cai · 2020 [cited by examiner]
US 20210151034A1 · Hasan · 2021 [cited by examiner]
US 20210264921A1 · Reece et al. · 2021 [cited by applicant]
US 20220138472A1 · Mittal et al. · 2022 [cited by applicant]
US 20220270636A1 · Tao · 2022 [cited by examiner]
US 20230154172A1 · Wasnik · 2023 [cited by examiner]
US 20240257554A1 · Xu · 2024 [cited by examiner]
US 20250078569A1 · Zhang · 2025 [cited by examiner]
CN 111164601A · 2020 [cited by applicant]
CN 111563422A · 2020 [cited by applicant]
CN 112699774A · 2021 [cited by applicant]
EP 3942552A1 · 2022 [cited by applicant]
Complex Emotion Recognition via Facial Expressions with Label Noises Self-Cure Relation Networks Xiaoqing Wang , 1,2 Yaocheng Wang , (Year: 2023). [cited by examiner]
Chen et al., “Self-Cure Network with Two-Stage Method for Facial Expression Recognition,” 2021 27th International Conference on Mechatronics and Machine Vision in Practice, Jan. 2022, 6 pages. EFS Web 2.1.18 (Year: 2022… [cited by examiner]
International Search Report and Written Opinion of the International Searching Authority dated Oct. 19, 2023 in connection with International Patent Application No. PCT/KR2023/009600, 10 pages. [cited by applicant]
Zhao et al., “An End-to-End Visual-Audio Attention Network for Emotion Recognition in User-Generated Videos,” AAAI 2020, Feb. 2020, 9 pages. [cited by applicant]
Chen et al., “Emotion Recognition with Audio, Video, EEG, and EMG: A Dataset and Baseline Approaches,” Digital Object Identifier 10.1109, Jan. 2022, 14 pages. [cited by applicant]
Schoneveld et al., “Leveraging Recent Advances in Deep Learning for Audio-Visual Emotion Recognition,” Elsevier, Pattern Recognition Letters, Sep. 2021, 8 pages. [cited by applicant]
Ma et al., “Acceleration of multi-task cascaded convolutional networks,” The Institute of Engineering and Technology Journals, Image Processing, Jul. 2020, 7 pages. [cited by applicant]
Chen et al., “Self-Cure Network with Two-Stage Method for Facial Expression Recognition,” 2021 27th International Conference on Mechatronics and Machine Vision in Practice, Jan. 2022, 6 pages. [cited by applicant]
Riggs, “Creating Audio Features with PyAudio Analysis,” Dolbyio Blog, Dec. 2021, 11 pages. [cited by applicant]
Gong et al., “PSLA: Improving Audio Tagging with Pretraining, Sampling, Labeling, and Aggregation,” IEEE/ACM Transactions on Audio, Speech, And Language Processing, Nov. 2021, 15 pages. [cited by applicant]
Vaswani et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 2017, 15 pages. [cited by applicant]
Wang et al., “Distilled Dual-Encoder Model for Vision-Language Understanding,” https://arxiv.org/abs/2112.08723, Oct. 2022, 12 pages. [cited by applicant]
Li et al., “Pre-training Model Based on Parallel Cross-Modality Fusion Layer,” PLoS ONE 17(2), Feb. 2022, 14 pages. [cited by applicant]
Supplementary European Search Report dated Aug. 18, 2025 in connection with European Patent Application No. 23894708.9, 9 pages. [cited by applicant]
Zhou et al., “Information Fusion in Attention Networks Using Adaptive and Multi-Level Factorized Bilinear Pooling for Audio-Visual Emotion Recognition,” IEEE/ACM Transactions on Audio, Speech, And Language Processing, N… [cited by applicant]
Wang et al., “Unlocking the Emotional World of Visual Media: An Overview of the Science, Research, and Impact of Understanding Emotion,” Proceedings of the IEEE, Jul. 2023, 49 pages. [cited by applicant]
Ortega et al., “Emotion Recognition Using Fusion of Audio and Video Features,” Jun. 2019, 6 pages. [cited by applicant]
Wang et al., “A systematic review on affective computing: emotion models, databases, and recent advances,” Mar. 2022, 48 pages. [cited by applicant]