IP Library › Granted Patent US 12,265,783
Granted Patent B2
US 12,265,783 · App. 18/054,511 · Granted Apr 1, 2025

Systems and methods for multi-modal conversation summarization on a conversation platform

Inventors: Divyansh Agarwal (San Francisco, CA); Chien-Sheng Wu (Mountain View, CA); Tian Xie (San Jose, CA)
Assignee: Salesforce, Inc.
G06F40/166G06F16/3344G06F16/345G06F40/205H04L51/216G06V20/47
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,783
App. No.
18/054,511
Granted
Apr 1, 2025
Kind
B2
Abstract

Embodiments described herein provide a multi-modal search-and-summarize tool for message platforms. Specifically, the multi-modal search-and-summarize tool may monitor conversational content of different formats, e.g., text, image, video, etc., and use multi-modal summarization models to generate a summary of the conversation channel. The summarization may be conducted via a search-and-summarize process in response to a specific user query, e.g., a user may enter “what did John and Josh say about the presentation tomorrow?” The multi-modal summarization model would first search for relevant conversation messages between user John and user Josh, identify communication files of different format (e.g., text messages, emojis, multimedia attachments, etc.), and then input the communication files to respective text or image encoders to generate a summary of the communication content.

Claims (56)

1. A method of multi-modal summarization of communication on a messaging platform, the method comprising:

receiving, via a user interface, a user request for summarizing communication messages relating to a topic between a first user and a second user on a messaging platform;

searching, via a search engine, for messages between the first user and the second user on the messaging platform;

filtering the messages based on the topic from the user request by predicting, via a topic classification model, whether the messages are related to the topic and excluding a subset of messages that are predicted to be unrelated to the topic;

generating, by a text encoder of a multi-modal summarization model, a text representation from an input sequence corresponding to textual content from the filtered messages;

generating, by an image encoder of the multi-modal summarization model, an image representation of visual features in a multimedia attachment file from the filtered messages;

generating, via a decoder of the multi-modal summarization model, a text summary summarizing both the filtered messages and the multimedia attachment file based on a combination of the text representation generated by the text encoder and the image representation generated by the image encoder,

wherein the text summary comprises a text that references a time and/or a sender of the multimedia attachment file generated from metadata of the multimedia attachment file; and

transmitting, via the user interface, the generated text summary relating to the topic in response to the user request.

2. The method of claim 1 , wherein the user request takes a form of a natural language question, and wherein the method further comprises:

identifying, by parsing the natural language question, the first user, the second user and the topic.

3. The method of claim 1 , wherein the user request further comprises one or more parameters entered via the user interface, and wherein the one or more parameters indicate any of:

a time range;

one or more conversation channels on the messaging platform; and

a file format.

4. The method of claim 1 , wherein the topic classification model is trained on a dataset of messages annotated with respective topic labels.

5. The method of claim 1 , wherein the multimedia attachment file includes any of a spreadsheet file, a presentation slide file, an audio file, a video file and an image file.

6. The method of claim 5 , further comprising:

converting the audio file into a text document; or

extracting one or more video frames from the video file.

7. The method of claim 1 , further comprising:

receiving, from the user interface, user feedback relating to the text summary;

generating an updated text summary according to the user feedback; and

incorporating the updated text summary into a training dataset for training the multi-modal summarization model.

8. The method of claim 1 , wherein the predicting, via the topic classification model, whether the messages are related to the topic comprises:

extracting one or more entities from the user request; and

predicting, one or more keywords related to the extracted one or more entities.

9. A system of multi-modal summarization of communication on a messaging platform, the system comprising:

a user interface that receives a user request for summarizing communication messages relating to a topic between a first user and a second user on a messaging platform;

a memory storing a plurality of processor-executable instructions; and

one or more processors executing the plurality of processor-executable instructions to perform operations comprising:

searching, via a search engine, for messages between the first user and the second user on the messaging platform;

filtering the messages based on the topic from the user request by predicting, via a topic classification model, whether the messages are related to the topic and excluding a subset of messages that are predicted to be unrelated to the topic;

generating, by a text encoder of a multi-modal summarization model, a text representation from an input sequence corresponding to textual content from the filtered messages;

generating, by an image encoder of the multi-modal summarization model, an image representation of visual features in a multimedia attachment file from the filtered messages; and

generating, via a decoder of the multi-modal summarization model, a text summary summarizing both the filtered messages and the multimedia attachment file based on a combination of the text representation generated by the text encoder and the image representation generated by the image encoder, wherein the text summary comprises a text that references a time and/or a sender of the multimedia attachment file generated from metadata of the multimedia attachment file;

wherein the user interface transmits the generated text summary relating to the topic in response to the user request.

10. The system of claim 9 , wherein the user request takes a form of a natural language question, and wherein the operations further comprise:

identifying, by parsing the natural language question, the first user, the second user and the topic.

11. The system of claim 9 , wherein the user request further comprises one or more parameters entered via the user interface, and wherein the one or more parameters indicate any of:

a time range;

one or more conversation channels on the messaging platform; and

a file format.

12. The system of claim 9 , wherein the topic classification model is trained on a dataset of messages annotated with respective topic labels.

13. The system of claim 9 , wherein the multimedia attachment file includes any of a spreadsheet file, a presentation slide file, an audio file, a video file and an image file.

14. The system of claim 13 , wherein the operations further comprise:

converting the audio file into a text document; or

extracting one or more video frames from the video file.

15. A non-transitory processor-readable storage medium storing a plurality of processor-executed instructions for multi-modal summarization of communication on a messaging platform, the instructions being executed by one or more processors to perform operations comprising:

receiving, via a user interface, a user request for summarizing communication messages relating to a topic between a first user and a second user on a messaging platform;

searching, via a search engine, for messages between the first user and the second user on the messaging platform;

filtering the messages based on the topic from the user request by predicting, via a topic classification model, whether the messages are related to the topic and excluding a subset of messages that are predicted to be unrelated to the topic;

generating, by a text encoder of a multi-modal summarization model, a text representation from an input sequence corresponding to textual content from the filtered messages;

generating, by an image encoder of the multi-modal summarization model, an image representation of visual features in a multimedia attachment file from the filtered messages;

generating, via a decoder of the multi-modal summarization model, a text summary summarizing both the filtered messages and the multimedia attachment file based on a combination of the text representation generated by the text encoder and the image representation generated by the image encoder, wherein the text summary comprises a text that references a time and/or a sender of the multimedia attachment file generated from metadata of the multimedia attachment file; and

transmitting, via the user interface, the generated text summary relating to the topic in response to the user request.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2022
From: AGARWAL, DIVYANSH; WU, CHIEN-SHENG; XIE, TIAN
To: SALESFORCE, INC.
Reel/Frame 061811/0166 →
Continuity (1)
Related Publication 20240160837A1 · May 16, 2024
References Cited (16)
US 7299261B1 · Oliver · 2007 [cited by examiner]
US 11244755B1 · Syeda-Mahmood · 2022 [cited by examiner]
US 11809480B1 · Cheng · 2023 [cited by examiner]
US 20030004966A1 · Bolle · 2003 [cited by examiner]
US 20070244975A1 · Dillon · 2007 [cited by examiner]
US 20110320543A1 · Bendel · 2011 [cited by examiner]
US 20130166280A1 · Quast · 2013 [cited by examiner]
US 20140079197A1 · Hirschberg · 2014 [cited by examiner]
US 20160147387A1 · Rahman · 2016 [cited by examiner]
US 20180367483A1 · Rodriguez · 2018 [cited by examiner]
US 20210263959A1 · Bastide · 2021 [cited by examiner]
US 20210374552A1 · Mallya · 2021 [cited by examiner]
Cox et al., “Scanning the Technology: On the Applications of Multimedia Processing to Communications” Proceedings of the IEEE, copyright 1998 IEEE, 70 pages. (Year: 1998). [cited by examiner]
O'Day et al., “Text Message Corpus: Applying Natural Language Processing To Mobile Device Forensics” 2013 IEEE International Conference on Multimedia and Expo Workshops, 6 pages. (Year: 2013). [cited by examiner]
Feng et al., “Automatic Caption Generation for News Images,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, No. 4, pp. 797-812, Apr. 2013. (Year: 2013). [cited by examiner]
Ritter et al., “Toward Application Integration with Multimedia Data,” 2017 IEEE 21st International Enterprise Distributed Object Computing Conference (EDOC), Quebec City, QC, Canada, 2017, pp. 103-112. (Year: 2017). [cited by examiner]