IP Library Patent Application 18116291
Patent Application
App. No. 18/116,291

METHOD AND APPARATUS FOR IDENTIFYING KEY INFORMATION IN A MULTI-PARTY MULTIMEDIA COMMUNICATION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/116,291
Abstract

An apparatus for identifying key information in a multi-party multimedia communication includes a processor, and a memory storing instructions that, when executed by the processor, configure the apparatus to perform a method. The method includes receiving multi-modal data including video data and audio data for each of multiple participants, and presentation data presented on one or more multimedia devices. Vision information and at least one of tonal information or text information is used to determine a representative participation score (RPS) for one or more participants. Content from presentation data presented during or proximate to pronounced RPS movement is identified and sent for display.

Claims (73)

1 . A computer implemented method for identifying key information in a multi-party multimedia communication, the method comprising:

receiving, at an analytics server, from a plurality of multimedia devices in a multimedia communication between a plurality of participants, multi-modal data comprising video data and audio data for at least one of the plurality of participants, and presentation data comprising data presented on at least one of the plurality of multimedia devices, wherein the plurality of multimedia devices are remote to the analytics server;

extracting, for each of the plurality of participants, vision information from the video data, at least one of tonal information or text information from the audio data;

determining a representative participation score (RPS) based on the vision information and the at least one of the tonal information or the text information, for each of the plurality of participants, for a plurality of consecutive time intervals;

determining a pronounced RPS movement

for at least one of the plurality of participants during a first plurality of time intervals,

for at least two of the plurality of participants during a second plurality of time intervals, or

for at least a predefined duration for at least one of the plurality of participants during a third plurality of time intervals;

identifying, from the presentation data, content relevant to the pronounced RPS movement, the content comprising at least one of a text element, an audio element, a graphic element, or a frame; and

sending the identified content for display.

2 . The method of claim 1 , wherein the identifying the content relevant to the pronounced RPS movement comprises identifying content presented

during the pronounced RPS movement, and

at least one of

a predefined duration or a predefined number of speaker turns before the pronounced RPS movement, or

at least one of a predefined duration or a predefined number of speaker turns after the pronounced RPS movement.

3 . The method of claim 2 , wherein the at least one of predefined duration is 5 s, and the at least one of predefined turns is 2 .

4 . The method of claim 2 , further comprising:

determining, from the identified content, a change from a first content to a second content, the second content presented at a time after the first content; and

sending the first content, the second content, and the determined change for display.

5 . The method of claim 2 , further comprising:

identifying a hyper-relevant text keyphrase in the identified content; and

sending the identified hyper-relevant text keyphrase for display.

6 . The method of claim 1 , wherein the identifying content comprises optical character recognition (OCR).

7 . The method of claim 1 , wherein the determining the change comprises comparing consecutive frames.

8 . A computing apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, configure the apparatus to:

receive, at an analytics server, from a plurality of multimedia devices in a multimedia communication between a plurality of participants, multi-modal data comprising video data and audio data for at least one of the plurality of participants, and presentation data comprising data presented on at least one of the plurality of multimedia devices, wherein the plurality of multimedia devices are remote to the analytics server;

extract, for each of the plurality of participants, vision information from the video data, at least one of tonal information or text information from the audio data;

determine a representative participation score (RPS) based on the vision information and the at least one of the tonal information or the text information, for each of the plurality of participants, for a plurality of consecutive time intervals;

determine a pronounced RPS movement

for at least one of the plurality of participants during a first plurality of time intervals,

for at least two of the plurality of participants during a second plurality of time intervals, or

for at least a predefined duration for at least one of the plurality of participants during a third plurality of time intervals;

identify, from the presentation data, content relevant to the pronounced RPS movement, the content comprising at least one of a text element, an audio element, a graphic element, or a frame; and

send the identified content for display.

9 . The computing apparatus of claim 8 , wherein the identifying the content relevant to the pronounced RPS movement comprises identifying content presented

during the pronounced RPS movement, and

at least one of

a predefined duration or a predefined number of speaker turns before the pronounced RPS movement, or

at least one of a predefined duration or a predefined number of speaker turns after the pronounced RPS movement.

10 . The computing apparatus of claim 9 , wherein the at least one of predefined duration is 5 s, and the at least one of predefined turns is 2.

11 . The computing apparatus of claim 9 , wherein the instructions further configure the apparatus to:

determine, from the identified content, a change from a first content to a second content, the second content presented at a time after the first content; and

send the first content, the second content, and the determined change for display.

12 . The computing apparatus of claim 9 , wherein the instructions further configure the apparatus to:

identify a hyper-relevant text keyphrase in the identified content; and

send the identified hyper-relevant text keyphrase for display.

13 . The computing apparatus of claim 8 , wherein the identifying content comprises optical character recognition (OCR).

14 . The computing apparatus of claim 8 , wherein the determining the change comprises comparing consecutive frames.

15 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to:

receive, at an analytics server, from a plurality of multimedia devices in a multimedia communication between a plurality of participants, multi-modal data comprising video data and audio data for at least one of the plurality of participants, and presentation data comprising data presented on at least one of the plurality of multimedia devices, wherein the plurality of multimedia devices are remote to the analytics server;

extract, for each of the plurality of participants, vision information from the video data, at least one of tonal information or text information from the audio data;

determine a representative participation score (RPS) based on the vision information and the at least one of the tonal information or the text information, for each of the plurality of participants, for a plurality of consecutive time intervals;

determine a pronounced RPS movement

for at least one of the plurality of participants during a first plurality of time intervals,

for at least two of the plurality of participants during a second plurality of time intervals, or

for at least a predefined duration for at least one of the plurality of participants during a third plurality of time intervals;

identify, from the presentation data, content relevant to the pronounced RPS movement, the content comprising at least one of a text element, an audio element, a graphic element, or a frame; and

send the identified content for display.

16 . The computer-readable storage medium of claim 15 , wherein the identifying the content relevant to the pronounced RPS movement comprises identifying content presented

during the pronounced RPS movement, and

at least one of

a predefined duration or a predefined number of speaker turns before the pronounced RPS movement, or

at least one of a predefined duration or a predefined number of speaker turns after the pronounced RPS movement.

17 . The computer-readable storage medium of claim 16 , wherein the instructions further configure the computer to:

determine, from the identified content, a change from a first content to a second content, the second content presented at a time after the first content; and

send the first content, the second content, and the determined change for display.

18 . The computer-readable storage medium of claim 16 , wherein the instructions further configure the computer to:

identify a hyper-relevant text keyphrase in the identified content; and

send the identified hyper-relevant text keyphrase for display.

19 . The computer-readable storage medium of claim 15 , wherein the identifying content comprises optical character recognition (OCR).

20 . The computer-readable storage medium of claim 15 , wherein the determining the change comprises comparing consecutive frames.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Oct 2, 2025
From: SILICON VALLEY BANK, A DIVISION OF FIRST-CITIZENS BANK & TRUST COMPANY
To: UNIPHORE TECHNOLOGIES INC.
Reel/Frame 072454/0763 →
SECURITY INTEREST Recorded Dec 24, 2024
From: UNIPHORE TECHNOLOGIES INC.
To: FIRST-CITIZENS BANK & TRUST COMPANY
Reel/Frame 069674/0415 →
SECURITY INTEREST Recorded Aug 20, 2024
From: UNIPHORE TECHNOLOGIES INC.; UNIPHORE TECHNOLOGIES NORTH AMERICA INC.; UNIPHORE SOFTWARE SYSTEMS INC.; COLABO, INC.
To: HSBC VENTURES USA INC.
Reel/Frame 068335/0563 →