IP Library › Granted Patent US 12,500,994
Granted Patent B2
US 12,500,994 · App. 18/536,059 · Granted Dec 16, 2025

Method for recording video conference and video conferencing system

Inventors: Shuo-Yu Wang (Taichung, TW); Chan-An Wang (Taichung, TW); Hsiu-Ling Lin (Taichung, TW); Chung-Yi Huang (Taichung, TW)
Assignee: Merry Electronics(Shenzhen) Co., Ltd.
H04N5/91G06F3/0484G06T7/70G06V40/10G10L17/02G10L17/14G11B27/34G06T2207/10016G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,500,994
App. No.
18/536,059
Granted
Dec 16, 2025
Kind
B2
Abstract

A method for recording a video conference and a video conferencing system are provided. The method includes: providing a user interface to a display device, in which the user interface includes a first area, a second area, and a timeline; in response to obtaining an image corresponding to each of multiple participants from a video signal through a person recognition algorithm, displaying the image of each participant in the first area; in response to converting an audio segment of one of the participants obtained from an audio signal into text content through a voice processing algorithm, associating the text content with the corresponding one of the participants, and based on an order of speaking, displaying the text content in the second area; and adjusting a time length of the timeline according to a recording time of the video conference.

Claims (47)

1 . A method for recording a video conference, performed by a processor in response to the video conference being started, the method comprising:

providing a user interface to a display device, wherein the user interface comprises a first area, a second area, and a timeline;

in response to obtaining an image corresponding to each of a plurality of participants from a video signal through a person recognition algorithm, displaying the image of each of the participants in the first area;

in response to converting an audio segment of one of the participants obtained from an audio signal into text content through a voice processing algorithm, associating the text content with the corresponding one of the participants, and displaying the text content in the second area based on an order of speaking; and

adjusting a time length of the timeline according to a recording time of the video conference,

wherein the method performed by the processor in response to the video conference being started further comprises:

identifying each of the participants included in the video signal through the person recognition algorithm, and obtaining a relative position of each of the participants in a conference space;

extracting a voice from the audio signal through a voiceprint recognition module;

determining a source position of the voice in the conference space through a sound source positioning algorithm;

matching the voice with the corresponding one of the participants based on the relative position and the source position; and

converting the audio segment corresponding to the voice into the text content through the voice processing algorithm.

2 . The method for recording the video conference according to claim 1 , wherein the first area provides an editing function, and after displaying the image of each of the participants in the first area, the method further comprises:

renaming a name corresponding to the image through the editing function.

3 . The method for recording the video conference according to claim 1 , wherein displaying the text content in the second area comprises:

extracting from the first area the image of one of the participants that matches the voice and a name corresponding thereto after converting the audio segment corresponding to the voice into the text content through the voice processing algorithm; and

displaying the image, the name, the text content, and a reception time of the audio segment in the second area.

4 . The method for recording the video conference according to claim 3 , wherein after converting the audio segment corresponding to the voice into the text content through the voice processing algorithm, the method further comprises:

associating a time section corresponding to the audio segment on the timeline with the text content.

5 . The method for recording the video conference according to claim 1 , wherein the user interface provides a marking function, and the method further comprises:

based on a time point when the marking function is enabled, putting a focus mark on the text content corresponding to the time point in the second area.

6 . The method for recording the video conference according to claim 1 , wherein the text content presented in the second area has a playback function, and the method further comprises:

playing the audio segment corresponding to the text content in response to the playback function being enabled.

7 . A video conferencing system, comprising:

a display device;

a storage comprising an application program; and

a processor coupled to the display device and the storage, and configured to execute the application program to start a video conference, and in response to the video conference being started, the processor being configured to:

provide a user interface to the display device, wherein the user interface comprises a first area, a second area, and a timeline;

in response to obtaining an image corresponding to each of a plurality of participants from a video signal through a person recognition algorithm, display the image of each of the participants in the first area;

in response to converting an audio segment of one of the participants obtained from an audio signal into text content through a voice processing algorithm, associate the text content with the corresponding one of the participants, and display the text content in the second area based on an order of speaking; and

adjust a time length of the timeline according to a recording time of the video conference,

wherein the storage further comprises a person recognition module, a voiceprint recognition module, and a voice processing module, and the processor is configured to:

identify each of the participants included in the video signal by executing the person recognition algorithm through the person recognition module, and obtain a relative position of each of the participants in a conference space;

extract a voice from the audio signal through the voiceprint recognition module;

determine a source position of the voice in the conference space by executing a sound source positioning algorithm through the voiceprint recognition module;

match the voice with the corresponding one of the participants based on the relative position and the source position; and

convert the audio segment corresponding to the voice into the text content by executing the voice processing algorithm through the voice processing module.

8 . The video conferencing system according to claim 7 , wherein the first area provides an editing function, and the processor is configured to:

rename a name corresponding to the image through the editing function after displaying the image of each of the participants in the first area.

9 . The video conferencing system according to claim 7 , wherein the processor is configured to:

extract from the first area the image of one of the participants that matches the voice and a name corresponding thereto after converting the audio segment corresponding to the voice into the text content through the voice processing module; and

display the name, the text content, and a reception time of the audio segment in the second area.

10 . The video conferencing system according to claim 9 , wherein the processor is configured to:

associate a time section corresponding to the audio segment on the timeline with the text content after converting the audio segment corresponding to the voice into the text content through the voice processing module.

11 . The video conferencing system according to claim 7 , wherein the user interface provides a marking function, and the processor is configured to:

based on a time point when the marking function is enabled, put a focus mark on the text content corresponding to the time point in the second area.

12 . The video conferencing system according to claim 7 , wherein the text content presented in the second area has a playback function, and the processor is configured to:

play the audio segment corresponding to the text content in response to the playback function being enabled.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2023
From: WANG, SHUO-YU; WANG, CHAN-AN; LIN, HSIU-LING; HUANG, CHUNG-YI
To: MERRY ELECTRONICS(SHENZHEN) CO., LTD.
Reel/Frame 065892/0388 →
Priority Claims (1)
TW 112143049 · Nov 8, 2023 · national
Continuity (1)
Related Publication 20250150550A1 · May 8, 2025
References Cited (13)
US 8351581B2 · Mikan · 2013 [cited by examiner]
US 9443518B1 · Gauci · 2016 [cited by examiner]
US 10798341B1 · Hegde · 2020 [cited by examiner]
US 10965909B2 · Tanaka · 2021 [cited by examiner]
US 11315569B1 · Talieh · 2022 [cited by examiner]
US 20120053936A1 · Marvit · 2012 [cited by examiner]
US 20220254348A1 · Tay · 2022 [cited by examiner]
US 20230208664A1 · Mese · 2023 [cited by examiner]
US 20230419966A1 · Roper · 2023 [cited by examiner]
US 20240163390A1 · Hutto · 2024 [cited by examiner]
US 20250005289A1 · Wang · 2025 [cited by examiner]
US 20250039335A1 · Jain · 2025 [cited by examiner]
US 20250140246A1 · Lee · 2025 [cited by examiner]