IP Library Patent Application 18661087
Patent Application
App. No. 18/661,087

METHOD AND SYSTEM FOR AUTOMATIC SPEAKER FRAMING IN VIDEO APPLICATIONS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/661,087
Abstract

For video applications, a method for dynamically switching from a current ROI to a target ROI is disclosed, wherein the target ROI only includes active speakers. Advantageously, if the target ROI crops a non-speaker, then the target ROI is expanded to include said non-speaker. Transitioning from the current ROI to the target ROI may be achieved based on a cutover transition technique, or a smooth transition technique. The cutover transition technique achieves the change from the current arrived to the target ROI in a single interval, whereas the smooth transition technique achieves the change over a number of intervals, wherein a percentage of the total change required is allocated to each interval. A system for implementing the above method is also disclosed.

Claims (42)

1 . A method for framing video, comprising:

associating a current region-of-interest (ROI) corresponding to a video of a scene being imaged;

determining active speakers in the imaged scene;

performing a reframing operation comprising dynamically calculating a target region of interest (ROI), wherein

the target ROI is calculated based on coordinates associated with calculated bounding box for the determined active speakers in the imaged scene, and

the target ROI is expanded to include a non-active speaker when the target ROI crops the non-active speaker; and

transitioning to the target ROI from the current ROI using one of a cutover transition technique and a smooth transition technique.

2 . The method of claim 1 , further comprising calculating a degree of overlap between the current ROI and the target ROI.

3 . The method of claim 2 , wherein the calculating of the degree of overlap between the current ROI and the target ROI comprises determining an intersection over union (IoU) between the current ROI and the target ROI.

4 . The method of claim 2 , wherein the transitioning comprises executing the cutover transition technique if the degree of overlap between the current ROI and the target ROI is greater than a threshold and executing the smooth transition technique if the degree of overlap is below the threshold.

5 . The method of claim 4 , wherein the smooth transition technique includes generation of a new target ROI based on intermediate frames between the current ROI and the target ROI.

6 . The method of claim 4 , wherein the smooth transition technique comprises:

calculating change required for the transition from the current ROI to the target ROI;

determining a number of intervals over which the current ROI is transited to the target ROI; and

allocating the calculated change to each interval of the number of intervals.

7 . The method of claim 6 , wherein the allocation of the calculated change across intervals is non-linear.

8 . The method of claim 6 , wherein the cutover transition technique comprises allocating the calculated change to a single interval of the number of intervals.

9 . The method of claim 1 , wherein bounding box is calculated based on convolutional neural network (CNN) features.

10 . The method of claim 1 , wherein the calculated bounding box for the determined active speakers is a face bounding box and a body bounding box.

11 . The method of claim 1 , wherein the target ROI is further calculated based on a size of the calculated bounding box and movements of the calculated bounding box.

12 . A system, comprising:

a camera configured to capture a video of a scene; and

a virtual director module configured to:

associate a current region-of-interest (ROI) corresponding to the captured video of the scene;

determine active speakers in the scene;

perform a reframing operation comprising dynamically calculating a target region of interest (ROI), wherein

the target ROI is calculated based on coordinates associated with calculated bounding box for the determined active speakers in the scene, and

the target ROI is expanded to include a non-active speaker when the target ROI crops the non-active speaker; and

transit the target ROI from the current ROI based on one of a cutover transition technique and a smooth transition technique.

13 . The system of claim 12 , wherein the virtual director module is further configured to calculate a degree of overlap between the current ROI and the target ROI.

14 . The system of claim 13 , wherein

the virtual director module is further configured to determine an intersection over union (IoU) between the current ROI and the target ROI, and

the degree of overlap between the current ROI of the target ROI is calculated based on the determination of the IoU between the current ROI and the target ROI.

15 . The system of claim 13 , wherein the transition comprises execution of the cutover transition technique if the degree of overlap between the current ROI and the target ROI is greater than a threshold and execution of the smooth transition technique if said degree of overlap is below said threshold.

16 . The system of claim 15 , wherein the smooth transition technique comprising:

calculating change required for the transition from the current ROI to the target ROI;

determining a number of intervals over which the current ROI is transited to the target ROI; and

allocating the calculated change to each interval of the number of intervals.

17 . The system of claim 16 , wherein the allocation of the calculated change across intervals is non-linear.

18 . The system of claim 16 , wherein the cutover transition technique includes allocation of the calculated change to a single interval of the number of intervals.

19 . The system of claim 12 , wherein bounding box is calculated based on convolutional neural network (CNN) features.

20 . The system of claim 12 , wherein the target ROI is further calculated based on a size of the calculated bounding box and movements of the calculated bounding box.

Assignments (1)
MERGER Recorded Mar 30, 2026
From: GN AUDIO A/S
To: GN HEARING A/S
Reel/Frame 075299/0225 →