IP Library Granted Patent US 12,634,579
Granted Patent B2
US 12,634,579 · App. 18/464,039 · Granted May 19, 2026

Autonomous video conferencing system with virtual director assistance

Inventors: Jon Tore Hafstad (Oslo, NO); Aida C. Lopez (Oslo, NO); Elena You (Oslo, NO); Kai Alexander Wig (Oslo, NO); Lars Erling Stensen (Oslo, NO); Mona Kleven Lauritzen (Oslo, NO); Stian Selbek (Koppang, NO); Tamás Becsei (Oslo, NO); Niklas Schmidt (Oslo, NO); Therese Byhring (Oslo, NO); Vebjørn Boge Nilssen (Oslo, NO); Patrik Kvarme Hansen (Oslo, NO); Knut Helge Teppan (Asker, NO); Stein Ove Erikson (Oslo, NO); Vegard Hammer (Tåmåsen, NO); Bendik Kvamstad (Oslo, NO); Jan Tore Korneliussen (Oslo, NO); Håvard Pederson Alstad (Oslo, NO); Oleg Jakobsen (Vinterbro, NO)
Assignee: HUDDLY AS
H04N23/67G06N3/04H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,579
App. No.
18/464,039
Granted
May 19, 2026
Kind
B2
Abstract

Systems and methods are provided to power video conferencing and remote collaboration with subsymbolic and symbolic artificial intelligence. The autonomous video conferencing systems of this disclosure include one main smart camera and multiple peripheral smart cameras, optionally coupled with one or more smart sensors. Each smart camera is equipped with a vision pipeline supported by machine learning to detect objects and their interactions as well as related changes in gesture and posture, and a virtual director adapted to apply a predetermined rule set consistent with television studio production principles. The main camera is adapted to select and update a focus video stream in real time under the direction of its virtual director and stream the updated focus stream to a user computer. Methods for creating an automated television studio production for a variety of conferencing spaces and special-purpose scenarios with virtual director assistance are provided.

Claims (57)

1 . A multi-camera videoconferencing system, comprising:

two or more cameras, wherein each camera of the multi-camera videoconferencing system is located in different areas of a video conferencing space and includes an image sensor configured to capture an overview video stream; and

at least one video processing unit configured to:

cause an overview video stream from a first camera to be shown on a display during a first time period; and

cause a focus video stream, derived from an overview video stream associated with a second camera, to be shown on the display during a second time period, wherein a transition is applied between the display of the overview video stream from the first camera and the display of the focus video stream;

wherein the focus video stream derived from the overview video stream associated with the second camera features an object that is also represented in the overview video stream of the first camera, and

wherein the first time period, the transition, and the second time period are determined based on a production scenario corresponding to the video conferencing space.

2 . The multi-camera videoconferencing system of claim 1 , wherein the at least one video processing unit is further configured to cause the transition between the overview video stream from the first camera and the focus video stream.

3 . The multi-camera videoconferencing system of claim 2 , wherein the transition is triggered by a detection that the object has started to speak.

4 . The multi-camera videoconferencing system of claim 2 , wherein the transition is triggered by a detection of movement associated with the object.

5 . The multi-camera videoconferencing system of claim 2 , wherein the transition is triggered by a detection of a change in a gaze direction associated with the object.

6 . The multi-camera videoconferencing system of claim 2 , wherein the transition is triggered by a detection of a reaction made by the object.

7 . The multi-camera videoconferencing system of claim 2 , wherein the at least one video processing unit is further configured to cause a secondary transition between the focus video stream and a listener video stream, wherein the listener video stream frames a non-speaking object, and wherein the secondary transition occurs after a predefined length of time.

8 . The multi-camera videoconferencing system of claim 7 , wherein the predefined length of time is associated with a time duration over which the object featured in the focus video stream is determined to be speaking.

9 . The multi-camera videoconferencing system of claim 1 , wherein the focus video stream frames a subject determined to be speaking.

10 . The multi-camera videoconferencing system of claim 1 , wherein the object is a meeting participant.

11 . The multi-camera videoconferencing system of claim 1 , wherein the focus video stream includes an interest shot, a listening shot, or a presenter shot.

12 . The multi-camera videoconferencing system of claim 1 , wherein production scenario is a classroom, a workshop, or a meeting room.

13 . The multi-camera videoconferencing system of claim 1 , wherein the first camera is different from the second camera.

14 . The multi-camera videoconferencing system of claim 1 , wherein the first camera is the same as the second camera.

15 . A multi video conferencing camera system, comprising:

a plurality of image sensors located in different areas of a video conferencing space, wherein at least one of the plurality of image sensors is configured to capture an overview video stream; and

a video processing unit configured to:

receive input from a user indicative of a selection by the user of a video stream mode from among two or more video stream modes, wherein each of the two or more video stream modes is associated with a unique video framing scheme for automatically transitioning from an overview video stream to a focus video stream and for automatically framing the focus video stream, wherein the unique video framing scheme is associated with a production scenario corresponding to the video conferencing space;

generate the focus video stream from the overview video stream in accordance with the user's video stream mode selection; and

in accordance with the user's video stream mode selection, cause the overview video stream to be shown on a display for a first time period and cause the focus video stream to be shown on the display for a second time period following a transition from the overview video stream to the focus video stream, wherein the first time period, transition, and second time period are determined based on the production scenario.

16 . The camera system of claim 15 , wherein the two or more video stream modes include two or more of a classroom mode, a workshop mode, a meeting room mode, a presentation mode, a courtroom mode, a news broadcast mode, or a panel discussion mode.

17 . The camera system of claim 15 , wherein the focus video stream includes an interest shot, a listening shot, or a presenter shot.

18 . The camera system of claim 15 , wherein the selection by the user may include a modification to a rule or at least one parameter associated with a rule.

19 . A multi-camera videoconferencing system, comprising:

at least one sensor configured to monitor one or more aspects of a videoconference environment;

two or more cameras located in different areas of the videoconference environment, wherein each camera of the multi-camera videoconferencing system includes an image sensor configured to capture an overview video stream; and

at least one video processing unit configured to:

cause an overview video stream from a first camera to be shown on a display during a first time period; and

cause a focus video stream, derived from an overview video stream associated with a second camera, to be shown on the display during a second time period, wherein a transition is applied between the display of the overview video stream from the first camera and the display of the focus video stream;

wherein the focus video stream is derived from the overview video stream based on an output of the at least one sensor and based upon a frame transition trigger detected based on analysis of at least one overview video stream of the two or more cameras, and

wherein the first time period, the transition, the second time period, and the frame transition trigger are determined based on a production scenario corresponding to the videoconference environment.

20 . The multi-camera videoconferencing system of claim 19 , wherein the at least one sensor includes one or more of a touchpad, a microphone, a smartphone, a GPS tracker, an echolocation sensor, a thermometer, a humidity sensor, and a biometric sensor.

21 . The multi-camera videoconferencing system of claim 19 , wherein the at least one sensor and the two or more cameras are connected via one or more of an Ethernet connection, a local area network, or a wireless network.

22 . The multi-camera videoconferencing system of claim 19 , wherein the at least one video processing unit is further configured to cause the transition between the overview video stream displayed during the first time period and the focus video stream displayed during the second time period.

23 . The multi-camera videoconferencing system of claim 22 , wherein the transition occurs after the first time period at least equals a predefined length of time.

24 . The multi-camera videoconferencing system of claim 19 , wherein the focus video stream frames an object detected to be speaking.

25 . The multi-camera videoconferencing system of claim 19 , wherein the at least one video processing unit is further configured to:

aggregate and process audio signals received from at least one audio device to identify subjects associated with two or more voices; and

generate the focus video stream based on detection of one of the two or more voices in the audio signals received from the at least one audio device and a correlation of the identified subjects associated with the two or more voices.

26 . The multi-camera videoconferencing system of claim 19 , wherein the at least one video processing unit is further configured to:

process audio signals received from at least one audio device;

determine whether a detected voice represented by the audio signals is associated with at least one subject present in the videoconference environment; and

generate the focus video stream based at least in part on the detected voice.

27 . The multi-camera videoconferencing system of claim 19 , wherein the at least one video processing unit is further configured to:

process audio signals received from at least one audio device;

determine whether a detected voice represented by the audio signals is not associated with at least one subject present in the videoconference environment; and

forego generation of the focus video stream based on the detected voice where the at least one video processing unit determines that the detected voice represented by the audio signals is not associated with at least one subject present in the videoconference environment.

28 . The multi-camera videoconferencing system of claim 27 , wherein the at least one audio device includes a microphone array, and the at least one video processing unit is configured to determine a direction of audio (DOA) associated with the detected voice, and wherein the determination of whether the detected voice represented by the audio signals is not associated with at least one subject present in the videoconference environment is based on the DOA associated with the detected voice.

29 . The multi-camera videoconferencing system of claim 27 , wherein the at least one video processing unit is configured to run one or more trained neural networks.

30 . The multi-camera videoconferencing system of claim 19 , wherein the first camera is different from the second camera.

31 . The multi-camera videoconferencing system of claim 19 , wherein the first camera is the same as the second camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2024
From: HAFSTAD, JON TORE; LOPEZ, AIDA C; YOU, ELENA; STENSEN, LARS ERLING; WIG, KAI ALEXANDER; LAURITZEN, MONA KLEVEN; SELBEK, STIAN; BECSEI, TAMAS; SCHMIDT, NIKLAS; BYHRING, THERESE; NILSSEN, VEBJORN BOGE; HANSEN, PATRIK KVARME; TEPPAN, KNUT HELGE; ERIKSEN, STEIN OVE; HAMMER, VEGARD; KVAMSTAD, BENDIK; KORNELIUSSEN, JAN TORE; ALSTAD, HAVARD PEDERSEN; JAKOBSEN, OLEG
To: HUDDLY AS
Reel/Frame 066459/0568 →
Continuity (2)
Continuation 17678954 · Feb 23, 2022
Related Publication 20230421899A1 · Dec 28, 2023
References Cited (13)
US 9270941B1 · Lavelle · 2016 [cited by examiner]
US 10659731B2 · Harrison · 2020 [cited by examiner]
US 20060005136A1 · Wallick · 2006 [cited by examiner]
US 20070279494A1 · Aman · 2007 [cited by examiner]
US 20180167584A1 · Nimri · 2018 [cited by examiner]
US 20190313058A1 · Harrison · 2019 [cited by examiner]
US 20200267427A1 · Rogers et al. · 2020 [cited by applicant]
US 20210019982A1 · Todd · 2021 [cited by examiner]
WO WO2020208038 · 2020 [cited by applicant]
WO WO2020208038A1 · 2020 [cited by examiner]
International Search Report and Written Opinion in PCT Application No. PCT/US2022/025427 dated Jun. 29, 2022 (25 pages). [cited by applicant]
Afshari et al., “Short Paper: The Smart Conference Room: An Integrated System Tested for Efficient, Occupancy-Aware Lighting Control,” ACM, Nov. 5, 2015, retrieved on [Sep. 6, 2022]. Retrieved from the Internet <URL: ht… [cited by applicant]
Extended European Search Report dated Apr. 11, 2025, for corresponding European Patent Application No. 22930097.5 (10 pgs). [cited by applicant]