IP Library Granted Patent US 12,659,429
Granted Patent B2
US 12,659,429 · App. 18/534,023 · Granted Jun 16, 2026

Video-conference device and method

Inventor: Lisa Rørbæk Kamstrup (Ballerup, DK)
Assignee: GN HEARING A/S
H04N7/152G06T3/40G06T5/70G06T7/50G06T7/70G06T11/60G06V10/761H04N7/157G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,659,429
App. No.
18/534,023
Granted
Jun 16, 2026
Kind
B2
Abstract

A video-conference device for providing an augmented view of in-room participants in a video-conference. The image sensor is configured to capture images comprising the room and the in-room participants. The device comprises a depth sensor configured for measuring the distance to each of the in-room participants having a processing unit configured for providing the augmented view of the in-room. The processing unit is configured to performing an image processing of the captured images from the image sensor to virtually identify each of the in-room participants. The processing unit is configured to determining a desired virtual position for each of the in-room participants in the augmented view. The processing unit is configured to performing a virtual scaling of each of the in-room participants to the desired virtual size.

Claims (89)

1 . A video-conference device for providing an augmented view of in-room participants in a video-conference, wherein the device is configured to be arranged in a room, a first number of participants includes a number of in-room participants present while the video-conference is held, and a second number of participants includes a number of far-end participants present at one or more different locations other than the room, wherein the device comprises:

an image sensor configured to capture images comprising the room and the in-room participants;

a depth sensor configured to measure a distance from the device to each of the in-room participants;

processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;

wherein the processing unit is configured to:

obtain the measurements of the distance from the device to each of the in-room participants by the depth sensor;

obtain the captured images by the images sensor;

perform an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;

determine a desired virtual position for each of the in-room participants in the augmented view;

determine a scaling factor for a desired virtual size of each of the in-room participants;

perform a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;

perform a virtual positioning of each of the in-room participants in the augmented view;

determine a distance A to each of the in-room participants thereby obtaining an actual position (x1,y1,z1) of each in-room participant;

determine a distance B between the in-room participants;

determine a relative location of the in-room participants based on the determined distance A and distance B;

determine the scaling factor N for the desired virtual size of each of the in-room participants based on the determined distance A and distance B; and

determine the virtual position (x2,y2,z2) of each of the in-room participants in the augmented view based on the determined relative location of the in-room participants.

2 . The device of claim 1 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the in-room participants are virtually positioned closer to each other and closer to the image sensor capturing the images.

3 . The device of claim 1 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the relative positions of the in-room participants relative to each other are maintained.

4 . The device of claim 1 , wherein the device is configured to be located at a non-central end of the room and the image sensor, the depth sensor and the processing unit of the device are also located in the room.

5 . The device of claim 1 , wherein the depth sensor comprises one or more of:

a time of flight (ToF) sensor configured for determining path lengths;

an infra-red sensor configured for determining distances based on reflected light; and

an acoustic sensor configured for determining distances based on reflected audio.

6 . The device according to claim 1 , wherein the image processing of the captured images to identify the in-room participants is performed by using one or more of:

virtual image segmentation;

virtual cut-out to virtually separate the in-room participants from the background;

virtual image resizing; and

virtual seam carving.

7 . The device of claim 1 , wherein the image processing is performed in two-dimensions (2D) and/or in three-dimensions (3D).

8 . The device to of claim 1 ,

wherein the depth sensor is configured to measure the distance to objects/parts of the room, and

wherein the image sensor is configured to capture images comprising the room with no in-room participants, and

wherein the processing unit is configured to:

map the room with no in-room participants, by determining distances C between the depth sensor and one or more objects/parts of the room.

9 . The device of claim 1 , wherein the processing unit is further configured to define a suitable illumination in the augmented view based on the time of day of the video-conference.

10 . The device of claim 1 , wherein the processing unit is further configured to perform blurring of objects in the augmented view of the room.

11 . A video-conference device for providing an augmented view of in-room participants in a video-conference, wherein the device is configured to be arranged in a room, a first number of participants includes a number of in-room participants present while the video-conference is held, and a second number of participants includes a number of far-end participants present at one or more different locations other than the room, wherein the device comprises:

an image sensor configured to capture images comprising the room and the in-room participants;

a depth sensor configured to measure a distance from the device to each of the in-room participants;

processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;

wherein the processing unit is configured to:

obtain the measurements of the distance from the device to each of the in-room participants by the depth sensor;

obtain the captured images by the images sensor;

perform an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;

determine a desired virtual position for each of the in-room participants in the augmented view;

determine a scaling factor for a desired virtual size of each of the in-room participants;

perform a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;

perform a virtual positioning of each of the in-room participants in the augmented view; and

in accordance with a determination that a criterion for the virtual positioning of an in-room participant is not satisfied:

forgo performing the virtual positioning of the in-room participant in the augmented view, and

virtually place the in-room participant in a dedicated picture/box in the augmented view.

12 . The device of claim 1 , wherein the device further comprises at least one of:

a display configured to display the augmented view and the far-end participants; or

an audio output transducer configured to transmit audio from the far-end participants to the room.

13 . A system comprising at least two video-conference devices according to claim 1 .

14 . A method, performed in a video-conference device, for providing an augmented view of in-room participants in the video-conference, wherein:

the device is configured to be arranged in a room,

a first number of participants is the in-room participants present in the room while the video-conference is held,

a second number of participants is far-end participants present at one or more different locations than the room,

the device comprises;

an image sensor configured to capture images comprising the room and the in-room participants;

a depth sensor configured to measure the distance to each of the in-room participants; and

a processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;

wherein the method comprises, by the processing unit:

obtaining the measurements of the distance to each of the in-room participants by the depth sensor;

obtaining the captured images by the images sensor;

performing an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;

determining a desired virtual position for each of the in-room participants in the augmented view;

determining a scaling factor for a desired virtual size of each of the in-room participants;

performing a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;

performing a virtual positioning of each of the in-room participants in the augmented view;

determining a distance A to each of the in-room participants thereby obtaining an actual position (x1,y1,z1) of each in-room participant;

determining a distance B between the in-room participants;

determining a relative location of the in-room participants based on the determined distance A and distance B;

determining the scaling factor N for the desired virtual size of each of the in-room participants based on the determined distance A and distance B; and

determining the virtual position (x2,y2,z2) of each of the in-room participants in the augmented view based on the determined relative location of the in-room participants.

15 . The device of claim 11 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the in-room participants are virtually positioned closer to each other and closer to the image sensor capturing the images.

16 . The device of claim 11 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the relative positions of the in-room participants relative to each other are maintained.

17 . The device of claim 11 , wherein the device is configured to be located at a non-central end of the room and the image sensor, the depth sensor and the processing unit of the device are also located in the room.

18 . The device of claim 11 , wherein the depth sensor comprises one or more of:

a time of flight (ToF) sensor configured for determining path lengths;

an infra-red sensor configured for determining distances based on reflected light; and

an acoustic sensor configured for determining distances based on reflected audio.

19 . The device according to claim 11 , wherein the image processing of the captured images to identify the in-room participants is performed by using one or more of:

virtual image segmentation;

virtual cut-out to virtually separate the in-room participants from the background;

virtual image resizing;

virtual seam carving.

Assignments (2)
MERGER Recorded Mar 30, 2026
From: GN AUDIO A/S
To: GN HEARING A/S
Reel/Frame 075299/0225 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2023
From: KAMSTRUP, LISA RØRBÆK KAMSTRUP
To: GN AUDIO A/S
Reel/Frame 065844/0856 →
Priority Claims (1)
EP 22215413 · Dec 21, 2022 · regional
Continuity (1)
Related Publication 20240214522A1 · Jun 27, 2024
References Cited (14)
US 9524588B2 · Barzuza · 2016 [cited by examiner]
US 9621795B1 · Whyte · 2017 [cited by examiner]
US 20110153738A1 · Bedingfield · 2011 [cited by examiner]
US 20130100235A1 · Vitsnudel · 2013 [cited by examiner]
US 20130182905A1 · Myers · 2013 [cited by examiner]
US 20160277712A1 · Michot · 2016 [cited by examiner]
US 20200294317A1 · Segal · 2020 [cited by applicant]
US 20230244799A1 · Sharma · 2023 [cited by examiner]
US 20230351915A1 · Chopdekar · 2023 [cited by examiner]
US 20240029472A1 · Shen · 2024 [cited by examiner]
EP 3975553A1 · 2022 [cited by applicant]
WO 2018005235A1 · 2018 [cited by applicant]
The extended European search report issued in European Application No. 22215413.0, dated Jun. 19, 2023, 12 pages provided. [cited by applicant]
Sae-Woon Ryu et al: “Tangible video teleconference system using real-time image-based relighting”, IEEE Transactions on Consumer Electronics, IEEE Service Center, New York, NY, US, vol. 55, No. 3, Aug. 1, 2009 (Aug. 1, … [cited by applicant]