IP Library › Granted Patent US 12,483,674
Granted Patent B2
US 12,483,674 · App. 18/198,687 · Granted Nov 25, 2025

Displaying video conference participants in alternative display orientation modes

Inventors: Ryan Fedyk (Brooklyn, NY); Stéphane Hervé Loïc Hulaud (Stockholm, SE)
Assignee: Google LLC
H04N7/152G06T7/20G06V40/171H04L65/403H04N7/147H04N7/15G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,483,674
App. No.
18/198,687
Granted
Nov 25, 2025
Kind
B2
Abstract

Systems and methods for determining whether to display video conference participants in first display mode are provided. A plurality of video streams from a plurality of client devices of a plurality of participants of a video conference are received. One or more visual features of one or more objects in each of the plurality of video streams is identified. Based on the identified one or more visual features, a determination is made whether to use a first display mode or a second display mode for one or more visual items of a plurality of visual items corresponding to the plurality of video streams in a rendered composition. The rendered composition of the plurality of visual items is caused to be displayed in a user interface of a client device of the plurality of client devices in accordance to the determined first display mode or second display mode.

Claims (62)

1 . A method comprising:

receiving a plurality of video streams from a plurality of client devices of a plurality of participants of a video conference;

identifying one or more visual features of one or more objects in each of the plurality of video streams;

determining, based on the identified one or more visual features, whether to use a first display mode of a plurality of display modes or a second display mode of the plurality of display modes for one or more visual items of a plurality of visual items corresponding to the plurality of video streams in a rendered composition, wherein the plurality of display modes comprises at least one of a portrait display mode or a landscape display mode; and

causing the rendered composition of the plurality of visual items to be displayed in a user interface of a client device of the plurality of client devices in accordance to the determined first display mode or second display mode of the plurality of display modes.

2 . The method of claim 1 , wherein the one or more objects comprise an image of a participant of the plurality of participants, and the one or more visual features comprise at least one of one or more body features or one or more facial features of the participant.

3 . The method of claim 2 , wherein the one or more facial features of the participant comprise at least one of an eyeline, a lower face region, a nose, or an upper face region of the participant.

4 . The method of claim 2 , further comprising:

determining, based on the identified one or more visual features, at least one of an estimated face size or an estimated head size of the participant.

5 . The method of claim 2 , further comprising:

identifying, over a period of time, movements of a first participant of the plurality of participants; and

responsive to determining that the movements satisfy a condition, removing a background of a first visual item corresponding to a first video stream comprising an image of the first participant prior to providing the rendered composition.

6 . The method of claim 2 , further comprising:

determining that a first video stream of the plurality of video streams comprises images of a subset of the plurality of participants; and

generating, for a first participant in the subset, based on one or more visual features associated with the first participant, an additional visual item comprising a cropped section of a first visual item corresponding to the first video stream in accordance to the determined first display mode or second display mode.

7 . The method of claim 1 , wherein determining whether to use the first display mode or the second display mode for the one or more visual items in the rendered composition is based on a set of rules or an output of a trained machine learning model.

8 . The method of claim 1 , further comprising:

cropping, based on the identified one or more visual features, the one or more visual items according to the determined first display mode or second display mode;

aligning, based on the identified one or more visual features, the one or more visual items; and

generating the rendered composition comprising the one or more visual items.

9 . The method of claim 1 , further comprising:

identifying a display size corresponding to the user interface;

identifying a set of layout templates corresponding to the display size;

selecting, based on the determination to use the first display mode or the second display mode for the one or more visual items in the rendered composition, a layout template of the set of layout templates; and

generating the rendered composition according to the selected layout template.

10 . A system comprising:

a memory device; and

a processing device coupled to the memory device, the processing device to perform operations comprising:

receiving a plurality of video streams from a plurality of client devices of a plurality of participants of a video conference;

identifying one or more visual features of one or more objects in each of the plurality of video streams;

determining, based on the identified one or more visual features, whether to use a first display mode of a plurality of display modes or a second display mode of the plurality of display modes for one or more visual items of a plurality of visual items corresponding to the plurality of video streams in a rendered composition, wherein the plurality of display modes comprises at least one of a portrait display mode or a landscape display mode; and

causing the rendered composition of the plurality of visual items to be displayed in a user interface of a client device of the plurality of client devices in accordance to the determined first display mode or second display mode of the plurality of display modes.

11 . The system of claim 10 , wherein the one or more objects comprise an image of a participant of the plurality of participants, and the one or more visual features comprise at least one of one or more body features or one or more facial features of the participant.

12 . The system of claim 11 , wherein the one or more facial features of the participant comprise at least one of an eyeline, a lower face region, a nose, or an upper face region of the participant, and wherein the processing device is to perform operations further comprising:

determining, based on the identified one or more visual features, at least one of an estimated face size or an estimated head size of the participant.

13 . The system of claim 10 , wherein determining whether to use the first display mode or the second display mode for the one or more visual items in the rendered composition is based on a set of rules or an output of a trained machine learning model.

14 . The system of claim 10 , wherein the processing device is to perform operations further comprising:

cropping, based on the identified one or more visual features, the one or more visual items according to the determined first display mode or second display mode;

aligning, based on the identified one or more visual features, the one or more visual items; and

generating the rendered composition comprising the one or more visual items.

15 . The system of claim 10 , wherein the processing device is to perform operations further comprising:

identifying a display size corresponding to the user interface;

identifying a set of layout templates corresponding to the display size;

selecting, based on the determination to use the first display mode or the second display mode for the one or more visual items in the rendered composition, a layout template of the set of layout templates; and

generating the rendered composition according to the selected layout template.

16 . A non-transitory computer readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform operations comprising:

receiving a plurality of video streams from a plurality of client devices of a plurality of participants of a video conference;

identifying one or more visual features of one or more objects in each of the plurality of video streams;

determining, based on the identified one or more visual features, whether to use a first display mode of a plurality of display modes or a second display mode of the plurality of display modes for one or more visual items of a plurality of visual items corresponding to the plurality of video streams in a rendered composition, wherein the plurality of display modes comprises at least one of a portrait display mode or a landscape display mode; and

causing the rendered composition of the plurality of visual items to be displayed in a user interface of a client device of the plurality of client devices in accordance to the determined first display mode or second display mode of the plurality of display modes.

17 . The non-transitory computer readable storage medium of claim 16 , wherein the one or more objects comprise an image of a participant of the plurality of participants, and the one or more visual features comprise at least one of one or more body features or one or more facial features of the participant; wherein the one or more facial features of the participant comprise at least one of an eyeline, a lower face region, a nose, or an upper face region of the participant; and wherein the processing device is to perform operations further comprising:

determining, based on the identified one or more visual features, at least one of an estimated face size or an estimated head size of the participant.

18 . The non-transitory computer readable storage medium of claim 16 , wherein determining whether to use the first display mode or the second display mode for the one or more visual items in the rendered composition is based on a set of rules or an output of a trained machine learning model.

19 . The non-transitory computer readable storage medium of claim 16 , wherein the processing device is to perform operations further comprising:

cropping, based on the identified one or more visual features, the one or more visual items according to the determined first display mode or second display mode;

aligning, based on the identified one or more visual features, the one or more visual items; and

generating the rendered composition comprising the one or more visual items.

20 . The non-transitory computer readable storage medium of claim 16 , wherein the processing device is to perform operations further comprising:

identifying a display size corresponding to the user interface;

identifying a set of layout templates corresponding to the display size;

selecting, based on the determination to use the first display mode or the second display mode for the one or more visual items in the rendered composition, a layout template of the set of layout templates; and

generating the rendered composition according to the selected layout template.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE MIDDLE NAME OF THE SECOND INVENTOR FROM STEPHANIE HULAUD TO STEPHAINE HERVE LOIC HULAUD PREVIOUSLY RECORDED AT REEL: 063816 FRAME: 0865. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Oct 5, 2023
From: FEDYK, RYAN; HULAUD, STÉPHANE HERVÉ LOÏC
To: GOOGLE LLC
Reel/Frame 065155/0508 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 31, 2023
From: FEDYK, RYAN; HULAUD, STÉPHANE
To: GOOGLE LLC
Reel/Frame 063816/0865 →
Continuity (1)
Related Publication 20240388675A1 · Nov 21, 2024
References Cited (10)
US 11082661B1 · Pollefeys · 2021 [cited by examiner]
US 20190065895A1 · Wang · 2019 [cited by examiner]
US 20190215464A1 · Kumar · 2019 [cited by examiner]
US 20200099889A1 · Sugihara · 2020 [cited by examiner]
US 20210203879A1 · Faulkner et al. · 2021 [cited by applicant]
US 20230121654A1 · Tangeland · 2023 [cited by examiner]
US 20240054786A1 · Andresen · 2024 [cited by examiner]
WO 2022078656A1 · 2022 [cited by applicant]
WO 2023064153A1 · 2023 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2024/029531, mailed Sep. 25, 2024, 22 Pages. [cited by applicant]