IP Library › Granted Patent US 12,279,065
Granted Patent B2
US 12,279,065 · App. 17/725,468 · Granted Apr 15, 2025

Systems and methods for multi-user video communication with engagement detection and adjustable fidelity

Inventors: Dean N. Reading (Sunnyvale, CA); Marc Estruch Tena (San Jose, CA); Lin Sun (San Jose, CA); Fulya Yilmaz (San Francisco, CA); Cathy Kim (Palo Alto, CA); Hanna Fuhrmann (San Francisco, CA); Catherine S. Kim (San Jose, CA); Jun Yeon Cho (San Jose, CA); Imran Mohammed (Santa Clara, CA); Curtis D. Aumiller (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
H04N5/265G06T7/70G06V40/10H04L65/60H04N5/2622H04N5/2628H04N9/646G06T2207/10016G06T2207/30196G06T2207/30242
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,279,065
App. No.
17/725,468
Granted
Apr 15, 2025
Kind
B2
Abstract

In one embodiment, a method includes maintaining a video communication between client devices with each being associated with a respective video stream, which is associated with a respective match scale measured based on a height of frames of the video stream and a depth of subjects within the frames, determining a respective scaling factor and cropping for each video stream, wherein the respective scaling factor is calculated based on the match scale associated with the corresponding video stream and a target match scale determined from the match scales associated with all video streams, and wherein the respective cropping is determined based on a uniformity of positions of the subjects within the frames across all video streams, generating a merged video stream from each video stream based on the respective scaling factor and cropping, and sending instructions for presenting the merged video stream to one or more of the client devices.

Claims (57)

1. A method comprising, by an electronic device:

maintaining a video communication between two or more client devices, wherein each client device is associated with a respective video stream in the video communication, and wherein each video stream is associated with a respective match scale measured based on a height of frames of the video stream and a depth of subjects within the frames;

determining, for each of the video streams, a respective scaling factor and a respective cropping, wherein the respective scaling factor is calculated based on the match scale associated with the corresponding video stream and a target match scale determined from the match scales associated with the video streams associated with the two or more client devices, and wherein the respective cropping is determined based on a uniformity of positions of the subjects within the frames across the video streams associated with the two or more client devices;

generating, based on the respective scaling factor and cropping of each video stream, a merged video stream from each of the video streams for the video communication; and

sending, to one or more of the client devices, instructions for presenting the merged video stream.

2. The method of claim 1 , further comprising:

determining, for each of the video streams, a respective fidelity based on one or more of a date and time associated with the corresponding video stream, a calendar of a participant associated with the corresponding video stream, a command from the participant associated with the corresponding video stream, or a degree of engagement of the participant associated with the video stream, wherein generating the merged video stream is further based on the respective fidelity associated with each of the video streams.

3. The method of claim 1 , further comprising:

determining, for each of the video streams, a respective degree of engagement of a participant associated with the corresponding video stream based on one or more of the participant's presence, the participant's location, the participant's body direction, the participant's head direction, the participant's pose, or the participant's voice, wherein generating the merged video stream is further based on the respective degree of engagement of the participant associated with each of the video streams.

4. The method of claim 1 , further comprising:

determining, for each of the video streams, a respective adjustment of perspective associated with the corresponding video stream to center the subjects within the frames of the corresponding video stream, wherein generating the merged video stream is further based on the respective adjustment of perspective associated with each of the video streams.

5. The method of claim 1 , further comprising:

determining, for each of the video streams, a respective adjustment of color and lighting associated with the corresponding video stream based on a consistency of visual characteristics across the video streams, wherein generating the merged video stream is further based on the respective adjustment of color and lighting associated with each of the video streams.

6. The method of claim 1 , further comprising:

identifying borders between any two of the video streams, wherein generating the merged video stream comprises blurring or blending the identified borders.

7. The method of claim 1 , further comprising:

determining, for each of the video streams, a respective adjustment of space allocation associated with the corresponding video stream based on one or more of a count of participants present in the corresponding video stream, a degree of engagement of a participant in the corresponding video stream, or an activity level of a participant in the corresponding video stream, wherein generating the merged video stream is further based on the respective adjustment of space allocation associated with each of the video streams.

8. An electronic device comprising:

one or more displays;

one or more non-transitory computer-readable storage media including instructions; and

one or more processors coupled to the storage media, the one or more processors configured to execute the instructions to:

maintain a video communication between two or more client devices, wherein each client device is associated with a respective video stream in the video communication, and wherein each video stream is associated with a respective match scale measured based on a height of frames of the video stream and a depth of subjects within the frames;

determine, for each of the video streams, a respective scaling factor and a respective cropping, wherein the respective scaling factor is calculated based on the match scale associated with the corresponding video stream and a target match scale determined from the match scales associated with the video streams associated with the two or more client devices, and wherein the respective cropping is determined based on a uniformity of positions of the subjects within the frames across the video streams associated with the two or more client devices;

generate, based on the respective scaling factor and cropping of each video stream, a merged video stream from each of the video streams for the video communication; and

send, to one or more of the client devices, instructions for presenting the merged video stream.

9. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

determine, for each of the video streams, a respective fidelity based on one or more of a date and time associated with the corresponding video stream, a calendar of a participant associated with the corresponding video stream, a command from the participant associated with the corresponding video stream, or a degree of engagement of the participant associated with the video stream, wherein generating the merged video stream is further based on the respective fidelity associated with each of the video streams.

10. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

determine, for each of the video streams, a respective degree of engagement of a participant associated with the corresponding video stream based on one or more of the participant's presence, the participant's location, the participant's body direction, the participant's head direction, the participant's pose, or the participant's voice, wherein generating the merged video stream is further based on the respective degree of engagement of the participant associated with each of the video streams.

11. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

determine, for each of the video streams, a respective adjustment of perspective associated with the corresponding video stream to center the subjects within the frames of the corresponding video stream, wherein generating the merged video stream is further based on the respective adjustment of perspective associated with each of the video streams.

12. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

determine, for each of the video streams, a respective adjustment of color and lighting associated with the corresponding video stream based on a consistency of visual characteristics across the video streams, wherein generating the merged video stream is further based on the respective adjustment of color and lighting associated with each of the video streams.

13. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

identify borders between any two of the video streams, wherein generating the merged video stream comprises blurring or blending the identified borders.

14. The electronic device of claim 8 , wherein the processors are further configured to execute the instructions to:

determine, for each of the video streams, a respective adjustment of space allocation associated with the corresponding video stream based on one or more of a count of participants present in the corresponding video stream, a degree of engagement of a participant in the corresponding video stream, or an activity level of a participant in the corresponding video stream, wherein generating the merged video stream is further based on the respective adjustment of space allocation associated with each of the video streams.

15. A computer-readable non-transitory storage media comprising instructions executable by a processor to:

maintain a video communication between two or more client devices, wherein each client device is associated with a respective video stream in the video communication, and wherein each video stream is associated with a respective match scale measured based on a height of frames of the video stream and a depth of subjects within the frames;

determine, for each of the video streams, a respective scaling factor and a respective cropping, wherein the respective scaling factor is calculated based on the match scale associated with the corresponding video stream and a target match scale determined from the match scales associated with the video streams associated with the two or more client devices, and wherein the respective cropping is determined based on a uniformity of positions of the subjects within the frames across the video streams associated with the two or more client devices;

generate, based on the respective scaling factor and cropping of each video stream, a merged video stream from each of the video streams for the video communication; and

send, to one or more of the client devices, instructions for presenting the merged video stream.

16. The media of claim 15 , wherein the instructions are further executable by the processor to:

determine, for each of the video streams, a respective fidelity based on one or more of a date and time associated with the corresponding video stream, a calendar of a participant associated with the corresponding video stream, a command from the participant associated with the corresponding video stream, or a degree of engagement of the participant associated with the video stream, wherein generating the merged video stream is further based on the respective fidelity associated with each of the video streams.

17. The media of claim 15 , wherein the instructions are further executable by the processor to:

determine, for each of the video streams, a respective degree of engagement of a participant associated with the corresponding video stream based on one or more of the participant's presence, the participant's location, the participant's body direction, the participant's head direction, the participant's pose, or the participant's voice, wherein generating the merged video stream is further based on the respective degree of engagement of the participant associated with each of the video streams.

18. The media of claim 15 , wherein the instructions are further executable by the processor to:

determine, for each of the video streams, a respective adjustment of perspective associated with the corresponding video stream to center the subjects within the frames of the corresponding video stream, wherein generating the merged video stream is further based on the respective adjustment of perspective associated with each of the video streams.

19. The media of claim 15 , wherein the instructions are further executable by the processor to:

determine, for each of the video streams, a respective adjustment of color and lighting associated with the corresponding video stream based on a consistency of visual characteristics across the video streams, wherein generating the merged video stream is further based on the respective adjustment of color and lighting associated with each of the video streams.

20. The media of claim 15 , wherein the instructions are further executable by the processor to:

identify borders between any two of the video streams, wherein generating the merged video stream comprises blurring or blending the identified borders.

21. A method comprising, by an electronic device:

maintaining a video communication between two or more client devices, wherein each client device is associated with a respective video stream in the video communication; and

automatically adjusting, for each of the video streams, a respective amount of applied obfuscation to the video stream based on a degree of engagement, with the video communication, of a participant shown in the video stream, by (1) increasing the amount of applied obfuscation as the engagement of the participant with the video communication decreases and (2) decreasing the amount of applied obfuscation as the engagement of the participant with the video communication increases.

22. The method of claim 21 , further comprising:

determining, for each of the video streams, a respective degree of engagement of a participant associated with the corresponding video stream based on one or more of the participant's presence, the participant's location, the participant's body direction, the participant's head direction, the participant's pose, or the participant's voice.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2022
From: READING, DEAN N; ESTRUCH TENA, MARC; SUN, LIN; YILMAZ, FULYA; KIM, CATHY; FUHRMANN, HANNA; KIM, CATHERINE S; CHO, JUN YEON; MOHAMMED, IMRAN; AUMILLER, CURTIS D.
To: SAMSUNG ELECTRONICS COMPANY, LTD.
Reel/Frame 059656/0601 →
Continuity (1)
Related Publication 20230344956A1 · Oct 26, 2023
References Cited (31)
US 6771303B2 · Zhang · 2004 [cited by applicant]
US 8253852B2 · Bilbrey · 2012 [cited by applicant]
US 9088694B2 · Navon · 2015 [cited by applicant]
US 9167275B1 · Daily · 2015 [cited by examiner]
US 10075491B2 · Smus · 2018 [cited by applicant]
US 10362272B1 · van Os · 2019 [cited by applicant]
US 10462420B2 · Cranfill · 2019 [cited by applicant]
US 10721440B2 · Sugihara · 2020 [cited by applicant]
US 10917608B1 · Faulkner · 2021 [cited by applicant]
US 10951947B2 · Faulkner · 2021 [cited by applicant]
US 20110063440A1 · Neustaedter · 2011 [cited by examiner]
US 20120026277A1 · Malzbender · 2012 [cited by applicant]
US 20130194375A1 · Michrowski · 2013 [cited by applicant]
US 20180032224A1 · Cornell · 2018 [cited by applicant]
US 20180343294A1 · Rands · 2018 [cited by examiner]
US 20190026874A1 · Jin · 2019 [cited by applicant]
US 20190349512A1 · Bently · 2019 [cited by applicant]
US 20200098096A1 · Moloney · 2020 [cited by applicant]
US 20210081003A1 · Bristol · 2021 [cited by applicant]
US 20210120157A1 · Xu · 2021 [cited by applicant]
US 20210185208A1 · Djakovic · 2021 [cited by applicant]
US 20220116435A1 · Hartnett · 2022 [cited by examiner]
US 20220116546A1 · Chowdary · 2022 [cited by applicant]
JP 2020202567A · 2020 [cited by applicant]
KR 1020200108196A · 2020 [cited by applicant]
KR 1020210047112A · 2021 [cited by applicant]
PCT Search Report and written decision in PCT/KR2023/002595, May 24, 2023. [cited by applicant]
Non-final office action in U.S. Appl. No. 17/725,466, Feb. 27, 2024. [cited by applicant]
Final office action in U.S. Appl. No. 17/725,466, May 31, 2024. [cited by applicant]
Notice of Allowance in U.S. Appl. No. 17/725,466, Aug. 5, 2024. [cited by applicant]
Extended European Search Report in Application No. 23792004.6-1207 / 4413725 PCT/KR2023002595, Dec. 16, 2024. [cited by applicant]