IP Library Granted Patent US 12,738,093
Granted Patent B2
US 12,738,093 · App. 17/870,618 · Granted Sep 15, 2026

Depth assisted images refinement

Inventors: Yezhi Shen (West Lafayette, IN); Weichen Xu (West Lafayette, IN); Qian Lin (Palo Alto, CA); Jan P. Allebach (West Lafayette, IN); Fengqing Zhu (West Lafayette, IN)
Assignees: Hewlett-Packard Development Company, L.P.; Purdue Research Foundation
G06V40/172G06T7/194G06T7/50G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,738,093
App. No.
17/870,618
Granted
Sep 15, 2026
Kind
B2
Abstract

In some examples in accordance with the present description, an electronic device is provided. The electronic device includes a controller to implement an image segmentation process. The controller is to obtain color information of an image. The controller also is to obtain depth information of the image. The controller also is to determine a depth of a face represented in the color information. The controller also is to segment a foreground of the image from a background of the image according to the color information and the depth information based on the depth of the face.

Claims (51)

1 . An electronic device, comprising:

an image sensor;

a depth sensor; and

a controller to:

receive, from the image sensor, red-green-blue (RGB) information of an image in separate R, G, and B channels;

receive, from the depth sensor, depth information of the image;

perform facial detection to identify a face in the RGB information;

truncate the depth information to exclude information for depth points not within a threshold distance from a depth of the identified face;

process the R, G, and B channels of the image and the truncated depth information according to a trained machine learning process to separate a foreground of the image from a background of the image, the trained machine learning process to receive the R, G, and B channels of the image and the truncated depth information as input and to output a segmentation result;

form an image mask by processing the RGB information and the truncated depth information according to the machine learning process;

apply the image mask to the image to obtain a separate representation of a foreground;

manipulate a background of the image by at least one of blurring or replacing the background;

overlay the separate foreground representation on the manipulated background to provide video data; and

transmit, via a network interface, the video data to another electronic device participating in a video conferencing session.

2 . The electronic device of claim 1 , wherein the controller is to truncate the depth information according to Euclidean distance clustering.

3 . The electronic device of claim 1 , wherein the controller is to sample the depth information within a bounding box that bounds the identified face to determine the depth of the identified face.

4 . The electronic device of claim 1 , wherein the controller is to process the R, G, and B channels of the image and the truncated depth information according to a convolutional neural network.

5 . An electronic device, comprising:

a controller to implement an image segmentation process to:

obtain color information of an image in separate color channels;

obtain depth information of the image;

determine a depth of a face represented in the color information;

provide the color information of the image in the separate color channels and the depth of the face to a trained machine learning process to separate a foreground of the image from a background of the image, the trained machine learning process to receive the color channels of the image and the truncated depth information as input and to output a segmentation result,

perform a depth cutoff of points of the depth information and separate the foreground of the image from the background of the image by processing the image and the cutoff depth information according to the trained machine learning process to form an image mask;

concatenate the separate color channels with the truncated depth information to form a four-channel input to the trained machine learning process;

process the four-channel input to generate the segmentation result; and

generate an updated face bounding box based on the segmentation result, the updated face bounding box for use in performing depth sampling for a subsequently received image.

6 . The electronic device of claim 5 , wherein the controller is to perform facial detection on the color channels of the image to define a region of the image including the face and sample the depth information of the image within the region to determine the depth of the face.

7 . The electronic device of claim 6 , wherein the controller is to perform a depth cutoff of points of the depth information that have a greater distance from a viewpoint than the depth of the face plus a threshold value.

8 . The electronic device of claim 5 , wherein the controller is to apply the image mask to the image to separate the foreground of the image from the background of the image.

9 . A non-transitory computer-readable medium storing machine-readable instructions which, when executed by a controller of an electronic device, cause the controller to:

obtain color information of an image in separate red, green, and blue color channels;

obtain depth information of the image;

determine a depth of a face present in the image;

perform a depth cutoff of the depth information for points having greater than a threshold distance from the depth of the face;

provide a 4-channel input to a machine learning process to process the image according to the cutoff depth information and the color information to separate a foreground of the image from a background of the image, the 4-channel input including the cutoff depth information, the red color channel, the green color channel, and the blue color channel, the trained machine learning process to receive the 4-channel input and to output a segmentation result,

form an image mask by processing the 4-channel input according to the machine learning process;

apply the image mask to the image to obtain a separate representation of a foreground;

manipulate a background of the image by at least one of blurring or replacing the background; and

overlay the separate foreground representation on the manipulated background to provide video data for a video conferencing session.

10 . The computer-readable medium of claim 9 , wherein execution of the executable code causes the controller to determine a bounding box surrounding the face in the color information of the image and determine the depth of the face by sampling the depth information at points with a region bounded by the bounding box.

11 . The computer-readable medium of claim 9 , wherein execution of the executable code causes the controller to perform the depth cutoff according to Euclidean distance clustering.

12 . The computer-readable medium of claim 9 , wherein execution of the executable code causes the controller to overlay the foreground over a manipulated representation of the image.

13 . The computer-readable medium of claim 12 , wherein the manipulated representation is a blurring of the image or a replacement of the image.

14 . The electronic device of claim 1 , wherein the controller is to concatenate the RGB information with the truncated depth information to generate a 4-channel input for the trained machine learning process.

15 . The electronic device of claim 1 , wherein the controller is to manipulate the RGB information to overlay the RGB information of the foreground of the image on a manipulated representation of the RGB information of the background of the image.

16 . The electronic device of claim 5 , wherein the controller is to manipulate the color information to overlay the color information of the foreground of the image on a manipulated representation of the color information of the background of the image.

17 . The computer-readable medium of claim 9 , wherein execution of the executable code causes the controller to concatenate the red, green, and blue color channels with the truncated depth information to generate the 4-channel input for the trained machine learning process.

18 . The computer-readable medium of claim 9 , wherein execution of the executable code causes the controller to manipulate the red, green, and blue color channels to overlay the color information of the foreground of the image on a manipulated representation of the color information of the background of the image.

19 . The electronic device of claim 1 , wherein applying the image mask comprises separating foreground pixels corresponding to the face from background pixels prior to manipulating the background.

20 . The electronic device of claim 5 , wherein the truncated depth information is generated by excluding depth points having distances greater than the determined depth of the face plus a threshold value.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2023
From: ALLEBACH, JAN P; ZHU, FENGQING MAGGIE; SHEN, YEZHI; XU, WEICHEN
To: PURDUE RESEARCH FOUNDATION
Reel/Frame 062717/0320 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: LIN, QIAN
To: HEWLETT-PACKARD DEVELOPMENT COMPANY, L.P.
Reel/Frame 061458/0196 →
Continuity (1)
Related Publication 20240029472A1 · Jan 25, 2024
References Cited (66)
US 5914748A · Parulski et al. · 1999 [cited by applicant]
US 6141435A · Naoi et al. · 2000 [cited by applicant]
US 6184858B1 · Christian et al. · 2001 [cited by applicant]
US 6430303B1 · Naoi et al. · 2002 [cited by applicant]
US 6546115B1 · Ito et al. · 2003 [cited by applicant]
US 7783075B2 · Zhang et al. · 2010 [cited by applicant]
US 8073247B1 · Sandrew et al. · 2011 [cited by applicant]
US 9078048B1 · Gargi et al. · 2015 [cited by applicant]
US 10121229B2 · Gray et al. · 2018 [cited by applicant]
US 10311579B2 · Yun et al. · 2019 [cited by applicant]
US 11276177B1 · Tsai et al. · 2022 [cited by applicant]
US 20030090751A1 · Itokawa et al. · 2003 [cited by applicant]
US 20040151342A1 · Venetianer et al. · 2004 [cited by applicant]
US 20060285747A1 · Blake et al. · 2006 [cited by applicant]
US 20070133880A1 · Sun et al. · 2007 [cited by applicant]
US 20070286520A1 · Zhang et al. · 2007 [cited by applicant]
US 20090087096A1 · Eaton et al. · 2009 [cited by applicant]
US 20110285748A1 · Slatter et al. · 2011 [cited by applicant]
US 20120050323A1 · Baron et al. · 2012 [cited by applicant]
US 20120140066A1 · Lin · 2012 [cited by applicant]
US 20120207345A1 · Tang · 2012 [cited by applicant]
US 20120327172A1 · El-Saban et al. · 2012 [cited by applicant]
US 20130101208A1 · Feris et al. · 2013 [cited by applicant]
US 20140003720A1 · Seow et al. · 2014 [cited by applicant]
US 20140132789A1 · Koyama · 2014 [cited by applicant]
US 20150003675A1 · Nakagami · 2015 [cited by applicant]
US 20150221066A1 · Kobayashi · 2015 [cited by applicant]
US 20150310297A1 · Li et al. · 2015 [cited by applicant]
US 20160125268A1 · Ebiyama · 2016 [cited by applicant]
US 20170046827A1 · Meyer · 2017 [cited by examiner]
US 20170161905A1 · Reyzin · 2017 [cited by applicant]
US 20170263037A1 · Maruyama et al. · 2017 [cited by applicant]
US 20170330050A1 · Baltsn · 2017 [cited by applicant]
US 20170358094A1 · Sun et al. · 2017 [cited by applicant]
US 20180144476A1 · Smith · 2018 [cited by applicant]
US 20180255326A1 · Oya · 2018 [cited by applicant]
US 20180365809A1 · Cutler et al. · 2018 [cited by applicant]
US 20190052868A1 · Na et al. · 2019 [cited by applicant]
US 20190104253A1 · Kawai · 2019 [cited by applicant]
US 20190180107A1 · Pham · 2019 [cited by applicant]
US 20190347776A1 · Chou et al. · 2019 [cited by applicant]
US 20200273152A1 · Lin · 2020 [cited by applicant]
US 20200402243A1 · Benou et al. · 2020 [cited by applicant]
US 20210004943A1 · Fukasawa et al. · 2021 [cited by applicant]
US 20210166399A1 · Stephens et al. · 2021 [cited by applicant]
US 20210209734A1 · Simhadri · 2021 [cited by examiner]
US 20210243383A1 · Huang et al. · 2021 [cited by applicant]
US 20220109838A1 · Guruva et al. · 2022 [cited by applicant]
US 20220215531A1 · Azernikov et al. · 2022 [cited by applicant]
US 20220256116A1 · Chu et al. · 2022 [cited by applicant]
US 20230014805A1 · Verma et al. · 2023 [cited by applicant]
US 20230026038A1 · Iwakiri · 2023 [cited by applicant]
US 20230044969A1 · Yang et al. · 2023 [cited by applicant]
US 20230063678A1 · Janus et al. · 2023 [cited by applicant]
US 20230224582A1 · Awasthi et al. · 2023 [cited by applicant]
US 20240031517A1 · Liu et al. · 2024 [cited by applicant]
US 20240054608A1 · Shen et al. · 2024 [cited by applicant]
US 20240104742A1 · Xu et al. · 2024 [cited by applicant]
US 20240153038A1 · Duan et al. · 2024 [cited by applicant]
US 20240404017A1 · Xiang et al. · 2024 [cited by applicant]
CN 110232326A · 2019 [cited by examiner]
A machine translated English version of document CN 110232326 (Year: 2019). [cited by examiner]
Jiang et al, “Robust RGB-D Face Recognition Using Attribute-Aware Loss”, 2020 (Year: 2020). [cited by examiner]
Goswami et al, “RGB-D Face Recognition With Texture and Attribute Features”, 2014 (Year: 2014). [cited by examiner]
Howard, Andrew et al, “Searching for MobileNetV3”, arXiv:1905.02244v5, Nov. 20, 2019, 11 pages. [cited by applicant]
Mingyang et al, “An End-to-End Atrous Spatial Pyramid Pooling and Skip-Connections Generative Adversarial Segmentation Network for Building Extraction from High-Resolution Aerial Images”, vol. 12, 5151, Applied Sciences… [cited by applicant]