IP Library Granted Patent US 12,670,542
Granted Patent B2
US 12,670,542 · App. 17/981,153 · Granted Jun 30, 2026

Joint objects image signal processing in temporal domain

Inventors: Marat Ravilevich Gilmutdinov (Saint Petersburg, RU); Elena Alexandrovna Alshina (Munich, DE); Nickolay Dmitrievich Egorov (Saint Petersburg, RU); Dmitry Vadimovich Novikov (Saint Petersburg, RU); Anton Igorevich Veselov (Saint Petersburg, RU); Kirill Aleksandrovich Malakhov (Saint Petersburg, RU); Nikita Vyacheslavovich Ustiuzhanin (Saint Petersburg, RU)
Assignee: Huawei Technologies Co., Ltd.
G06T3/4015G06F18/2414G06V10/30G06V10/761G06V10/82G06V20/41H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,542
App. No.
17/981,153
Filed
Nov 4, 2022
Granted
Jun 30, 2026
Kind
B2
Examiner
NIU, FENG
Art Unit
2669
USPC
382/156
Abstract

The present disclosure relates to pre-processing of video images. In particular, the video images are pre-processed in an object-based manner, i.e., by applying different pre-processing to different objects detected in the image. Moreover, the pre-processing is applied to a group of images. As such, object detection is performed in a plurality of images and the pre-processing for the plurality of images may be adapted to the decoded images and is applied to the decoded images.

Claims (73)

1 . A method for processing frames of a video sequence in raw image format, the method comprising:

performing pre-processing on at least two respective frames of the video sequence, wherein the pre-processing includes filtering the at least two respective frames with a filter of which parameters are set according to a size of an object to be detected in the at least two respective frames;

identifying the object included in image regions of the at least two respective frames of the video sequence, wherein identifying the object comprises:

detecting a location of the object within the at least two respective frames by distinguishing the object from other parts of the at least two respective frames; and

recognizing an identity of the object in the at least two respective frames based on detecting the location of the object within the at least two respective frames, wherein the performing pre-processing on the at least two respective frames of the video sequence is performed before the recognizing the identity of the object; and

processing the image regions in the at least two respective frames that include the object by a first image processing adapted to the object and different from a second image processing applied to image regions in the at least two respective frames not including the object, wherein the first image processing includes de-noising with a de-noising filter of which at least one parameter is determined based on an object bounding box of the object and a size of the object bounding box included in the at least two respective frames, and wherein the at least one parameter includes a volume in a number of samples of the object bounding box.

2 . The method according to claim 1 , wherein the recognizing the identity of the object comprises:

computing a plurality of feature vectors for a plurality of image regions in the at least two respective frames, wherein computing of a feature vector of the plurality of feature vectors for a given frame includes determining a value of at least one feature of the image region of the given frame; and

forming a cluster based on the feature vectors, wherein the cluster includes image regions of the at least two respective frames, the image regions including the object with the same recognized identity.

3 . The method according to claim 2 ,

wherein the forming the cluster is performed by K-means approach; and/or

wherein the forming the cluster is based on determining a similarity measure of feature vectors calculated for the image regions in different frames among the at least two respective frames, wherein the similarity measure employed is one of Euclidean distance, Chebyschev distance, or cosine similarity.

4 . The method according to claim 1 ,

wherein the identifying the object further comprises:

detecting one or more classes of the object; and

wherein the recognizing of the identity of the object is based on at least one of the one or more detected classes of the object.

5 . The method according to claim 4 , wherein the detecting of the location of the object and the detecting the one or more classes of the object is performed by a YOLO (You Only Look Once) object detection algorithm, a mobileNet object detection algorithm, a SSD (Single Shot Multibox Detector) object detection algorithm, a SSH (Single Stage Headless) face detection algorithm, or a MTCNN (Multi-task Cascaded Convolutional Neural Network) face detection algorithm.

6 . The method according to claim 1 , wherein the performing the pre-processing further includes:

filtering the at least two respective frames with a filter adapted to a type of the identity of the object.

7 . The method according to claim 1 , further comprising:

obtaining the at least two respective frames from an image sensor;

wherein the first image processing and/or the second image processing comprises performing at least one of:

defect pixel correction,

white balance,

de-noising,

demosaicing,

color space correction,

color enhancement,

contrast enhancement,

sharpening, or

color transformation.

8 . The method according to claim 1 , wherein the raw image format is a Bayer pattern and the performing pre-processing includes conversion of the at least two respective frames into an RGB (red-green-blue) image format.

9 . The method according to claim 1 , wherein the at least two respective frames are:

temporally adjacent frames; or

more than two frames equally spaced in a time domain.

10 . The method according to claim 1 , further comprising:

encoding the at least two respective frames of the video sequence by applying lossy and/or lossless compression.

11 . A non-transitory computer-readable storage medium that stores a computer program that, when executed on one or more processors, causes the one or more processors to execute operations comprising:

performing pre-processing on at least two respective frames of a video sequence, wherein the pre-processing includes filtering the at least two respective frames with a filter of which parameters are set according to a size of an object to be detected in the at least two respective frames;

identifying the object included in image regions of the at least two respective frames of the video sequence, wherein identifying the object comprises:

detecting a location of the object within the at least two respective frames by distinguishing the object from other parts of the at least two respective frames; and

recognizing an identity of the object in the at least two respective frames based on detecting the location of the object within the at least two respective frames, wherein the performing pre-processing on the at least two respective frames of the video sequence is performed before the recognizing the identity of the object; and

processing the image regions in the at least two respective frames that include the object by a first image processing adapted to the object and different from a second image processing applied to image regions in the at least two respective frames not including the object, wherein the first image processing includes de-noising with a de-noising filter of which at least one parameter is determined based on an object bounding box of the object and a size of the object bounding box included in the at least two respective frames, and wherein the at least one parameter includes a volume in a number of samples of the object bounding box.

12 . An apparatus for processing frames of a video sequence in raw image format, the apparatus comprising:

processing circuitry configured to:

perform pre-processing on at least two respective frames of the video sequence, wherein the pre-processing includes filtering the at least two respective frames with a filter of which parameters are set according to a size of an object to be detected in the at least two respective frames;

identify the object in image regions of the at least two respective frames of the video sequence, wherein identifying the object comprises:

detecting a location of the object within the at least two respective frames by distinguishing the object from other parts of the at least two respective frames; and

recognizing an identity of the object in the at least two respective frames based on detecting the location of the object within the at least two respective frames, wherein the performing pre-processing on the at least two respective frames of the video sequence is performed before the recognizing the identity of the object; and

process the image regions in the at least two respective frames that include the object by a first image processing adapted to the object and different from a second image processing applied to image regions in the at least two respective frames not including the object, wherein the first image processing includes de-noising with a de-noising filter of which at least one parameter is determined based on an object bounding box of the object and a size of the object bounding box included in the at least two respective frames, and wherein the at least one parameter includes a volume in a number of samples of the object bounding box.

13 . The apparatus according to claim 12 , further comprising:

an image sensor for capturing the video sequence in the raw image format.

14 . The method according to claim 1 , wherein the location of the object is a location of a bounding box framing the object or a pixel map.

15 . The computer-readable storage medium according to claim 11 , wherein the recognizing the identity of the object comprises:

computing a plurality of feature vectors for a plurality of image regions in the at least two respective frames, wherein computing of a feature vector of the plurality of feature vectors for a given frame includes determining a value of at least one feature of the image region of the given frame; and

forming a cluster based on the feature vectors, wherein the cluster includes image regions of the at least two respective frames, the image regions including the object with the same recognized identity.

16 . The computer-readable storage medium according to claim 15 ,

wherein the forming the cluster is performed by K-means approach; and/or

wherein the forming the cluster is based on determining a similarity measure of feature vectors calculated for the image regions in different frames among the at least two respective frames, wherein the similarity measure employed is one of Euclidean distance, Chebyschev distance, or cosine similarity.

17 . The computer-readable storage medium according to claim 11 ,

wherein the identifying the object further comprises:

detecting one or more classes of the object; and

wherein the recognizing of the identity of the object is based on at least one of the one or more detected classes of the object.

18 . The apparatus according to claim 12 , wherein the recognizing the identity of the object comprises:

computing a plurality of feature vectors for a plurality of image regions in the at least two respective frames, wherein computing of a feature vector of the plurality of feature vectors for a given frame includes determining a value of at least one feature of the image region of the given frame; and

forming a cluster based on the feature vectors, wherein the cluster includes image regions of the at least two respective frames, the image regions including the object with the same recognized identity.

19 . The apparatus according to claim 18 ,

wherein the forming the cluster is performed by K-means approach; and/or

wherein the forming the cluster is based on determining a similarity measure of feature vectors calculated for the image regions in different frames among the at least two respective frames, wherein the similarity measure employed is one of Euclidean distance, Chebyschev distance, or cosine similarity.

20 . The apparatus according to claim 12 ,

wherein the identifying the object further comprises:

detecting one or more classes of the object; and

wherein the recognizing of the identity of the object is based on at least one of the one or more detected classes of the object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 19, 2026
From: GILMUTDINOV, MARAT RAVILEVICH; EGOROV, NICKOLAY DMITRIEVICH; NOVIKOV, DMITRY VADIMOVICH; VESELOV, ANTON IGOREVICH; USTIUZHANIN, NIKITA VYACHESLAVOVICH; ALSHINA, ELENA ALEXANDROVNA; MALAKHOV, KIRILL ALEKSANDROVICH
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 074698/0158 →
Priority Claims (1)
WO PCT/EP2020/062557 · May 6, 2020 · international
Continuity (2)
Continuation PCTRU2021050113 · Apr 28, 2021
Related Publication 20230127009A1 · Apr 27, 2023
References Cited (79)
US 5878415A · Olds · 1999 [cited by examiner]
US 8228995B2 · Kondo · 2012 [cited by examiner]
US 8254717B2 · Velthoven · 2012 [cited by examiner]
US 8374392B2 · Hu · 2013 [cited by examiner]
US 8374457B1 · Wang · 2013 [cited by examiner]
US 8786625B2 · Cote · 2014 [cited by examiner]
US 8855434B2 · Sato · 2014 [cited by examiner]
US 8948506B2 · Saito · 2015 [cited by examiner]
US 9042599B2 · Du · 2015 [cited by examiner]
US 9147230B2 · Saito · 2015 [cited by examiner]
US 9225897B1 · Sehn · 2015 [cited by examiner]
US 9386287B2 · Hatano · 2016 [cited by examiner]
US 9398280B2 · Nikkanen · 2016 [cited by examiner]
US 9639762B2 · Chakraborty · 2017 [cited by examiner]
US 10055821B2 · Glotzbach · 2018 [cited by examiner]
US 10511846B1 · Chen et al. · 2019 [cited by applicant]
US 10521885B2 · Takahashi · 2019 [cited by examiner]
US 10893283B2 · Chen · 2021 [cited by examiner]
US 11113969B2 · Avedisov · 2021 [cited by examiner]
US 11657481B2 · Sytnik · 2023 [cited by examiner]
US 11677918B2 · Cho · 2023 [cited by examiner]
US 11995895B2 · Shin · 2024 [cited by examiner]
US 12022189B2 · Jung · 2024 [cited by examiner]
US 20050133708A1 · Eberhard · 2005 [cited by examiner]
US 20070165140A1 · Kondo · 2007 [cited by examiner]
US 20080219582A1 · Kirenko · 2008 [cited by examiner]
US 20090237563A1 · Doser · 2009 [cited by examiner]
US 20090278953A1 · Velthoven · 2009 [cited by examiner]
US 20100027906A1 · Hara · 2010 [cited by examiner]
US 20100265353A1 · Koyama · 2010 [cited by examiner]
US 20100296702A1 · Hu · 2010 [cited by examiner]
US 20110058609A1 · Chaudhury et al. · 2011 [cited by applicant]
US 20110249142A1 · Brunner · 2011 [cited by applicant]
US 20120081385A1 · Cote · 2012 [cited by examiner]
US 20120114172A1 · Du · 2012 [cited by examiner]
US 20130028531A1 · Sato · 2013 [cited by examiner]
US 20130114713A1 · Bossen · 2013 [cited by examiner]
US 20130216130A1 · Saito · 2013 [cited by examiner]
US 20130272605A1 · Saito · 2013 [cited by examiner]
US 20130273968A1 · Rhoads · 2013 [cited by examiner]
US 20140078338A1 · Hatano · 2014 [cited by examiner]
US 20140176548A1 · Green · 2014 [cited by examiner]
US 20150054980A1 · Nikkanen et al. · 2015 [cited by applicant]
US 20150161773A1 · Takahashi · 2015 [cited by examiner]
US 20160070963A1 · Chakraborty · 2016 [cited by examiner]
US 20170221186A1 · Glotzbach · 2017 [cited by examiner]
US 20180089906A1 · Alkouh · 2018 [cited by examiner]
US 20200029075A1 · Sato · 2020 [cited by examiner]
US 20200099944A1 · Chen · 2020 [cited by examiner]
US 20200178809A1 · Wang · 2020 [cited by examiner]
US 20200241100A1 · Rehwald · 2020 [cited by examiner]
US 20200286382A1 · Avedisov · 2020 [cited by examiner]
US 20200380274A1 · Shin · 2020 [cited by examiner]
US 20220182590A1 · Cho · 2022 [cited by examiner]
US 20230063005A1 · Nakao · 2023 [cited by examiner]
US 20230091780A1 · Jung · 2023 [cited by examiner]
US 20230281764A1 · Sytnik · 2023 [cited by examiner]
JP 2009145991A · 2009 [cited by examiner]
JP 2019083467A · 2019 [cited by examiner]
KR 101848507B1 · 2018 [cited by examiner]
JP-2019083467-A (machine translation) (Year: 2019). [cited by examiner]
KR-101848507-B1 (machine translation) (Year: 2018). [cited by examiner]
JP-2009145991-A (machine translation) (Year: 2009). [cited by examiner]
Nasser et al., “A novel generic dictionary-based denoising method for improving noisy and densely packed nuclei segmentation in 3D time-lapse fluorescence microscopy images.” Sci Rep. Apr. 4, 2019;9(1):5654. doi: 10.103… [cited by examiner]
Lloyd “Least Squares Quantization in Pcm,” IEEE Transactions on Information Theory, vol. 28, No. 2, pp. 129-137, Institute of Electrical and Electronics Engineers, New York, New York (Mar. 1982). [cited by applicant]
Park “Architectural Analysis of a Baseline ISP Pipeline,” in KYUNG (Ed.) “Theory and Applications of Smart Cameras,” KAIST Research Series, pp. 21-45, Springer, Dordrecht, Netherlands (Jul. 2015). [cited by applicant]
Redmon et al., “YOLOv3: An Incremental Improvement,” Computer Vision and Pattern Recognition, arXiv:1804.02767v1 [cs.CV], Total 6 pages (Apr. 2018). [cited by applicant]
Zhang et al., “Joint Face Detection and Alignment using Multi-task Cascaded Convolutional Networks,” IEEE Signal Processing Letters, vol. 23, No. 10, pp. 1-5, Institute of Electrical and Electronics Engineers, New York,… [cited by applicant]
Schroff et al., “FaceNet: A Unified Embedding for Face Recognition and Clustering,” 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Total 10 pages, Institute of Electrical and Electronics Enginee… [cited by applicant]
Van De Weijer et al., “Color Constancy based on the Grey-Edge Hypothesis,” IEEE International Conference on mage Processing 2005, Total 4 pages, Institute of Electrical and Electronics Engineers, New York, New York (Sep… [cited by applicant]
He et al., “Guided Image Filtering,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 35, No. 6, pp. 1397-1409, DOI: 10.1109/TPAMI.2012.213, Institute of Electrical and Electronics Engineers, New Yor… [cited by applicant]
Buades et al., “A non-local algorithm for image denoising,” 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR'05), Total 6 pages, Institute of Electrical and Electronics Engineers, N… [cited by applicant]
Hirakawa et al., “Adaptive Homogeneity-Directed Demosaicing Algorithm,” IEEE Transactions on Image Processing, vol. 14, No. 3, pp. 360-369, DOI: 10.1109/TIP.2004.838691, Institute of Electrical and Electronics Engineers… [cited by applicant]
Menon et al., “Demosaicing With Directional Filtering and a posteriori Decision,” IEEE Transactions on Image Processing, vol. 16, No. 1, pp. 132-141, DOI: 10.1109/TIP.2006.884928, Institute of Electrical and Electronics… [cited by applicant]
Remez et al., “Deep Class-Aware Image Denoising,” 2017 International Conference on Sampling Theory and Applications (SampTA), Tallin, Estonia, pp. 138-142, DOI: 10.1109/SAMPTA.2017.8024474, Institute of Electrical and E… [cited by applicant]
Liao et al., “Video-based Person Re-identification via 3D Convolutional Networks and Non-local Attention,” Computer Vision and Pattern Recognition, pp. 1-9, Asian Conference on Computer Vision (ACCV) (last revised: Apr.… [cited by applicant]
Ramanath et al., “Color Image Processing Pipeline in Digital Still Cameras,” IEEE Signal Processing Magazine, pp. 1-20 (2004). [cited by applicant]
Martinez-Ponte et al., “Robust Human Face Hiding Ensuring Privacy,” 6th International Workshop on Image Analysis for Multimedia Interactive Services, XP055012519, Montreux, Switzerland, Total 4 pages (Apr. 13, 2005). [cited by applicant]
Hukkelas, “DeepPrivacy: A Generative Adversarial Network for Face Anonymization,” ISVC 2019, https://doi.org/10.1007/978-3-030-33720-9_44, XP047545421, Total 14 pages (Oct. 21, 2019). [cited by applicant]