IP Library › Granted Patent US 12,652,373
Granted Patent B2
US 12,652,373 · App. 18/123,577 · Granted Jun 9, 2026

Image processing method and apparatus, device, and medium

Inventors: Mingliang Chen (Beijing, CN); Jun Ouyang (Beijing, CN); Zhicheng Li (Beijing, CN)
Assignee: TENCENT CLOUD COMPUTING (BEIJING) CO., LTD
H04N9/64G06T5/00G06T5/40G06T5/92G06T7/90G06T2207/10016G06T2207/10024G06T2207/20072G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,652,373
App. No.
18/123,577
Granted
Jun 9, 2026
Kind
B2
Abstract

An image processing method is provided. In the method, a target video frame set is acquired from video data of a plurality of video frames. The target video frame set includes a subset of the video frames that is selected based on characteristics of the subset of the video frames. A global color feature of a reference video frame is acquired. An image semantic feature of the reference video frame is acquired. An enhancement parameter of the reference video frame is acquired for each of at least one image information dimension according to the global color feature and the image semantic feature. Image enhancement is separately performed on the video frames in the target video frame set according to each enhancement parameter of the reference video frame to obtain target image data for each of the video frames in the target video frame set.

Claims (75)

1 . An image processing method, comprising:

acquiring a target video frame set from video data of a plurality of video frames, the target video frame set including a subset of the plurality of video frames that is selected based on characteristics of the subset of the plurality of video frames;

acquiring a global color feature of a reference video frame, the reference video frame being one of the subset of the plurality of video frames in the target video frame set;

acquiring an image semantic feature of the reference video frame;

acquiring an enhancement parameter of the reference video frame for each of a plurality of image information dimensions by splicing the global color feature and the image semantic feature, the enhancement parameter of the reference video frame for each of the plurality of image information dimensions including a brightness enhancement parameter, a contrast enhancement parameter, and a saturation enhancement parameter, each of the brightness enhancement parameter, the contrast enhancement parameter, and the saturation enhancement parameter being predicted by a respective regression network; and

separately performing image enhancement on the plurality of video frames in the target video frame set according to each enhancement parameter of the reference video frame to obtain target image data for each of the plurality of video frames in the target video frame set, the target image data for a first video frame of the plurality of video frames being determined based on a product of (i) the saturation enhancement parameter and (ii) a second difference between a pixel value of a second candidate video frame and a channel average pixel value of the second candidate video frame, the second candidate video frame being determined based on a product of (i) the contrast enhancement parameter and (ii) a first difference between a pixel value of a first candidate video frame and a global average pixel value of the first candidate video frame, and the first candidate video frame being determined based on a product of the first video frame and the brightness enhancement parameter.

2 . The method according to claim 1 , wherein the characteristics of the subset of the plurality of video frames correspond to a same scene attribute.

3 . The method according to claim 1 , wherein the acquiring the target video frame set comprises:

acquiring a color histogram of each of the plurality of video frames of the video data;

acquiring a similarity distance between each of adjacent video frames in the plurality of video frames of the video data according to the color histograms of the respective adjacent video frames;

dividing the plurality of video frames into a plurality of video frame sets; and

selecting the target video frame set from the plurality of video frame sets, the similarity distance between each of the adjacent video frames in the target video frame set being less than a distance threshold.

4 . The method according to claim 3 , wherein the acquiring the color histogram of each of the plurality of video frames of the video data comprises:

counting, according to a color space of pixels in a video frame, a pixel quantity of pixels included in each of a plurality of image color ranges, the video frame being one of the plurality of video frames of the video data, the plurality of image color ranges being obtained by dividing the color space; and

generating, according to the pixel quantity corresponding to each of the image color ranges, the color histogram corresponding to the video frame.

5 . The method according to claim 1 , wherein the acquiring the global color feature of the reference video frame comprises:

reducing a size of the reference video frame to obtain a candidate video frame with a target size, and acquiring, according to a color histogram corresponding to the candidate video frame, the global color feature of the reference video frame.

6 . The method according to claim 1 , wherein the acquiring the image semantic feature of the reference video frame comprises:

performing a convolution operation on the reference video frame by using a convolutional layer in a feature extraction model, to obtain an image convolution feature of the reference video frame; and

performing a residual operation on the image convolution feature by using a residual layer in the feature extraction model, to obtain the image semantic feature of the reference video frame.

7 . The method according to claim 1 , wherein the acquiring the enhancement parameter of the reference video frame for each of the plurality of image information dimensions comprises:

splicing the global color feature and the image semantic feature to obtain a target image feature; and

processing the target image feature by using a target generation model to predict the enhancement parameter of the target image feature for each of the plurality of image information dimensions.

8 . The method according to claim 7 , wherein

when the enhancement parameter for each of the plurality of image information dimensions includes the brightness enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a first regression network in the target generation model, to obtain the brightness enhancement parameter of the reference video frame;

when the enhancement parameter for each of the plurality of image information dimensions includes the contrast enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a second regression network in the target generation model, to obtain the contrast enhancement parameter of the reference video frame; and

when the enhancement parameter for each of the plurality of image information dimensions includes the saturation enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a third regression network in the target generation model, to obtain the saturation enhancement parameter of the reference video frame.

9 . The method according to claim 8 , wherein

the separately performing includes:

multiplying the first video frame of the target video frame set by the brightness enhancement parameter to determine the first candidate video frame corresponding to the first video frame;

acquiring the global average pixel value corresponding to the first candidate video frame, acquiring the first difference between the pixel value in the first candidate video frame and the global average pixel value, and determining, according to the global average pixel value and the product of the first difference and the contrast enhancement parameter, the second candidate video frame corresponding to the first video frame; and

acquiring the channel average pixel value corresponding to the second candidate video frame, acquiring the second difference between the pixel value in the second candidate video frame and the channel average pixel value, and determining, according to the channel average pixel value and the product of the second difference and the saturation enhancement parameter, the target image data corresponding to the first video frame.

10 . The method according to claim 7 , further comprising:

processing, by using an initial generation model, a sample global color feature and a sample image semantic feature of a sample video frame, and outputting a predicted enhancement parameter of the sample video frame, the sample video frame being obtained by performing a random degradation operation on a color of the reference video frame;

determining, according to label information of the sample video frame and the predicted enhancement parameter of the sample video frame, a loss function of the initial generation model, the label information being determined according to a coefficient of the random degradation operation; and

performing iterative training on the initial generation model according to the loss function, and when the initial generation model meets a convergence condition, determining the initial generation model meeting the convergence condition as the target generation model.

11 . An image processing apparatus, comprising:

processing circuitry configured to:

acquire a target video frame set from video data of a plurality of video frames, the target video frame set including a subset of the plurality of video frames that is selected based on characteristics of the subset of the plurality of video frames;

acquire a global color feature of a reference video frame, the reference video frame being one of the subset of the plurality of video frames in the target video frame set;

acquire an image semantic feature of the reference video frame;

acquire an enhancement parameter of the reference video frame for each of a plurality of image information dimensions by splicing the global color feature and the image semantic feature, the enhancement parameter of the reference video frame for each of the plurality of image information dimensions including a brightness enhancement parameter, a contrast enhancement parameter, and a saturation enhancement parameter, each of the brightness enhancement parameter, the contrast enhancement parameter, and the saturation enhancement parameter being predicted by a respective regression network; and

separately perform image enhancement on the plurality of video frames in the target video frame set according to each enhancement parameter of the reference video frame to obtain target image data for each of the plurality of video frames in the target video frame set, the target image data for a first video frame of the plurality of video frames being determined based on a product of (i) the saturation enhancement parameter and (ii) a second difference between a pixel value of a second candidate video frame and a channel average pixel value of the second candidate video frame, the second candidate video frame being determined based on a product of (i) the contrast enhancement parameter and (ii) a first difference between a pixel value of a first candidate video frame and a global average pixel value of the first candidate video frame, and the first candidate video frame being determined based on a product of the first video frame and the brightness enhancement parameter.

12 . The image processing apparatus according to claim 11 , wherein the characteristics of the subset of the plurality of video frames correspond to a same scene attribute.

13 . The image processing apparatus according to claim 11 , wherein the processing circuitry is configured to:

acquire a color histogram of each of the plurality of video frames of the video data;

acquire a similarity distance between each of adjacent video frames in the plurality of video frames of the video data according to the color histograms of the respective adjacent video frames;

divide the plurality of video frames into a plurality of video frame sets; and

select the target video frame set from the plurality of video frame sets, the similarity distance between each of the adjacent video frames in the target video frame set being less than a distance threshold.

14 . The image processing apparatus according to claim 13 , wherein the processing circuitry is configured to:

count, according to a color space of pixels in a video frame, a pixel quantity of pixels included in each of a plurality of image color ranges, the video frame being one of the plurality of video frames of the video data, the plurality of image color ranges being obtained by dividing the color space; and

generate, according to the pixel quantity corresponding to each of the image color ranges, the color histogram corresponding to the video frame.

15 . The image processing apparatus according to claim 11 , wherein the processing circuitry is configured to:

reduce a size of the reference video frame to obtain a candidate video frame with a target size, and acquire, according to a color histogram corresponding to the candidate video frame, the global color feature of the reference video frame.

16 . The image processing apparatus according to claim 11 , wherein the processing circuitry is configured to:

perform a convolution operation on the reference video frame by using a convolutional layer in a feature extraction model, to obtain an image convolution feature of the reference video frame; and

perform a residual operation on the image convolution feature by using a residual layer in the feature extraction model, to obtain the image semantic feature of the reference video frame.

17 . The image processing apparatus according to claim 11 , wherein the processing circuitry is configured to:

splice the global color feature and the image semantic feature to obtain a target image feature; and

process the target image feature by using a target generation model to predict the enhancement parameter of the target image feature for each of the plurality of image information dimensions.

18 . The image processing apparatus according to claim 17 , wherein

when the enhancement parameter for each of the plurality of image information dimensions includes the brightness enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a first regression network in the target generation model, to obtain the brightness enhancement parameter of the reference video frame;

when the enhancement parameter for each of the plurality of image information dimensions includes the contrast enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a second regression network in the target generation model, to obtain the contrast enhancement parameter of the reference video frame; and

when the enhancement parameter for each of the plurality of image information dimensions includes the saturation enhancement parameter, the target image feature is weighted by using a weight matrix corresponding to a third regression network in the target generation model, to obtain the saturation enhancement parameter of the reference video frame.

19 . The image processing apparatus according to claim 18 , wherein

the processing circuitry is configured to:

multiply the first video frame of the target video frame set by the brightness enhancement parameter to determine the first candidate video frame corresponding to the first video frame;

acquire the global average pixel value corresponding to the first candidate video frame, acquiring the first difference between the pixel value in the first candidate video frame and the global average pixel value, and determine, according to the global average pixel value and the product of the first difference and the contrast enhancement parameter, the second candidate video frame corresponding to the first video frame; and

acquire the channel average pixel value corresponding to the second candidate video frame, acquire the second difference between the pixel value in the second candidate video frame and the channel average pixel value, and determine, according to the channel average pixel value and the product of the second difference and the saturation enhancement parameter, the target image data corresponding to the first video frame.

20 . A non-transitory computer-readable storage medium, storing instructions which when executed by a processor cause the processor to perform:

acquiring a target video frame set from video data of a plurality of video frames, the target video frame set including a subset of the plurality of video frames that is selected based on characteristics of the subset of the plurality of video frames;

acquiring a global color feature of a reference video frame, the reference video frame being one of the subset of the plurality of video frames in the target video frame set;

acquiring an image semantic feature of the reference video frame;

acquiring an enhancement parameter of the reference video frame for each of a plurality of image information dimensions by splicing the global color feature and the image semantic feature, the enhancement parameter of the reference video frame for each of the plurality of image information dimensions including a brightness enhancement parameter, a contrast enhancement parameter, and a saturation enhancement parameter, each of the brightness enhancement parameter, the contrast enhancement parameter, and the saturation enhancement parameter being predicted by a respective regression network; and

separately performing image enhancement on the plurality of video frames in the target video frame set according to each enhancement parameter of the reference video frame to obtain target image data for each of the plurality of video frames in the target video frame set, the target image data for a first video frame of the plurality of video frames being determined based on a product of (i) the saturation enhancement parameter and (ii) a second difference between a pixel value of a second candidate video frame and a channel average pixel value of the second candidate video frame, the second candidate video frame being determined based on a product of (i) the contrast enhancement parameter and (ii) a first difference between a pixel value of a first candidate video frame and a global average pixel value of the first candidate video frame, and the first candidate video frame being determined based on a product of the first video frame and the brightness enhancement parameter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 20, 2023
From: CHEN, MINGLIANG; OUYANG, JUN; LI, ZHICHENG
To: TENCENT CLOUD COMPUTING (BEIJING) CO., LTD
Reel/Frame 063034/0001 →
Priority Claims (1)
CN 202110468372.7 · Apr 28, 2021 · national
Continuity (2)
Continuation PCTCN2021108468 · Jul 26, 2021
Related Publication 20230230215A1 · Jul 20, 2023
References Cited (17)
US 11776273B1 · Chen · 2023 [cited by examiner]
US 20120293660A1 · Murakami · 2012 [cited by examiner]
US 20140072284A1 · Avrahami et al. · 2014 [cited by applicant]
US 20180270489A1 · Maymon · 2018 [cited by examiner]
US 20200329217A1 · Sun · 2020 [cited by examiner]
US 20210112226A1 · Abou Shousha · 2021 [cited by examiner]
US 20210289186A1 · Peng · 2021 [cited by examiner]
US 20210358087A1 · Huang · 2021 [cited by examiner]
US 20220141455A1 · Cricri · 2022 [cited by examiner]
US 20230260155A1 · Zhang · 2023 [cited by examiner]
CN 103516989A · 2014 [cited by applicant]
CN 109525901A · 2019 [cited by applicant]
CN 111031346A · 2020 [cited by applicant]
CN 111583138A · 2020 [cited by applicant]
International Search Report PCT/CN2021/108468, mailed Jan. 19, 2022, 5 pages. [cited by applicant]
Park et al., “Distort-and-Recover: Color Enhancement using Deep Reinforcement Learning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 5928-5936, 9 pages. [cited by applicant]
Written Opinion in PCT/CN2021/108468, mailed Jan. 19, 2022, 4 pages. [cited by applicant]