IP Library Granted Patent US 12664616
Granted Patent B2
US 12664616 · App. 18/252,979 · Granted Jun 23, 2026

Image processing method, model training method, apparatus, medium and device

Inventors: Jia Sun (Beijing, CN); Zehuan Yuan (Beijing, CN); Changhu Wang (Beijing, CN)
Assignee: Beijing Bytedance Network Technology Co., Ltd.
G06T5/50G06T3/40G06V10/751G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12664616
App. No.
18/252,979
Granted
Jun 23, 2026
Kind
B2
Abstract

The present disclosure relates to an image processing method and apparatus, a model training method and apparatus, a medium, and a device. The image processing method extracts a first target object image from an image to be processed; inputs the first target object image into a target object image processing model to obtain a second target object image output by the target object image processing model; and performs image fusion according to the second target object image and the image to be processed, so as to obtain a target image.

Claims (54)

1 . An image processing method, the method comprises:

extracting a first target object image from an image to be processed;

inputting the first target object image into a target object image processing model to obtain a second target object image output by the target object image processing model, wherein a resolution of the second target object image is higher than that of the first target object image; and

performing image fusion according to the second target object image and the image to be processed, so as to obtain a target image,

wherein the target object image processing model further includes a discriminator; and

determining whether the target object image processing model is completely trained according to target difference information between a high-resolution image and an original training sample image includes:

fusing each type of difference information included in the target difference information to obtain fused difference information; and

determining that the target object image processing model is completely trained if a degree of a difference characterized by the fused difference information is less than a preset fusion difference threshold and the discriminator judges an authenticity of the high-resolution image is real.

2 . The method according to claim 1 , wherein the target difference information further includes overall difference information between the high-resolution image adjusted to a preset resolution and the original training sample image adjusted to the preset resolution.

3 . The method according to claim 1 , wherein

the determining whether the target object image processing model is completely trained according to target difference information between the high-resolution image and the original training sample image includes:

determining that the target object image processing model is completely trained if a degree of a difference characterized by each type of difference information included in the target difference information is less than their corresponding difference threshold, and the discriminator judges an authenticity of the high-resolution image is real.

4 . The method according to claim 1 , wherein the performing image fusion according to the second target object image and the image to be processed so as to obtain the target image includes:

performing resolution enhancement processing on the image to be processed to obtain the target image to be processed; and

performing image fusion according to the second target object image and the target image to be processed to obtain the target image.

5 . The method according to claim 1 ,

wherein the target object image processing model is a generative adversarial network model including a generator, and the target object image processing model is trained by:

using a low-resolution image of an original training sample image as an input to the generator, to obtain the high-resolution image output by the generator after processing the low-resolution image;

determining whether the target object image processing model is completely trained according to the target difference information between the high-resolution image and the original training sample image, wherein the target difference information includes at least one of: feature point difference information between the high-resolution image and the original training sample image, and difference information in a specified feature region in the high-resolution image and the original training sample image; and

in response to the target object image processing model being completely trained, obtaining the target object image processing model.

6 . A training method for a target object image processing model, wherein the method comprises:

using a low-resolution image of an original training sample image as an input to a generator, to obtain a high-resolution image output by the generator after processing the low-resolution image, wherein the target object image processing model is a generative adversarial network model including the generator;

determining whether the target object image processing model is completely trained according to target difference information between the high-resolution image and the original training sample image, wherein the target difference information includes at least one of: feature point difference information between the high-resolution image and the original training sample image, and difference information in a specified feature region in the high-resolution image and the original training sample image; and

in response to the target object image processing model being completely trained, obtaining the target object image processing model,

wherein the target object image processing model further comprises a discriminator; and

the determining whether the target object image processing model is completely trained according to target difference information between the high-resolution image and the original training sample image includes:

fusing each type of difference information included in the target difference information to obtain fused difference information; and

determining that the target object image processing model is completely trained if a degree of a difference characterized by the fused difference information is less than a preset fusion difference threshold and the discriminator judges an authenticity of the high-resolution image is real.

7 . The method according to claim 6 , wherein the target difference information further includes overall difference information between the high-resolution image adjusted to a preset resolution and the original training sample image adjusted to the preset resolution.

8 . The method according to claim 6 , wherein

the determining whether the target object image processing model is completely trained according to target difference information between the high-resolution image and the original training sample image includes:

determining that the target object image processing model is completely trained if a degree of the difference characterized by each type of difference information included in the target difference information is less than their corresponding difference threshold, and the discriminator judges an authenticity of the high-resolution image is real.

9 . An electronic device comprising:

a storage apparatus having a computer program stored thereon;

a processing apparatus configured to execute the computer program in the storage apparatus to implement an image processing method, the method comprises:

extracting a first target object image from an image to be processed;

inputting the first target object image into a target object image processing model to obtain a second target object image output by the target object image processing model, wherein a resolution of the second target object image is higher than that of the first target object image; and

performing image fusion according to the second target object image and the image to be processed, so as to obtain a target image,

wherein the target object image processing model further includes a discriminator;

determining whether the target object image processing model is completely trained according to target difference information between a high-resolution image and an original training sample image includes:

fusing each type of difference information included in the target difference information to obtain fused difference information; and

determining that the target object image processing model is completely trained if a degree of a difference characterized by the fused difference information is less than a preset fusion difference threshold and the discriminator judges an authenticity of the high-resolution image is real.

10 . The device according to claim 9 , wherein the target difference information further includes overall difference information between the high-resolution image adjusted to a preset resolution and the original training sample image adjusted to the preset resolution.

11 . The device according to claim 9 , wherein

the determining whether the target object image processing model is completely trained according to target difference information between the high-resolution image and the original training sample image includes:

determining that the target object image processing model is completely trained if a degree of a difference characterized by each type of difference information included in the target difference information is less than their corresponding difference threshold, and the discriminator judges an authenticity of the high-resolution image is real.

12 . The device according to claim 9 , wherein the performing image fusion according to the second target object image and the image to be processed so as to obtain the target image includes:

performing resolution enhancement processing on the image to be processed to obtain the target image to be processed; and

performing image fusion according to the second target object image and the target image to be processed to obtain the target image.

13 . The device according to claim 9 , wherein

the target object image processing model is a generative adversarial network model including a generator, and the target object image processing model is trained by:

using a low-resolution image of an original training sample image as an input to the generator, to obtain the high-resolution image output by the generator after processing the low-resolution image;

determining whether the target object image processing model is completely trained according to the target difference information between the high-resolution image and the original training sample image, wherein the target difference information includes at least one of: feature point difference information between the high-resolution image and the original training sample image, and difference information in a specified feature region in the high-resolution image and the original training sample image; and

in response to the target object image processing model being completely trained, obtaining the target object image processing model.