IP Library Granted Patent US 12,254,544
Granted Patent B2
US 12,254,544 · App. 17/634,002 · Granted Mar 18, 2025

Image-text fusion method and apparatus, and electronic device

Inventors: Wenjie Zhang (Nanjing, CN); Weicai Zhong (Xi'an, CN); Liang Hu (Shenzhen, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06T11/60G06T11/001G06V10/44G06V10/462G06V10/54G06V10/56G06V20/62G06V40/168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,544
App. No.
17/634,002
Granted
Mar 18, 2025
Kind
B2
Abstract

This application relates to the field of digital image processing technologies, and discloses an image-text fusion method and apparatus, and an electronic device, to minimize blockage of a saliency feature in an image by a text when the text is laid out in the image, and obtain a higher visual balance degree after the text is laid out in the first image, thereby achieving a better layout effect. According to the method of this application, first, a plurality of candidate text templates and layout positions of a plurality of corresponding texts in an image can be determined, so that a text laid out in the image does not block a visually salient object having a greater feature value, such as a human face or a building. Then based on magnitudes of feature values of pixels blocked by the text, a balance degree of feature value distribution of pixels in each region in the image in which the text is laid out, and the like when the text is laid out in the image at corresponding layout positions in the image by using different text templates, a final text template of the text and a layout position of the text in the image are determined.

Claims (71)

1. An image-text fusion method, wherein the method comprises:

obtaining a first image and a first text to be laid out in the first image;

determining a feature value of each pixel in the first image, wherein a feature value of a pixel is used to represent a probability that a user pays attention to the pixel, wherein the probability that the user pays attention to the pixel is higher for greater feature values of the pixel;

determining a plurality of first layout formats of the first text in the first image based on the first text and the feature value of each pixel in the first image, wherein when the first text is laid out in the first image based on each first layout format, the first text does not block a pixel whose feature value is greater than a first threshold;

determining a second layout format from the plurality of first layout formats based on cost parameters of the plurality of first layout formats, wherein a cost parameter of a first layout format is used to represent a magnitude of a feature value of a pixel blocked by the first text when the first text is laid out in the first image based on the first layout format, and a balance degree of feature value distribution of pixels in each region in the first image in which the first text is laid out; and

laying out the first text in the first image based on the second layout format to obtain a second image.

2. The image-text fusion method according to claim 1 , wherein the determining of the feature value of each pixel in the first image further comprises:

determining at least two parameters of a visual saliency parameter, a face feature parameter, an edge feature parameter, and a text feature parameter of each pixel in the first image, wherein the visual saliency parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a visual saliency feature, the face feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a face, the edge feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to an object contour, and the text feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a text; and

separately performing weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image.

3. The image-text fusion method according to claim 2 , wherein before the separately performing weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image, the method further comprises:

separately generating at least two feature maps based on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image, wherein a pixel value of each pixel in each feature map is a corresponding parameter of the corresponding pixel; and

the separately performing weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image comprises:

performing weighted summation on pixel values of each pixel in the at least two feature maps to determine the feature value of each pixel in the first image.

4. The image-text fusion method according to claim 1 , wherein the determining of the plurality of first layout formats of the first text in the first image based on the first text and the feature value of each pixel in the first image further comprises:

determining the plurality of first layout formats based on the feature value of each pixel in the first image and a size of a text box of the first text when the first text is laid out by using one or more text templates.

5. The image-text fusion method according to claim 4 , further comprising:

obtaining the one or more text templates, wherein each of the one or more text templates specifies at least one of a line spacing, a line width, a font size, a font, a character thickness, an alignment mode, a decorative line position, and a decorative line thickness of a text.

6. The image-text fusion method according to claim 1 , wherein determining of the second layout format from the plurality of first layout formats based on the cost parameters of the plurality of first layout formats further comprises:

determining a texture feature parameter of an image region that is in the first image and is blocked by the text box of the first text when the first text is laid out in the first image based on the plurality of first layout formats separately, wherein the texture feature parameter is used to represent a quantity of texture features corresponding to the image region in the image;

selecting, from the plurality of first layout formats, a plurality of first layout formats corresponding to image regions whose texture feature parameters are less than a second threshold; and

determining the second layout format from the plurality of selected first layout formats based on a cost parameter of each selected first layout format.

7. The image-text fusion method according to claim 1 , further comprising:

performing at least two of step a, step b, and step c, and step d for each of the plurality of first layout formats, to obtain a cost parameter of each first layout format:

step a: calculating a text intrusion parameter of the first text when the first text is laid out in the first image based on a first layout format, wherein the text intrusion parameter is a ratio of a first parameter to a second parameter, wherein the first parameter is a sum of feature values of pixels in an image region blocked by the first text in the first image, and the second parameter is an area of the image region, or the second parameter is a total quantity of pixels in the image region, or the second parameter is a product of a total quantity of pixels in the image region and a preset value;

step b: calculating a visual space occupation parameter of the first text when the first text is laid out in the first image based on the first layout format, wherein the visual space occupation parameter is used to represent a proportion of pixels whose feature values are less than a third threshold in the image region;

step c: calculating a visual balance parameter of the first text when the first text is laid out in the first image based on the first layout format, wherein the visual balance parameter is used to represent a degree of impact of the first text on the balance degree of feature value distribution of pixels in each region in the first image in which the first text is laid out; and

step d: calculating a cost parameter of the first layout format based on at least two of the calculated text intrusion parameter, the visual space occupation parameter, and the visual balance parameter of the first text.

8. The image-text fusion method according to claim 7 , wherein calculating of the cost parameter of the first layout format based on at least two of the calculated text intrusion parameter, the visual space occupation parameter, and the visual balance parameter of the first text further comprises:

using

T i =λ 1 *E s ( L i )+λ 2 *E u ( L i )+λ 3 *E n ( L i ); or

T i =(λ 1 *E s ( L i )+λ 2 *E u ( L i ))* E n ( L i ), or

T i =E s ( L i )* E u ( L i )* E n ( L i )

to calculate the cost parameter T i of the first layout format, wherein

E s (L i ) is the text intrusion parameter of the first text when the first text is laid out in the first image based on the first layout format, E u (L i ) is the visual space occupation parameter of the first text when the first text is laid out in the first image based on the first layout format, E n (L i ) is the visual balance parameter of the first text when the first text is laid out in the first image based on the first layout format, and λ 1 , λ 2 , and λ 3 are weight parameters corresponding to E s (L i ), and E u (L i ) E n (L i ).

9. The image-text fusion method according to claim 8 , wherein the determining of the second layout format from the plurality of first layout formats based on the cost parameters of the plurality of first layout formats further comprises:

determining that a first layout format corresponding to a smallest cost parameter among the cost parameters of the plurality of first layout formats is the second layout format.

10. The image-text fusion method according to claim 1 , further comprising:

determining a color parameter of the first text, wherein the color parameter of the first text is a derivative color of a dominant color of an image region blocked by the first text in the first image when the first text is laid out in the first image based on the second layout format, and the derivative color of the dominant color is a color having a same hue as the dominant color but having a tone, saturation, and brightness different from an HSV of the dominant color; and

coloring the first text in the second image based on the color parameter of the first text to obtain a third image.

11. The image-text fusion method according to claim 10 , wherein

the dominant color of the image region blocked by the first text in the first image is determined based on tones, saturation, and brightness of three primary colors RGB of the image region blocked by the first text in the first image in an HSV space when the first text is laid out in the first image based on the second layout format; and the dominant color is a hue with a highest proportion in the image region.

12. The image-text fusion method according to claim 1 , further comprising:

determining to perform rendering processing on the second image if at least one of the following condition 1 and condition 2 is met:

condition 1: a texture feature parameter of an image region that is in the first image and is blocked by the first text when the first text is laid out in the first image based on the second layout format is greater than a fourth threshold, wherein the texture feature parameter is used to represent a quantity of texture features corresponding to the image region in the image; and

condition 2: a proportion of a dominant color of the image region is less than a fifth threshold; and

covering the second image with a mask layer; or determining a mask parameter, and processing the second image based on the determined mask parameter; or performing projection rendering on the first text.

13. The image-text fusion method according to claim 10 , further comprising:

determining to perform rendering processing on the third image if at least one of the following condition 1 and condition 2 is met:

condition 1: a texture feature parameter of an image region that is in the first image and is blocked by the first text when the first text is laid out in the first image based on the second layout format is greater than a fourth threshold, wherein the texture feature parameter is used to represent a quantity of texture features corresponding to the image region in the image; and

condition 2: a proportion of a dominant color of the image region is less than a fifth threshold; and

covering the third image with a mask layer; or determining a mask parameter, and processing the third image based on the determined mask parameter; or performing projection rendering on the first text.

14. A non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processing circuit, the image-text fusion method according to claim 1 is implemented.

15. An image-text fusion apparatus comprising a memory configured to store one or more computer programs and a processor configured to execute the one or more computer programs stored in the memory, wherein the processor is configured to:

obtain a first image and a first text to be laid out in the first image;

determine a feature value of each pixel in the first image, wherein a feature value of a pixel is used to represent a probability that a user pays attention to the pixel, wherein the probability that the user pays attention to the pixel is higher for greater feature values of the pixel;

determine a plurality of first layout formats of the first text in the first image based on the first text and the feature value of each pixel in the first image, wherein when the first text is laid out in the first image based on each first layout format, the first text does not block a pixel whose feature value is greater than a first threshold; and

determine a second layout format from the plurality of first layout formats based on cost parameters of the plurality of first layout formats, wherein a cost parameter of a first layout format is used to represent a magnitude of a feature value of a pixel blocked by the first text when the first text is laid out in the first image based on the first layout format, and a balance degree of feature value distribution of pixels in each region in the first image in which the first text is laid out; and

lay out the first text in the first image based on the second layout format to obtain a second image.

16. The image-text fusion apparatus according to claim 15 , wherein the processor is further configured to determine at least two parameters of a visual saliency parameter, a face feature parameter, an edge feature parameter, and a text feature parameter of each pixel in the first image, wherein the visual saliency parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a visual saliency feature, the face feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a face, the edge feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to an object contour, and the text feature parameter of each pixel is used to represent a probability that the respective pixel is a pixel corresponding to a text; and

the processor is configured to separately performs weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image.

17. The image-text fusion apparatus according to claim 16 , wherein before the processor separately performs weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image, the processor is further configured to:

separately generate at least two feature maps based on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image, wherein a pixel value of each pixel in each feature map is a corresponding parameter of the corresponding pixel; and

the separately performing weighted summation on the determined at least two parameters of the visual saliency parameter, the face feature parameter, the edge feature parameter, and the text feature parameter of each pixel in the first image to determine the feature value of each pixel in the first image comprises:

performing weighted summation on pixel values of each pixel in the at least two feature maps to determine the feature value of each pixel in the first image.

18. The image-text fusion apparatus according to claim 15 , wherein the processor is further configured to determine the plurality of first layout formats based on the feature value of each pixel in the first image and a size of a text box of the first text when the first text is laid out by using one or more text templates.

19. The image-text fusion apparatus according to claim 18 , wherein the processor is further configured to:

obtain the one or more text templates, wherein each of the one or more text templates specifies at least one of a line spacing, a line width, a font size, a font, a character thickness, an alignment mode, a decorative line position, and a decorative line thickness of a text.

20. The image-text fusion apparatus according to claim 15 , wherein

the processor is further configured to determine a texture feature parameter of an image region that is in the first image and is blocked by the text box of the first text when the first text is laid out in the first image based on the plurality of first layout formats separately, wherein the texture feature parameter is used to represent a quantity of texture features corresponding to the image region in the image;

the processor is configured to select, from the plurality of first layout formats, a plurality of first layout formats corresponding to image regions whose texture feature parameters are less than a second threshold; and

the processor is configured to determine the second layout format from the plurality of selected first layout formats based on a cost parameter of each selected first layout format.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2022
From: ZHANG, WENJIE; ZHONG, WEICAI; HU, LIANG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 061047/0202 →
Priority Claims (1)
CN 201910783866.7 · Aug 23, 2019 · national
Continuity (1)
Related Publication 20220319077A1 · Oct 6, 2022
References Cited (20)
US 11189066B1 · Bylinskii · 2021 [cited by examiner]
US 20110173532A1 · Forman et al. · 2011 [cited by applicant]
US 20160093059A1 · Tumanov · 2016 [cited by examiner]
US 20160180161A1 · Novak et al. · 2016 [cited by applicant]
CN 101123002A · 2008 [cited by applicant]
CN 102890826A · 2013 [cited by applicant]
CN 107103635A · 2017 [cited by applicant]
CN 109117713A · 2019 [cited by applicant]
CN 109493399A · 2019 [cited by applicant]
CN 109643222A · 2019 [cited by applicant]
CN 110009712A · 2019 [cited by applicant]
CN 110706310A · 2020 [cited by applicant]
WO 2013005366A1 · 2013 [cited by applicant]
Yang et al., “Automatic Generation of Visual-Textual Presentation Layout”, ACM Transactions on Multimedia Computing Communications and Applications, Association for Computing Machinery, US, vol. 12, No. 2, Feb. 9, 2016,… [cited by applicant]
Malu et al., “An Approach to Optimal Text Placement on Images”, Jul. 21, 2013, arxiv. org, 7 pages. [cited by applicant]
Jia et al., “Image-Based Label Placement for Augmented Reality Browsers”, 2018 IEEE 4th International Conference on Computer and Communications, Dec. 7, 2018, 6 pages. [cited by applicant]
Grasset et al., “Image-Driven View Management for Augmented Reality Browsers”, 2012, IEEE International Symposium on Mixed and Augmented Reality, IEEE, Nov. 5, 2012, 10 pages. [cited by applicant]
Rosten et al., “Real-Time Video Annotations for Augmented Reality”, In: “Advances in Databases and Information Systems”, Jan. 1, 2005, Springer International Publishing, Cham 032682, vol. 3804, 9 pages. [cited by applicant]
Tanaka et al., “An information Layout Method for an Optical See-through Head Mounted Display Focusing on the Viewability”, 2008. ISMAR2008. 7th IEEE/ACM International Symposium on Mixed and Augmented Reality, IEEE, Pisc… [cited by applicant]
Rakholia et al., “Where to Place: A Real-Time Visual Saliency Based Label Placement for Augmented Reality Applications”, 2018, 25th IEEE International Conference on Image Processing, IEEE, Oct. 7, 2018, 5 pages. [cited by applicant]