IP Library › Granted Patent US 12,725,321
Granted Patent B2
US 12,725,321 · App. 18/457,521 · Granted Sep 1, 2026

Control font generation consistency

Inventors: Li Chen (Redmond, WA); Ji Li (San Jose, CA)
Assignee: Microsoft Technology Licensing, LLC
G06T11/10G06F40/40G06T5/70G06T7/50G06T2200/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,321
App. No.
18/457,521
Filed
Aug 29, 2023
Granted
Sep 1, 2026
Kind
B2
Art Unit
2618
USPC
345/467
Abstract

Systems and methods for generating custom art fonts with consistent style include receiving user input that identifies a base font style for a custom font and includes descriptive text that defies one or more text effects to use for the custom font. Depth maps are selected for characters to be included in the custom font. The depth maps are preprocessed to add noise to the depth maps. A generative model generates custom font images conditioned with the text prompt and the depth maps. The custom font images are then used to render text on a display screen of a computing device.

Claims (44)

1 . A font generation system comprising:

a processor; and

a memory in communication with the processor, the memory comprising executable instructions that, when executed by the processor alone or in combination with other processors, cause the font generation system to perform functions of:

receiving user input that identifies a base font style for a custom font and includes descriptive text that describes in a natural language format one or more text effects to use for the custom font;

selecting depth maps for characters to be included in the custom font, each depth map including an image of one of the characters for the custom font;

preprocessing the depth maps for the custom font using a predetermined function that adds noise to at least a character portion of each of the depth maps; and

providing the descriptive text and the preprocessed depth maps to a generative image model, the descriptive text being provided to the generative image model as a text prompt, the generative image model being trained to generate a custom font output image for each character to be included in the custom font conditioned by the text prompt and the preprocessed depth map associated with the character; and

receiving the custom font output images for each character included in the custom font from the generative image model and utilizing the custom font output images to render text on a display screen of a computing device.

2 . The font generation system of claim 1 , wherein the predetermined function for adding noise to the depth maps includes at least one of a Gaussian noise function, a point noise function, a Perlin noise function, and a texture effect function.

3 . The font generation system of claim 2 , wherein the generative image model is a latent diffusion model having a text encoder which generates text embeddings from the descriptive text, a noise predictor which is trained to perform a denoising process for each of the characters included in the custom font to generate a latent output image based on the text embeddings and the preprocessed depth maps associated with each of the characters of the custom font, and a text decoder which converts a latent output image in a latent space for each of the characters to a custom font image in a pixel space for each of the characters of the custom font.

4 . The font generation system of claim 3 , wherein the denoising process for a character in the custom font includes performing a predetermined number of denoising steps on an input latent image to generate a conditioned latent image based on the text embeddings and the preprocessed depth map for the character, and

wherein, for each of the denoising steps, the noise predictor predicts an amount of noise in the input latent image that should be subtracted to arrive at a desired custom font image for the character, the amount of noise being subtracted from the input latent image to generate a conditioned latent image, the conditioned latent image of a last denoising step corresponding to the latent output image.

5 . The font generation system of claim 1 , further comprising:

using a prompt engineering component to generate the text prompt from the descriptive text using a prompt engineering scheme for automatically generating the text prompt which takes into consideration at least one of a type of generative model used and a desired output of the generative image model and automatically generates the text prompt from the descriptive text by adding text, deleting text, replacing text, and/or formatting text.

6 . The font generation system of claim 1 , wherein the depth map includes a character portion and a background image portion, the character portion being depicted in a first grayscale shade and the background portion being depicted in a second grayscale shade.

7 . The font generation system of claim 1 , wherein the user input is received via a user interface of the font generation system, the user interface including user interface controls for receiving a font style selection designating the base font style, for receiving the descriptive text and for displaying the custom font output images.

8 . The font generation system of claim 1 , further comprising:

performing a postprocessing operation to remove a background from the custom font images.

9 . A method for generating custom art fonts using a generative image model, the method comprising:

receiving user input that identifies a base font style for a custom font and includes descriptive text that describes in a natural language format one or more text effects to use for the custom font;

selecting depth maps for characters to be included in the custom font, each depth map including an image of one of the characters for the custom font;

preprocessing the depth maps for the custom font using a predetermined function that adds noise to at least a character portion of each of the depth maps; and

providing the descriptive text and the preprocessed depth maps to the generative image model, the descriptive text being provided to the generative image model as a text prompt, the generative image model being trained to generate a custom font output image for each character to be included in the custom font conditioned by the text prompt and the preprocessed depth map associated with the character; and

receiving the custom font output images for each character included in the custom font from the generative image model and utilizing the custom font output images to render text on a display screen of a computing device.

10 . The method of claim 9 , wherein the predetermined function for adding noise to the depth maps includes at least one of a Gaussian noise function, a point noise function, a Perlin noise function, and a texture effect function.

11 . The method of claim 10 , wherein the generative image model is a latent diffusion model having a text encoder which generates text embeddings from the descriptive text, a noise predictor which is trained to perform a denoising process for each of the characters included in the custom font to generate a latent output image based on the text embeddings and the preprocessed depth maps associated with each of the characters of the custom font, and a text decoder which converts a latent output image in a latent space for each of the characters to a custom font image in a pixel space for each of the characters of the custom font.

12 . The method of claim 11 , wherein the denoising process for a character in the custom font includes performing a predetermined number of denoising steps on an input latent image to generate a conditioned latent image based on the text embeddings and the preprocessed depth map for the character, and

wherein, for each of the denoising steps, the noise predictor predicts an amount of noise in the input latent image that should be subtracted to arrive at a desired custom font image for the character, the amount of noise being subtracted from the input latent image to generate a conditioned latent image, the conditioned latent image of a last denoising step corresponding to the latent output image.

13 . The method of claim 9 , further comprising:

using a prompt engineering component to generate the text prompt from the descriptive text using a prompt engineering scheme for automatically generating the text prompt which takes into consideration at least one of a type of generative model used and a desired output of the generative image model and automatically generates the text prompt from the descriptive text by adding text, deleting text, replacing text, and/or formatting text.

14 . The method of claim 9 , wherein the depth map includes a character portion and a background image portion, the character portion being depicted in a first grayscale shade and the background portion being depicted in a second grayscale shade.

15 . The method of claim 9 , wherein the user input is received via a user interface of a font generation system, the user interface including user interface controls for receiving a font style selection designating the base font style, for receiving the descriptive text and for displaying the custom font output images.

16 . The method of claim 9 , further comprising:

performing a postprocessing operation to remove a background from the custom font images.

17 . A non-transitory computer readable medium on which are stored instructions that, when executed, cause a programmable device to perform functions of:

receiving user input that identifies a base font style for a custom font and includes descriptive text that describes in a natural language format one or more text effects to use for the custom font;

selecting depth maps for characters to be included in the custom font, each depth map including an image of one of the characters for the custom font;

preprocessing the depth maps for the custom font using a predetermined function that adds noise to at least a character portion of each of the depth maps; and

providing the descriptive text and the preprocessed depth maps to a generative image model, the descriptive text being provided to the generative image model as a text prompt, the generative image model being trained to generate a custom font output image for each character to be included in the custom font conditioned by the text prompt and the preprocessed depth map associated with the character; and

receiving the custom font output images for each character included in the custom font from the generative image model and utilizing the custom font output images to render text on a display screen of a computing device.

18 . The computer readable medium of claim 17 , wherein the predetermined function for adding noise to the depth maps includes at least one of a Gaussian noise function, a point noise function, a Perlin noise function, and a texture effect function.

19 . The computer readable medium of claim 18 , wherein the generative image model is a latent diffusion model having a text encoder which generates text embeddings from the descriptive text, a noise predictor which is trained to perform a denoising process for each of the characters included in the custom font to generate a latent output image based on the text embeddings and the preprocessed depth maps associated with each of the characters of the custom font, and a text decoder which converts a latent output image in a latent space for each of the characters to a custom font image in a pixel space for each of the characters of the custom font.

20 . The computer readable medium of claim 19 , wherein the denoising process for a character in the custom font includes performing a predetermined number of denoising steps on an input latent image to generate a conditioned latent image based on the text embeddings and the preprocessed depth map for the character, and

Wherein, for each of the denoising steps, the noise predictor predicts an amount of noise in the input latent image that should be subtracted to arrive at a desired custom font image for the character, the amount of noise being subtracted from the input latent image to generate a conditioned latent image, the conditioned latent image of a last denoising step corresponding to the latent output image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: CHEN, LI; LI, JI
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 064733/0680 →
Continuity (1)
Related Publication 20250078343A1 · Mar 6, 2025
References Cited (2)
US 20240127510A1 · Darabi · 2024 [cited by examiner]
US 20240355018A1 · Aggarwal · 2024 [cited by examiner]