IP Library › Granted Patent US 12,541,898
Granted Patent B2
US 12,541,898 · App. 18/649,466 · Granted Feb 3, 2026

Generative artificial intelligence visual effect generator

Inventors: Haifeng Gong (Fremont, CA); Ruoting Wan (Sunnyvale, CA); Dongdong Wang (Sunnyvale, CA); Yicong Tian (Mountain View, CA); Hang Qi (Mountain View, CA)
Assignee: Google LLC
G06T11/60
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,898
App. No.
18/649,466
Granted
Feb 3, 2026
Kind
B2
Abstract

A method and system for enhancing text with image-based visual effects is presented. An initial text prompt including a phrase for display is submitted to a language model. The language model outputs a candidate prompt. The candidate prompt may be further modified to create an image prompt. The image prompt is submitted to an image generating model which produces an image. A visual effect of the display phrase based on the image is displayed.

Claims (50)

1 . A method of visually enhancing text comprising:

receiving, by one or more processors, a set of text including a display phrase;

generating, by the one or more processors, a candidate prompt based on the set of text, wherein the candidate prompt differs from the set of text;

converting, by the one or more processors, the candidate prompt into an image prompt by removing one or more words from the candidate prompt;

submitting the image prompt into an image generating model to generate an image; and

rendering the display phrase with a visual effect of the generated image.

2 . The method of claim 1 , further comprising:

placing the display phrase in a template;

adding other information from the set of text to complete the template; and

using the completed template as the candidate prompt.

3 . The method of claim 1 , wherein converting the candidate prompt into the image prompt comprises:

submitting the candidate prompt to a large language model configured to produce, using the candidate prompt, an output phrase that excludes the one or more words; and

creating the image prompt based on the output phrase.

4 . The method of claim 1 , wherein rendering the display phrase with a visual effect comprises rendering the display phrase as a transparent stencil over the generated image.

5 . The method of claim 1 , wherein rendering the display phrase with a visual effect comprises imprinting the display phrase into the generated image.

6 . The method of claim 1 , wherein rendering the display phrase with a visual effect comprises both modifying the display phrase and modifying a background using the generated image.

7 . The method of claim 1 , further comprising:

generating a plurality of images prior to receiving the image prompt;

upon receipt of the image prompt, matching the image prompt with an image of the plurality of images by determining a similarity between an embedding of each of the plurality of images and an embedding of the image prompt, and returning images with a similarity greater than a threshold value.

8 . The method of claim 7 , wherein the plurality of generated images with text visual effects are ranked based on a visual appeal score.

9 . The method of claim 1 , wherein rendering the display phrase with a visual effect comprises:

forming the display phrase into a mask delineated by boundaries; and

generating the image within the boundaries of the mask.

10 . The method of claim 9 , further comprising selecting a font for the display phrase, wherein the mask boundaries are defined by a plurality of non-overlapping same-sized patches at a plurality of locations.

11 . An artificial intelligence system comprising:

one or more memory devices; and

one or more processors configured to execute code including a set of instructions, wherein execution of the set of instructions causes the one or more processors to perform operations comprising:

receiving, by the one or more processors, a set of text including a display phrase;

generating, by the one or more processors, a candidate prompt based on the set of text, wherein the candidate prompt differs from the set of text;

converting, by the one or more processors, the candidate prompt into an image prompt by removing one or more words from the candidate prompt;

submitting the image prompt into an image generating model to generate an image; and

rendering the display phrase with a visual effect of the generated image.

12 . The system of claim 11 , wherein the one or more processors further perform operations comprising:

placing the display phrase in a template;

adding other information from the set of text to complete the template; and

using the completed template as the candidate prompt.

13 . The system of claim 11 , wherein converting the candidate prompt into the image prompt comprises:

submitting the candidate prompt to a large language model configured to produce, using the candidate prompt, an output phrase that excludes the one or more words; and

creating the image prompt based on the output phrase.

14 . The system of claim 11 , wherein rendering the display phrase with a visual effect comprises rendering the display phrase as a transparent stencil over the generated image.

15 . The system of claim 11 , wherein rendering the display phrase with a visual effect comprises imprinting the display phrase into the generated image.

16 . The system of claim 11 , wherein rendering the display phrase with a visual effect comprises both modifying the display phrase and modifying a background using the generated image.

17 . The system of claim 11 , wherein the one or more processors further perform operations comprising:

generating a plurality of images prior to receiving the image prompt;

upon receipt of the image prompt, matching the image prompt with an image of the plurality of images by determining a similarity between an embedding of each of the plurality of images and an embedding of the image prompt, and returning images with a similarity greater than a threshold value.

18 . The system of claim 17 , wherein the plurality of generated images with text visual effects are ranked based on a visual appeal score.

19 . The system of claim 11 , wherein rendering the display phrase with a visual effect comprises:

forming the display phrase into a mask delineated by boundaries; and

generating the image within the boundaries of the mask.

20 . The system of claim 19 , wherein the one or more processors further perform operations comprising selecting a font for the display phrase, wherein the mask boundaries are defined by a plurality of non-overlapping same-sized patches at a plurality of locations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2024
From: GONG, HAIFENG; WAN, RUOTING; WANG, DONGDONG; TIAN, YICONG; QI, HANG
To: GOOGLE LLC
Reel/Frame 067373/0006 →
Continuity (1)
Related Publication 20250336122A1 · Oct 30, 2025
References Cited (5)
US 10235349B2 · Zhang et al. · 2019 [cited by applicant]
US 10949978B2 · Holzer et al. · 2021 [cited by applicant]
US 11212464B2 · Knorr et al. · 2021 [cited by applicant]
US 20240095275A1 · Tambi · 2024 [cited by examiner]
US 20240338870A1 · Iyer · 2024 [cited by examiner]