IP Library Granted Patent US 12688370
Granted Patent B2
US 12688370 · App. 18/387,671 · Granted Jul 21, 2026

Prompt based background replacement

Inventors: Yair Adato (Kfar Ben Nun, IL); Michael Feinstein (Tel Aviv, IL); Nimrod Sarid (Tel Aviv, IL); Ron Mokady (Ramat Hasaron, IL); Eyal Gutflaish (Beer Sheva, IL)
Assignee: BRIA ARTIFICIAL INTELLIGENCE LTD.
G06F40/30G06F40/279G06F40/40G06T7/194G06T7/70G06T11/10G06V10/764G06V10/774G10L15/063G10L15/18H04N5/272G06T2207/30196G10L2015/0631
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688370
App. No.
18/387,671
Filed
Nov 7, 2023
Granted
Jul 21, 2026
Kind
B2
Art Unit
2426
USPC
348/586
Abstract

Systems, methods and non-transitory computer readable media for prompt based background replacement are provided. A visual content including a background portion and at least one foreground object may be accessed. Further, a textual input indicative of a desire of an individual to modify the visual content may be received. The textual input and the visual content may be analyzed to generate a modified version of the visual content. The modified version may differ from the visual content in the background portion. Further, the modified version may include a depiction of the at least one foreground object substantially identical to a depiction of the at least one foreground object in the visual content. Further, a presentation of the modified version of the visual content to the individual may be caused.

Claims (51)

1 . A non-transitory computer readable medium storing a software program comprising data and computer implementable instructions that when executed by at least one processor cause the at least one processor to perform operations for prompt based background replacement, the operations comprising:

accessing a visual content including a background portion and at least one foreground object;

receiving a textual input indicative of a desire of an individual to modify the visual content;

analyzing the textual input and the visual content to generate a modified version of the visual content, the modified version differs from the visual content in the background portion, the modified version includes a depiction of the at least one foreground object object depicted in the visual content; and

causing a presentation of the modified version of the visual content to the individual, and wherein the operations comprise:

identifying a first mathematical object in a mathematical space, wherein the first mathematical object corresponds to a word of the textual input;

calculating a convolution of at least part of the visual content to obtain a numerical result value;

calculating a function of the first mathematical object and the numerical result value to identify a second mathematical object in the mathematical space; and

determining a pixel value of a pixel of the background portion of the modified version of the visual content based on the second mathematical object.

2 . The non-transitory computer readable medium of claim 1 , wherein the background portion of the visual content encloses the at least one foreground object in the visual content.

3 . The non-transitory computer readable medium of claim 1 , wherein the at least one foreground object includes at least one of a logo or a product.

4 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise analyzing the visual content to identify the background portion of the visual content.

5 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise analyzing the visual content to identify a portion of the visual content associated with the at least one foreground object.

6 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise analyzing the visual content to identify a category associated with the at least one foreground object, and wherein the modification of the background portion is based on the category associated with the at least one foreground object and the textual input.

7 . The non-transitory computer readable medium of claim 1 , wherein the at least one foreground object is a person, and wherein the modification of the background portion is based on the textual input and a demographic characteristic of the person.

8 . The non-transitory computer readable medium of claim 1 , wherein the at least one foreground object is an animal, and wherein the modification of the background portion is based on the textual input and a kind of the animal.

9 . The non-transitory computer readable medium of claim 1 , wherein a position and a spatial orientation of the at least one foreground object in the modified version of the visual content correspond to a position and a spatial orientation of the at least one foreground object in the visual content.

10 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise determining, based on the textual input, at least one of at least one of a position or a spatial orientation of the at least one foreground object in the modified version of the visual content.

11 . The non-transitory computer readable medium of claim 1 , wherein the at least part of the visual content includes at least one pixel of the at least one foreground object in the visual content and at least one pixel of the background portion of the visual content.

12 . The non-transitory computer readable medium of claim 1 , wherein the at least part of the visual content is entirely included in the at least one foreground object in the visual content.

13 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

identifying a third mathematical object in a mathematical space, wherein the third mathematical object corresponds to a word of the textual input;

calculating a convolution of at least part of a different visual content to obtain a second numerical result value;

calculating a function of the third mathematical object and the second numerical result value to identify a fourth mathematical object in the mathematical space; and

determining a pixel value of a second pixel of the background portion of the modified version of the visual content based on the fourth mathematical object.

14 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise using a machine learning model to analyze the textual input and the visual content to generate the modified version of the visual content.

15 . The non-transitory computer readable medium of claim 1 , wherein the operations further comprise:

analyzing the textual input and the visual content to generate a plurality of modified versions of the visual content, each specific modified version of the plurality of modified versions differs from the visual content in the background portion, and includes a respective depiction of the at least one foreground object depicted in the visual content; and

causing a presentation of the plurality of modified versions of the visual content to the individual.

16 . The non-transitory computer readable medium of claim 1 , wherein the textual input includes a noun and an adjective adjacent to the noun, wherein the at least one foreground object is identified based on the noun, and wherein the modification of the background portion is based on the adjective.

17 . The non-transitory computer readable medium of claim 1 , wherein the textual input includes a verb and an adverb adjacent to the verb, wherein the at least one foreground object is identified based on the verb, and wherein the modification of the background portion is based on the adverb.

18 . The non-transitory computer readable medium of claim 1 , wherein the at least part of the visual content is entirely included in the background portion of the visual content.

19 . A method for prompt based background replacement, the method comprising:

accessing a visual content including a background portion and at least one foreground object;

receiving a textual input indicative of a desire of an individual to modify the visual content;

analyzing the textual input and the visual content to generate a modified version of the visual content, the modified version differs from the visual content in the background portion, the modified version includes a depiction of the at least one foreground object depicted in the visual content; and

causing a presentation of the modified version of the visual content to the individual, and further comprising:

identifying a first mathematical object in a mathematical space, wherein the first mathematical object corresponds to a word of the textual input;

calculating a convolution of at least part of the visual content to obtain a numerical result value;

calculating a function of the first mathematical object and the numerical result value to identify a second mathematical object in the mathematical space; and

determining a pixel value of a pixel of the background portion of the modified version of the visual content based on the second mathematical object.

20 . A system for prompt based background replacement, the system comprising:

at least one processor configured to perform the operations of:

accessing a visual content including a background portion and at least one foreground object;

receiving a textual input indicative of a desire of an individual to modify the visual content;

analyzing the textual input and the visual content to generate a modified version of the visual content, the modified version differs from the visual content in the background portion, the modified version includes a depiction of the at least one foreground object depicted in the visual content; and

causing a presentation of the modified version of the visual content to the individual, and wherein the operations comprise:

identifying a first mathematical object in a mathematical space, wherein the first mathematical object corresponds to a word of the textual input;

calculating a convolution of at least part of the visual content to obtain a numerical result value;

calculating a function of the first mathematical object and the numerical result value to identify a second mathematical object in the mathematical space; and

determining a pixel value of a pixel of the background portion of the modified version of the visual content based on the second mathematical object.