IP Library › Granted Patent US 12,432,307
Granted Patent B2
US 12,432,307 · App. 17/984,392 · Granted Sep 30, 2025

Personalized semantic based image editing

Inventors: Evelyn Chee (Singapore, SG); Yubo Duan (Singapore, SG)
Assignee: Black Sesame Technologies Inc.
H04N1/387G10L15/1815G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,432,307
App. No.
17/984,392
Granted
Sep 30, 2025
Kind
B2
Abstract

This present invention discloses a personalized and user specific system for editing an image. The system utilizes voice commands that are supplemented by semantic learning to form semantic descriptors. In addition, the system also utilizes user preferences to term an edited image. The system is utilized especially by novice photo editors, to provide a desired filter effect to the image with relative ease.

Claims (59)

1. An image editing system comprising:

an imaging module to receive an image;

a speech module to receive a natural language speech from a user;

a speech extractor for extracting context from the received natural language speech;

an encoder-decoder for extracting a semantic feature from the context and for generating a tag based on the semantic feature, wherein the tag determines a filter to be applied based on natural language descriptors derived from the speech commands including voice instructions with emotions and to match the user's current mood; and

a filtering module comprising:

a first filter, wherein the first filter is extracted based on the tag and applied on the image to form a filtered image;

a second filter, wherein the second filter is applied on the filtered image based on one or more user preferences to generate an edited image for the purpose of updating the image; and

wherein the encoder-decoder passes the voice or typed phrases from the user and input image to their respective encoders and the encoded features are jointly processed by another neural network block before being passed to the decoder to generate the final output.

2. The system of claim 1 , wherein the imaging module is a lens or a camera.

3. The system of claim 1 , wherein the speech module is a microphone.

4. The system of claim 1 , wherein the semantic feature is a word or text descriptor.

5. The system of claim 1 , wherein the first filter is extracted from an image database.

6. The system of claim 5 , wherein the image database is from a web server, a social networking site, or a web page.

7. The system of claim 5 , wherein the image database is screened according to the tag and an aesthetic value.

8. The system of claim 1 , wherein the user preferences are an editing style data or an image quality data associated with the image.

9. The system of claim 8 , wherein the image quality data is saturation, contrast, size, dimension, or sharpness.

10. The system of claim 8 , wherein the editing style data is brightness, color hue, or color palettes.

11. The system of claim 8 , wherein the user preferences are collected from a questionnaire, a web history, or a user usage history.

12. The system of claim 1 , wherein the first filter and the second filter are based on an artificial intelligence based learning model that allows collecting, clustering, and accessing.

13. The system of claim 12 , wherein the artificial intelligence based learning model is a neural network module, a machine-learning module, or a deep convolutional neural network.

14. An image editing system comprising:

an image captured from an electronic device;

an imaging module to receive the image;

an input module to receive a textual input from a user;

an encoder-decoder for extracting a semantic feature from the textual input and for generating a tag based on the semantic feature, wherein the tag determines a filter to be applied based on natural language descriptors derived from the speech commands including voice instructions with emotions and to match the user's current mood;

a filtering module comprising:

a first filter, wherein the first filter is extracted by using the tag and applied on the image to form a filtered image; and

a second filter, wherein the second filter is applied on the filtered image based on one or more user preferences to generate an edited image for the purpose of updating the image;

wherein the encoder-decoder passes the voice or typed phrases from the user and input image to their respective encoders and the encoded features are jointly processed by another neural network block before being passed to the decoder to generate the final output; and

wherein application of each of the first filter and the second filter is based on an artificial intelligence based model.

15. The system of claim 14 , wherein the electronic device is a digital camera, a PDA, a mobile device, a tablet, or a laptop.

16. An image editing system comprising:

an imaging module to receive an image;

a speech module to receive speech from a user;

a speech extractor for extracting context from the speech;

an encoder-decoder for extracting a semantic feature from the context and for generating a tag based on the semantic feature, wherein the tag includes details about determines a filter to be applied based on natural language descriptors derived from the speech commands including voice instructions with emotions and to match the user's current mood; and

a filtering module comprising:

a first filter, wherein the first filter is extracted by using the tag and applied on the image to form a filtered image; and

a second filter, wherein the second filter is applied on the filtered image based on one or more user preferences to generate an edited image for the purpose of updating the image;

wherein the encoder-decoder passes the voice or typed phrases from the user and input image to their respective encoders and the encoded features are jointly processed by another neural network block before being passed to the decoder to generate the final output; and

wherein application of each of the first filter and the second filter is based on an artificial intelligence based model.

17. An image editing method comprising:

capturing an image;

receiving speech from a user;

extracting context from the speech;

extracting a semantic feature from the context;

generating a tag based on the semantic feature, wherein the tag determines a filter to be applied based on natural language descriptors derived from the speech commands including voice instructions with emotions and to match the user's current mood;

extracting a first filter based on the tag;

applying a first filter on the image to generate a filtered image; and

applying a second filter on the filtered image based on one or more user preferences to form an edited image for the purpose of updating the image.

18. An image editing method comprising:

capturing an image;

receiving a textual input from a user;

extracting a semantic feature from the textual input;

generating a tag based on the semantic feature, wherein the tag determines a filter to be applied based on natural language descriptors derived from the speech commands including voice instructions with emotions and to match the user's current mood;

extracting a first filter based on the tag;

applying a first filter on the image to generate a filtered image; and

applying a second filter on the filtered image based on one or more user preferences to form an edited image for the purpose of updating the image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2022
From: CHEE, EVELYN; DUAN, YUBO
To: BLACK SESAME TECHNOLOGIES INC.
Reel/Frame 061716/0968 →
Continuity (1)
Related Publication 20240163388A1 · May 16, 2024
References Cited (14)
US 5990901A · Lawton et al. · 1999 [cited by applicant]
US 6834124B1 · Lin et al. · 2004 [cited by applicant]
US 9407816B1 · Sehn · 2016 [cited by applicant]
US 9569697B1 · McNerney et al. · 2017 [cited by applicant]
US 9754355B2 · Chang et al. · 2017 [cited by applicant]
US 10776860B2 · Sakamoto · 2020 [cited by examiner]
US 11099469B1 · Selfe · 2021 [cited by examiner]
US 20130346920A1 · Morris · 2013 [cited by examiner]
US 20160093032A1 · Lei · 2016 [cited by examiner]
US 20190005034A1 · Ben-Yair · 2019 [cited by examiner]
US 20210133931A1 · Lee · 2021 [cited by examiner]
US 20210174193A1 · Pouran Ben Veyseh · 2021 [cited by examiner]
US 20230126177A1 · Xu · 2023 [cited by examiner]
US 20230283849A1 · Doshi · 2023 [cited by examiner]