IP Library › Granted Patent US 11,257,491
Granted Patent B2
US 11,257,491 · App. 16/205,126 · Granted Feb 22, 2022

Voice interaction for image editing

Inventors: Sarah Kong (Cupertino, CA); Yinglan Ma (San Jose, CA); Hyunghwan Byun (San Jose, CA); Chih-Yao Hsieh (San Jose, CA)
Assignee: ADOBE INC.
G10L15/22G06F3/04845G06F3/167G06T11/60G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,257,491
App. No.
16/205,126
Granted
Feb 22, 2022
Kind
B2
Abstract

This application relates generally to modifying visual data based on audio commands and more specifically, to performing complex operations that modify visual data based on one or more audio commands. In some embodiments, a computer system may receive an audio input and identify an audio command based on the audio input. The audio command may be mapped to one or more operations capable of being performed by a multimedia editing application. The computer system may perform the one or more operations to edit to received multimedia data.

Claims (97)

1. A computer-implemented method for editing an image based on voice interaction comprising:

receiving an image using a photo editing application;

identifying a first label for a first segment within the image and a second label for a second segment within the image, wherein the first segment and the second segment are operable portions of the image;

receiving a voice input of a user;

matching (a) a first portion of the voice input to the first label and not to the second label and (b) a second portion of the voice input to a vocal command;

identifying, based on the vocal command, a particular set of operations for modifying within the image, wherein the particular set of operations is previously mapped to the vocal command, wherein the particular set of operations comprises a plurality of sequential operations, wherein identifying the particular set of operations comprises:

determining a skill level of the user from a set of skill levels;

using a machine learning model, determining one or more sets of operations, each of the one or more sets of operations corresponding to a respective skill level of the set of skill levels; and

using the machine learning model, selecting the particular set of operations from the one or more sets of operations;

applying the particular set of operations to the first segment and not to the second segment to generate a modified image, wherein applying the particular set of operations to the first segment comprises:

identifying a first operation of the particular set of operations;

executing the first operation of the particular set of operations;

while executing the first operation of the particular set of operations, causing a first visual instruction associated with the first operation to be displayed in a graphical user interface, the first visual instruction including a visual indicator that instructs the user how to manually perform the first operation; and

causing the modified image to be displayed in the graphical user interface.

2. The computer-implemented method of claim 1 , further comprising:

receiving a subsequent voice input;

determining, based on the subsequent voice input, a subsequent vocal command;

identifying, based on the subsequent vocal command, a subsequent set of operations for modifying one or more segments within the image; and

modifying, based on the subsequent set of operations, the one or more segments within the modified image, wherein the particular set of operations and the subsequent set of operations modify at least a common segment within the image.

3. The computer-implemented method of claim 1 , further comprising:

storing, in an operation log, each completed operation of the particular set of operations; and

causing the operation log to be displayed in the graphical user interface.

4. The computer-implemented method of claim 1 , wherein the particular set of operations comprises the first operation, a second operation, and a third operation, wherein:

the first operation comprises instructions for identifying, within the image, a first object and a first location of the first object within the image;

the second operation comprises instructions for removing the first object from the first location; and

the third operation comprises instructions for modifying the first location based at least on a background within the image, wherein the background is a segment of the image.

5. The computer-implemented method of claim 4 , wherein the particular set of operations further comprises a fourth operation, wherein the fourth operation comprises instructions for placing the first object at a second location, wherein the second location is based on the vocal command.

6. The computer-implemented method of claim 1 , wherein applying the particular set of operations to the first segment further comprises:

identifying a second operation of the particular set of operations, the second operation being sequentially after the first operation;

executing the second operation of the particular set of operations;

while executing the second operation of the particular set of operations, causing a second visual instruction associated with the second operation to be displayed in the graphical user interface, the second visual instruction comprising visual indicators that instruct the user how to manually perform the second operation.

7. A non-transitory computer-readable storage medium having stored thereon instructions for causing at least one computer system to edit an image based on voice interaction, the instructions comprising:

receiving an image using a photo editing application;

identifying a first label for a first segment within the image and a second label for a second segment within the image, wherein the first segment and the second segment are operable portions of the image;

receiving a voice input of a user;

matching (a) a first portion of the voice input to the first label and not to the second label and (b) a second portion of the voice input to a vocal command;

identifying, based on the vocal command, a particular set of operations for modifying within the image, wherein the particular set of operations is previously mapped to the vocal command, wherein the particular set of operations comprises a plurality of sequential operations, wherein identifying the particular set of operations comprises:

determining a skill level of the user from a set of skill levels;

using a machine learning model, determining one or more sets of operations, each of the one or more sets of operations corresponding to a respective skill level of the set of skill levels:

using the machine learning model, selecting the particular set of operations from the one or more sets of operations;

applying the particular set of operations to the first segment and not to the second segment to generate a modified image, wherein applying the set of operations to the first segment comprises:

identifying a first operation of the particular set of operations;

executing the first operation of the particular set of operations;

while executing the first operation of the particular set of operations, causing a first visual instruction associated with the first operation to be displayed in a graphical user interface, the first visual instruction including a visual indicator that instructs the user how to manually perform the first operation; and

causing the modified image to be displayed in the graphical user interface.

8. The computer-readable storage medium of claim 7 , the instructions further comprising:

receiving a subsequent voice input;

determining, based on the subsequent voice input, a subsequent vocal command;

identifying, based on the subsequent vocal command, a subsequent set of operations for modifying one or more segments within the image; and

modifying, based on the subsequent set of operations, the one or more segments within the modified image, wherein the particular set of operations and the subsequent set of operations modify at least a common segment within the image.

9. The computer-readable storage medium of claim 7 , the instructions further comprising:

storing, in an operation log, each completed operation of the particular set of operations; and

causing the operation log to be displayed in the graphical user interface.

10. The computer-readable storage medium of claim 7 , wherein the particular set of operations comprises the first operation, a second operation, and a third operation, wherein:

the first operation comprises instructions for identifying, within the image, a first object and a first location of the first object;

the second operation comprises instructions for removing the first object from the first location; and

the third operation comprises instructions for modifying the first location based at least on a background within the image, wherein the background is a segment of the image.

11. The computer-readable storage medium of claim 10 , wherein the particular set of operations further comprises a fourth operation, wherein the fourth operation comprises instructions for placing the first object at a second location, wherein the second location is based on the vocal command.

12. The computer-readable storage medium of claim 7 , wherein applying the particular set of operations to the first segment further comprises:

identifying a second operation of the set of operations, the second operation being sequentially after the first operation;

executing the second operation of the particular set of operations;

while executing the second operation of the particular set of operations, causing a second visual instruction associated with the second operation to be displayed in the graphical user interface, the second visual instruction comprising visual indicators that instruct the user how to manually perform the second operation.

13. A system for editing an image based on voice interaction, comprising:

one or more processors; and

a memory coupled with the one or more processors, the memory configured to store instructions that when executed by the one or more processors cause the one or more processors to:

receive an image using a photo editing application;

identify a first label for a first segment within the image and a second label for a second segment within the image, wherein the first segment and the second segment are operable portions of the image;

receive a voice input of a user;

match (a) a first portion of the voice input to the first label and not to the second label and (b) a second portion of the voice input to a vocal command;

identify, based on the vocal command, a particular set of operations for modifying within the image, wherein the particular set of operations is previously mapped to the vocal command, wherein the particular set of operations comprises a plurality of sequential operations, wherein identifying the particular set of operations comprises:

determining a skill level of the user from a set of skill levels;

using a machine learning model, determining one or more sets of operations, each of the one or more sets of operations corresponding to a respective skill level of the set of skill levels; and

using the machine learning model, selecting the particular set of operations from the one or more sets of operations;

apply the particular set of operations to the first segment and not to the second segment to generate a modified image, wherein applying the particular set of operations to the first segment comprises:

identifying a first operation of the particular set of operations;

executing the first operation of the particular set of operations;

while executing the first operation of the particular set of operations, causing a first visual instruction associated with the first operation to be displayed in a graphical user interface, the first visual instruction including a visual indicator that instructs the user how to manually perform the first operation; and

cause the modified image to be displayed in the graphical user interface.

14. The system of claim 13 , wherein the instructions that when executed by the one or processors further cause the one or more processors to:

receive a subsequent voice input;

determine, based on the subsequent voice input, a subsequent vocal command;

identify, based on the subsequent vocal command, a subsequent set of operations for modifying one or more segments within the image; and

modify, based on the subsequent set of operations, the one or more segments within the modified image, wherein the particular set of operations and the subsequent set of operations modify at least a common segment within the image.

15. The system of claim 13 , wherein the instructions that when executed by the one or processors further cause the one or more processors to:

store, in an operation log, each completed operation of the particular set of operations; and

cause the operation log to be displayed in the graphical user interface.

16. The system of claim 13 , wherein the particular set of operations comprises the first operation, a second operation, and a third operation, wherein:

the first operation comprises instructions for identifying within the image, a first object and a first location of the first object;

the second operation comprises instructions for removing the first object from the first location; and

the third operation comprises instructions for modifying the first location based at least on a background within the image, wherein the background is a segment of the image.

17. The system of claim 16 , wherein the particular set of operations further comprises a fourth operation, wherein the fourth operation operations comprises instructions for placing the first object at a second location, wherein the second location is based on the vocal command.

18. The system of claim 13 , wherein applying the particular set of operations to the first segment further comprises:

identify a second operation of the particular set of operations, the second operation being sequentially after the first operation;

executing the second operation of the particular set of operations;

while executing the second operation of the particular set of operations, causing a second visual instruction associated with the second operation to be displayed in the graphical user interface, the second visual instruction comprising visual indicators that instruct the user how to manually perform the second operation.

19. The computer-implemented method of claim 1 , wherein the visual indicator is displayed at a location of the graphical user interface corresponding to an icon, wherein the user manually performs the first operation by selecting the icon.

20. The system of claim 13 , wherein the visual indicator is displayed at a location of the graphical user interface corresponding to an icon, wherein the user manually performs the first operation by selecting the icon.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2020
From: KONG, SARAH; MA, YINGLAN; BYUN, HYUNGHWAN; HSIEH, CHIH-YAO
To: ADOBE INC.
Reel/Frame 054309/0344 →
Continuity (1)
Related Publication 20200175975A1 · Jun 4, 2020