IP Library Granted Patent US 12,154,021
Granted Patent B1
US 12,154,021 · App. 17/578,753 · Granted Nov 26, 2024

Visual search and content display system

Inventors: Josh Sternberg (Seattle, WA); Akshad Viswanathan (Seattle, WA); Xiaopeng Zhang (Bellevue, WA); Ethan Alexander Smith (Seattle, WA); Lenworth Richard Rose (Seattle, WA); Joyce Huan Fu (Seattle, WA); Mengyun Lv (Seattle, WA); Rui Chen (Bellevue, WA); Alexandru Indrei (Bothell, WA); Reece Dano (Seattle, WA); Saeed Salahi (Kirkland, WA); Anqi Liang (Seattle, WA)
Assignee: Amazon Technologies, Inc.
G06N3/0464G06F3/04817G06F3/0482G06F3/0485G06F16/2468G06F16/532G06F16/538G06F16/58G06F18/24G06N20/20G06Q30/0631G06Q30/0643G06N3/045G06N3/08G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,154,021
App. No.
17/578,753
Granted
Nov 26, 2024
Kind
B1
Abstract

Devices and techniques are generally described for providing a graphical user interface. In some examples, first image data may be sent to a computing device, that when rendered is effective to display a first combination of items disposed together in a first environment and a first selectable control. In some examples, first data may be received indicating a selection of the first selectable control. In various examples, a first computing device may determine first feature data representing the first image data. In some examples, the first computing device may determine second image data using the first feature data. In some examples, the second image data may be sent to a second computing device. The second image data, when rendered, may be effective to display a second combination of items disposed together in a second environment.

Claims (87)

1. A method comprising:

sending first instructions, from a first computing device to a second computing device, to cause first image data to be displayed, wherein the first image data, when rendered, is effective to display:

a first image depicting a first combination of items arranged together in a first environment; and

a first selectable control associated with the first image data;

receiving first data indicating a selection of the first selectable control;

determining, by the first computing device based at least in part on the receiving the first data indicating the selection of the first selectable control, first feature vector data associated with a machine learning feature space, wherein the first feature vector data is generated by an encoder and visually represents the first image data;

determining, by the first computing device, second image data using the first feature vector data based at least in part on the selection of the first selectable control;

sending second instructions, from the first computing device to the second computing device, to cause the second image data to be displayed by the second computing device based at least in part on the selection of the first selectable control, wherein the second image data, when rendered is effective to display a second image depicting a second combination of items arranged together in a second environment;

determining, using an object recognition machine learning model, a first subset of the second image data corresponding to a first item of the second combination of items and a second subset of the second image data corresponding to a second item of the second combination of items;

generating, by the object recognition machine learning model based at least in part on detection of the first subset of the second image data, a first feature embedding for the first subset of the second image data visually representing the first item; and

generating, by the object recognition machine learning model based at least in part on detection of the second subset of the second image data, a second feature embedding for the second subset of the second image data visually representing the second item.

2. The method of claim 1 , further comprising:

determining the second image data by determining a distance between the first feature vector data and second feature vector data associated with the second image data in the machine learning feature space of the first feature vector data and the second feature vector data; and

determining that the distance is a within a tolerance of a minimum distance among image data stored in a repository of image data.

3. The method of claim 1 , further comprising:

determining the second image data by determining a distance between the first feature vector data and second feature vector data associated with the second image data in the machine learning feature space of the first feature vector data and the second feature vector data; and

determining that the distance is within a tolerance of a maximum distance among image data stored in a repository of image data.

4. The method of claim 1 , further comprising:

receiving a selection of a set of attributes for items associated with the second image data;

receiving a control input effective to dismiss a third item of the second combination of items;

displaying a different item in place of the third item, wherein the different item includes the set of attributes; and

modifying the first feature vector data based at least in part on the different item.

5. The method of claim 4 , further comprising determining that second feature vector data representing the different item has a distance in the machine learning feature space that is within a tolerance of a maximum distance from a fourth feature vector data representing the third item among a set of items including the set of attributes.

6. The method of claim 1 , wherein the first image data, when rendered is effective to further display a first icon associated with a first item of the first combination of items, the method further comprising:

receiving data indicating a selection of the first icon; and

sending third image data that, when rendered, is effective to display an enlarged image of the first item.

7. The method of claim 1 , wherein the first image data, when rendered is effective to further display:

a first icon associated with a first item of the first combination of items;

a second icon associated with a second item of the first combination of items; and

scrollable enlarged images of each item of the first combination of items, wherein the scrollable enlarged images are displayed on a different part of the display relative to the first combination of items arranged together in the first environment;

the method further comprising:

receiving first data indicating a selection of the first item in the scrollable enlarged images; and

sending second data to the second computing device effective to cause the first icon to change from a first appearance to a second appearance.

8. The method of claim 7 , further comprising:

receiving third data indicating a scroll operation from the first item in the scrollable enlarged images to the second item in the scrollable enlarged images;

sending fourth data to the second computing device effective to cause the first icon to change from the second appearance to the first appearance; and

sending fifth data to the second computing device effective to cause the second icon to change from the first appearance to the second appearance.

9. The method of claim 1 , wherein the encoder comprises a convolutional neural network, the method further comprising:

determining, using the convolutional neural network, the first feature vector data visually representing the first combination of items as arranged together in the first environment;

determining, using a second convolutional neural network, second feature vector data visually representing a first item of the first combination of items, wherein the object recognition machine learning model comprises the second convolutional neural network;

storing the first feature vector data in at least one non-transitory computer-readable memory in association with the first image data; and

storing the second feature vector data in the at least one non-transitory computer-readable memory in association with third image data representing the first item.

10. The method of claim 1 , further comprising:

determining that an enlarged representation of the first item is currently displayed by the second computing device; and

modifying an appearance of a graphical tag associated with the first item in the second image data in response to the enlarged representation of the first item being currently displayed.

11. The method of claim 1 , wherein the first image depicts the first combination of items arranged together within a room.

12. A system, comprising:

at least one processor; and

at least one non-transitory computer-readable memory storing instructions that, when executed by the at least one processor, are effective to program the at least one processor to:

send first instructions to a first computing device, the first instructions effective to cause first image data to be displayed, wherein the first image data, when rendered, is effective to display:

a first image depicting a first combination of items arranged together in a first environment; and

a first selectable control associated with the first image data;

receive first data indicating a selection of the first selectable control;

determine, based at least in part on receipt of the first data indicating the selection of the first selectable control, first feature vector data associated with a machine learning feature space, wherein the first feature vector data is generated by an encoder and visually represents the first image data;

determine second image data using the first feature vector data based at least in part on the selection of the first selectable control;

send second instructions to the first computing device, the second instructions effective to cause the second image data to be displayed based at least in part on the selection of the first selectable control, wherein the second image data, when rendered is effective to display a second image depicting a second combination of items arranged together in a second environment;

determine, using an object recognition machine learning model, a first subset of the second image data corresponding to a first item of the second combination of items and a second subset of the second image data corresponding to a second item of the second combination of items;

generate, by the object recognition machine learning model based at least in part on detection of the first subset of the second image data, a first feature embedding for the first subset of the second image data visually representing the first item; and

generate by the object recognition machine learning model based at least in part on detection of the second subset of the second image data, a second feature embedding for the second subset of the second image data visually representing the second item.

13. The system of claim 12 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

determine the second image data by determining a distance between the first feature vector data and second feature vector data associated with the second image data in the machine learning feature space of the first feature vector data and the second feature vector data; and

determine that the distance is a within a tolerance of a minimum distance among image data stored in a repository of image data.

14. The system of claim 13 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

determine the second image data by determining a distance between the first feature vector data and second feature vector data associated with the second image data in the machine learning feature space of the first feature vector data and the second feature vector data; and

determine that the distance is a within a tolerance of a maximum distance among image data stored in a repository of image data.

15. The system of claim 12 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

receive a selection of a set of attributes for items associated with the second image data;

receive a control input effective to dismiss a third item of the second combination of items; and

display a different item in place of the third item, wherein the different item includes the set of attributes.

16. The system of claim 15 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to determine that second feature vector data representing the different item has a distance in the machine learning feature space that is within a tolerance of a maximum distance from a fourth feature vector data representing the third item among a set of items including the set of attributes.

17. The system of claim 12 , wherein the first image data, when rendered is effective to further display a first icon associated with a first item of the first combination of items, the at least one non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

receive data indicating a selection of the first icon; and

send third image data that, when rendered, is effective to display an enlarged image of the first item.

18. The system of claim 12 , wherein the first image data, when rendered is effective to further display:

a first icon associated with a first item of the first combination of items;

a second icon associated with a second item of the first combination of items; and

scrollable enlarged images of each item of the first combination of items, wherein the scrollable enlarged images are displayed on a different part of the display relative to the first combination of items arranged together in the first environment;

wherein the at least one non-transitory computer-readable memory storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

receive first data indicating a selection of the first item in the scrollable enlarged images; and

send second data to the first computing device effective to cause the first icon to change from a first appearance to a second appearance.

19. The system of claim 18 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

receive third data indicating a scroll operation from the first item in the scrollable enlarged images to the second item in the scrollable enlarged images;

send fourth data to the first computing device effective to cause the first icon to change from the second appearance to the first appearance; and

send fifth data to the first computing device effective to cause the second icon to change from the first appearance to the second appearance.

20. The system of claim 12 , storing further instructions that, when executed by the at least one processor, are further effective to program the at least one processor to:

determining that an enlarged representation of the first item is currently displayed by the first computing device; and

modifying an appearance of a graphical tag associated with the first item in the second image data in response to the enlarged representation of the first item being currently displayed.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: STERNBERG, JOSH; VISWANATHAN, AKSHAD; ZHANG, XIAOPENG; SMITH, ETHAN ALEXANDER; ROSE, LENWORTH RICHARD; FU, JOYCE HUANG; LV, MENGYUN; CHEN, RUI; INDREI, ALEXANDRU; DANO, REECE; SALAHI, SAEED
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058694/0672 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 19, 2022
From: LIANG, ANQI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 058695/0144 →
Continuity (1)
Continuation 16825273 · Mar 20, 2020
Cited By (6)
US 1,079,726 US 1,142,414 US 1,145,771 US 12,613,892 US 12,614,218 US 12,626,298