IP Library Granted Patent US 12,450,837
Granted Patent B2
US 12,450,837 · App. 17/742,900 · Granted Oct 21, 2025

Contextual visual and voice search from electronic eyewear device

Inventors: David Meisenholder (Los Angeles, CA); Kameron Sheffield (South Jordan, UT); Joseph Timothy Fortier (Aspen, CO); Raymond Zeng (Long Island City, NY); Andrei Rybin (Lehi, UT); Jonathan Geddes (Saratoga Springs, UT)
Assignee: Snap Inc.
G06T19/006G06F3/04817G06F3/0482G06F3/167G06V10/40G06V10/74G06V10/768G06V20/20G06V20/60G10L15/08G10L15/22G06T2200/24G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,837
App. No.
17/742,900
Granted
Oct 21, 2025
Kind
B2
Abstract

Augmented reality features are selected for presentation to a display of an electronic eyewear device by using a camera of the electronic eyewear device to capture a scan image and processing the scan image to extract contextual signals. Simultaneously, voice data from the user is captured by a microphone of the electronic eyewear device and voice-to-text conversion of the captured voice data is performed to identify keywords in the voice data. The extracted contextual signals and the identified keywords are then used to select at least one augmented reality feature that matches the extracted contextual signals and the identified keywords, and the selected augmented reality feature is presented to the display for user selection. The contextual information thus refines the search results to provide the augmented reality feature best suited for the context of the scan image captured by the electronic eyewear device.

Claims (44)

1. An electronic eyewear device adapted to be worn on a head of a user, comprising:

a display;

at least one camera adapted to scan a scene in a viewing area around the user and

to capture a scan image;

a microphone adapted to capture voice data from the user;

a memory that stores instructions; and

a processor that executes the instructions to perform operations including:

initiating a scan by the at least one camera to capture the scan image;

processing the scan image or sending the scan image to an image processing device to extract at least one contextual signal from the scan image, wherein the at least one extracted contextual signal includes contextual data from the user and the viewing area around the user;

capturing, via the microphone, voice data from the user;

performing voice-to-text conversion of the captured voice data or sending the captured voice data to a voice data processing device to identify at least one keyword in the voice data;

processing the at least one extracted contextual signal and the at least one identified keyword or forwarding the at least one extracted contextual signal and the at least one identified keyword to an augmented reality feature storage to guide a search for an augmented reality feature having metadata that matches the user's search intent as determined from the at least one identified keyword and the at least one contextual signal and selecting at least one augmented reality feature from the augmented reality feature storage having metadata that matches the at least one extracted contextual signal and the at least one identified keyword;

presenting the selected at least one selected augmented reality feature to the display for user selection; and

applying an augmented reality feature selected by the user to the scene for display on the electronic eyewear device.

2. The electronic eyewear device of claim 1 , wherein the processor executes the instructions to perform additional operations including, for each contextual signal extracted from the scan image, presenting contextual signal descriptor text to the display.

3. The electronic eyewear device of claim 1 , wherein the at least one extracted contextual signal identifies at least one of a type of place or an object that is included in the scan image.

4. The electronic eyewear device of claim 1 , wherein the at least one extracted contextual signal identifies whether any tracking objects or markers are located in the scan image.

5. The electronic eyewear device of claim 1 , wherein initiating the scan is in response to a tap of a scan button or a press and hold of the scan button by the user.

6. The electronic eyewear device of claim 1 , wherein the processor executes the instructions to perform additional operations including presenting scan notifications to the display to indicate that a background scan has been initiated.

7. The electronic eyewear device of claim 1 , wherein the processor executes the instructions to present the at least one selected augmented reality feature to the display in a carousel of augmented reality features for user selection.

8. The electronic eyewear device of claim 7 , wherein the processor executes the instructions to badge the at least one selected augmented reality feature in the carousel with a scan icon that differentiates the at least one selected augmented reality feature from any other augmented reality feature in the carousel.

9. The electronic eyewear device of claim 1 , wherein the processor executes the instructions to perform additional operations including presenting the at least one identified keyword to the display.

10. The electronic eyewear device of claim 1 , wherein the processor executes the instructions to perform additional operations including determining whether the user has spoken after the initiation of a scan of the scene and, when the user has spoken after the initiation of a scan of the scene, capturing the voice data from the user and initiating a voice scan animation on the display indicating that the user will see scan results based on the user's voice data.

11. A method of selecting augmented reality features for presentation to a display of an electronic eyewear device, comprising:

initiating a scan of a scene in a viewing area around a user by at least one camera of the electronic eyewear device to capture a scan image;

processing the scan image or sending the scan image to an image processing device to extract at least one contextual signal from the scan image, wherein the at least one extracted contextual signal includes contextual data from a user and a viewing area around the user;

capturing voice data from the user;

performing voice-to-text conversion of the captured voice data or sending the captured voice data to a voice data processing device to identify at least one keyword in the voice data;

processing the at least one extracted contextual signal and the at least one identified keyword or forwarding the at least one extracted contextual signal and the at least one identified keyword to an augmented reality feature storage to guide a search for an augmented reality feature having metadata that matches the user's search intent as determined from the at least one identified keyword and the at least one contextual signal and selecting at least one augmented reality feature from the augmented reality feature storage having metadata that matches the at least one extracted contextual signal and the at least one identified keyword;

presenting the selected at least one selected augmented reality feature to the display for user selection; and

applying an augmented reality feature selected by the user to the scene for display on the electronic eyewear device.

12. The method of claim 11 , further comprising presenting at least one of contextual signal descriptor text for each contextual signal extracted from the scan image or the at least one identified keyword to the display of the electronic eyewear device.

13. The method of claim 11 , wherein presenting the selected at least one selected augmented reality feature to the display comprises presenting the at least one selected augmented reality feature in a carousel of augmented reality features for user selection.

14. The method of claim 13 , further comprising badging the at least one selected augmented reality feature in the carousel with a scan icon that differentiates the at least one selected augmented reality feature from any other augmented reality feature in the carousel.

15. The method of claim 11 , further comprising determining whether the user has spoken after the scan of the scene has been initiated and, when the user has spoken after the scan of the scene has been initiated, capturing the voice data from the user, and initiating a voice scan animation on the display indicating that the user will see scan results based on the user's voice data.

16. A non-transitory computer-readable storage medium that stores instructions that when executed by at least one processor cause the processor to select augmented reality features for presentation to a display of an electronic eyewear device by performing operations including:

initiating a scan of a scene in a viewing area around a user by at least one camera of the electronic eyewear device to capture a scan image;

processing the scan image or sending the scan image to an image processing device to extract at least one contextual signal from the scan image, wherein the at least one extracted contextual signal includes contextual data from a user and a viewing area around the user;

capturing voice data from the user;

performing voice-to-text conversion of the captured voice data or sending the captured voice data to a voice data processing device to identify at least one keyword in the voice data;

processing the at least one extracted contextual signal and the at least one identified keyword or forwarding the at least one extracted contextual signal and the at least one identified keyword to an augmented reality feature storage to guide a search for an augmented reality feature having metadata that matches the user's search intent as determined from the at least one identified keyword and the at least one contextual signal and selecting at least one augmented reality feature from the augmented reality feature storage having metadata that matches the at least one extracted contextual signal and the at least one identified keyword;

presenting the selected at least one selected augmented reality feature to the display for user selection; and

applying an augmented reality feature selected by the user to the scene for display on the electronic eyewear device.

17. The medium of claim 16 , further comprising instructions that when executed by the at least one processor causes the processor to present at least one of contextual signal descriptor text for each contextual signal extracted from the scan image or the at least one identified keyword to the display of the electronic eyewear device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2025
From: MEISENHOLDER, DAVID; SHEFFIELD, KAMERON; FORTIER, JOSEPH TIMOTHY; ZENG, RAYMOND; RYBIN, ANDREI; GEDDES, JONATHAN
To: SNAP INC.
Reel/Frame 072380/0046 →
Continuity (2)
Provisional Application 63190613 · May 19, 2021
Related Publication 20220375172A1 · Nov 24, 2022
References Cited (9)
US 10045001B2 · Anderson · 2018 [cited by examiner]
US 10394420B2 · Esinovskaya · 2019 [cited by examiner]
US 10712901B2 · Hwang et al. · 2020 [cited by applicant]
US 20180292907A1 · Katz · 2018 [cited by applicant]
US 20180307303A1 · Powderly · 2018 [cited by examiner]
US 20190289084A1 · Duan · 2019 [cited by examiner]
US 20200201514A1 · Murphy · 2020 [cited by examiner]
WO 2015073869A1 · 2015 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/028985, dated Aug. 26, 2022 (Aug. 26, 2022)—12 pages. [cited by applicant]