IP Library Granted Patent US 12,556,802
Granted Patent B2
US 12,556,802 · App. 18/788,885 · Granted Feb 17, 2026

Image capturing using a display free body wearable computing device

Inventors: Seungmi Lee (Singapore, SG); Si Fi Faye Li (Singapore, SG); Michiel Sebastiaan Emanuel Petrus Knoppert (Amsterdam, NL); Yan Yan (Singapore, SG); Prabu Selvaraj (Singapore, SG); Chin Leong Ong (Singapore, SG); Weiyi Wang (Singapore, SG)
Assignee: Dell Products L.P.
H04N23/611H04N13/156H04N13/239H04N23/62
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,556,802
App. No.
18/788,885
Filed
Jul 30, 2024
Granted
Feb 17, 2026
Kind
B2
Art Unit
2485
USPC
348/47
Abstract

Methods and systems for capturing images of a scene using a display free body wearable computing device are disclosed. The method may include obtaining any number and/or types of user inputs from a user of the display free body wearable computing device. The user inputs may include, for example, voice commands, gestures, and/or any other user inputs. The method may also include interpreting a user input of the user inputs to identify a portion of the scene that the user wishes to capture in an image of the images. Once identified, using at least two image sensors, the display free body wearable computing device may capture and combine a stereo image depicting the desired portion of the scene to obtain a desired image.

Claims (63)

1 . A method for capturing images of a scene using a display free body wearable computing device, the method comprising:

obtaining, using at least one sensor of the display free body wearable computing device, at least one portion of user input from a user;

interpreting the at least one portion of user input to identify a portion of the scene that the user wishes to capture in an image of the images;

obtaining, using at least two image sensors of the display free body wearable computing device, a stereo image depicting at least the portion of the scene; and

combining the stereo image to obtain the image.

2 . The method of claim 1 , wherein obtaining the at least one portion of user input comprises:

obtaining, using a microphone array of the at least one sensor, a voice command from the user.

3 . The method of claim 1 , wherein obtaining the at least one portion of user input comprises:

obtaining, using at least one of the at least two image sensors, a guidance image depicting a portion of the scene and a portion of the user of the display free body wearable computing device.

4 . The method of claim 3 , wherein interpreting the at least one portion of user input comprises:

identifying, using the guidance image, an object depicted in the scene that is of interest to the user.

5 . The method of claim 4 , wherein obtaining the stereo image comprises:

directing, to obtain a desirable field of view for the at least two image sensors, movement of the user to:

remove the portion of the user from a field of view of the at least two image sensors, and

retain the object in the field of view; and

while the desirable field of view is present, activating the at least two image sensors to capture the stereo image.

6 . The method of claim 3 , wherein the portion of the user depicts a recognizable gesture.

7 . The method of claim 6 , wherein the recognizable gesture is a pointing gesture used by the user to convey interest in an object in the scene to the display free body wearable computing device.

8 . The method of claim 6 , wherein the recognizable gesture is a framing gesture used by the user to convey interest in a portion of the scene.

9 . The method of claim 8 , further comprising:

identifying a distance between the portion of the user and the at least two image sensors; and

performing, using the distance and the guidance image, parallax collection to identify the portion of the scene.

10 . The method of claim 1 , wherein the display free body wearable computing device comprises:

an integrated sensing and interaction component adapted to:

be positioned symmetrically on two portions of a user's head,

be positioned between ears and eyes of the user, and

capture a stereo image of at least a portion of a scene present in a field of view of the user;

an integrated computing, powering, and securing portion; and

an adjustment member adapted to position the integrated sensing and interaction component with respect to the integrated computing, powering, and securing portion.

11 . The method of claim 10 , wherein the integrated sensing and interaction component comprises:

a pair of cameras;

speakers;

a microphone array; and

a touch pad.

12 . The method of claim 11 , wherein the integrated sensing and interaction component is adapted to:

obtain the stereo image from the pair of cameras;

at least partially process the stereo image to obtain an image processing result;

identify an action to be performed based, at least in part, on the image processing result and a derived result from a remote entity, the derived result being based, at least in part, on the stereo image and/or the image processing result; and

use at least the speakers to perform the action.

13 . The method of claim 11 , wherein the pair of cameras comprise lenses configured to:

establish a camera line of sight that is parallel to a line of sight of the user; and

establish a camera field of view that comprises the field of view of the user.

14 . The method of claim 11 , wherein the stereo image comprises a pair of images of the scene, each of the images being captured at different angles and/or positions with respect to the scene by the pair of cameras.

15 . A non-transitory machine-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations for capturing images of a scene using a display free body wearable computing device, the operations comprising:

obtaining, using at least one sensor of the display free body wearable computing device, at least one portion of user input from a user;

interpreting the at least one portion of user input to identify a portion of the scene that the user wishes to capture in an image of the images;

obtaining, using at least two image sensors of the display free body wearable computing device, a stereo image depicting at least the portion of the scene; and

combining the stereo image to obtain the image.

16 . The non-transitory machine-readable medium of claim 15 , wherein obtaining the at least one portion of user input comprises:

obtaining, using a microphone array of the at least one sensor, a voice command from the user.

17 . The non-transitory machine-readable medium of claim 15 , wherein obtaining the at least one portion of user input comprises:

obtaining, using at least one of the at least two image sensors, a guidance image depicting a portion of the scene and a portion of the user of the display free body wearable computing device.

18 . A data processing system, comprising:

a processor;

and a memory coupled to the processor to store instructions, which when executed by the processor, cause the processor to perform operations for capturing images of a scene using a display free body wearable computing device, the operations comprising:

obtaining, using at least one sensor of the display free body wearable computing device, at least one portion of user input from a user;

interpreting the at least one portion of user input to identify a portion of the scene that the user wishes to capture in an image of the images;

obtaining, using at least two image sensors of the display free body wearable computing device, a stereo image depicting at least the portion of the scene; and

combining the stereo image to obtain the image.

19 . The data processing system of claim 18 , wherein obtaining the at least one portion of user input comprises:

obtaining, using a microphone array of the at least one sensor, a voice command from the user.

20 . The data processing system of claim 18 , wherein obtaining the at least one portion of user input comprises:

obtaining, using at least one of the at least two image sensors, a guidance image depicting a portion of the scene and a portion of the user of the display free body wearable computing device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 31, 2024
From: LEE, SEUNGMI; LI, SI FI FAYE; KNOPPERT, MICHIEL SEBASTIAAN EMANUEL PETRUS; YAN, YAN; SELVARAJ, PRABU; ONG, CHIN LEONG; WANG, WEIYI
To: DELL PRODUCTS L.P.
Reel/Frame 068143/0861 →
Continuity (1)
Related Publication 20260039942A1 · Feb 5, 2026
References Cited (64)
US 4907296A · Blecha · 1990 [cited by examiner]
US 5856811A · Shih · 1999 [cited by examiner]
US 7810750B2 · Abreu · 2010 [cited by applicant]
US 8159519B2 · Kurtz · 2012 [cited by examiner]
US 8902315B2 · Fisher et al. · 2014 [cited by applicant]
US 9538072B2 · Stewart et al. · 2017 [cited by applicant]
US 10110805B2 · Pomerantz · 2018 [cited by applicant]
US 10163210B2 · Kim · 2018 [cited by applicant]
US 10389993B2 · MacMillan et al. · 2019 [cited by applicant]
US 10924651B2 · Chaudhri et al. · 2021 [cited by applicant]
US 11196863B2 · Spohrer · 2021 [cited by applicant]
US 11206325B1 · Dennis · 2021 [cited by applicant]
US 11431660B1 · Leeds et al. · 2022 [cited by applicant]
US 11489996B2 · Burton · 2022 [cited by applicant]
US 11523055B1 · Chaudhri et al. · 2022 [cited by applicant]
US 11523243B2 · Satongar et al. · 2022 [cited by applicant]
US 11567569B2 · Spencer · 2023 [cited by applicant]
US 11816269B1 · Chaudhri et al. · 2023 [cited by applicant]
US 11899911B2 · Kocienda et al. · 2024 [cited by applicant]
US 20090122161A1 · Bolkhovitinov · 2009 [cited by applicant]
US 20110279666A1 · Strombom · 2011 [cited by examiner]
US 20140146153A1 · Birnkrant · 2014 [cited by examiner]
US 20150009550A1 · Misago · 2015 [cited by examiner]
US 20160225192A1 · Jones · 2016 [cited by examiner]
US 20170007351A1 · Yu · 2017 [cited by examiner]
US 20170099479A1 · Browd · 2017 [cited by examiner]
US 20170181802A1 · Sachs · 2017 [cited by examiner]
US 20170322410A1 · Watson · 2017 [cited by examiner]
US 20180012413A1 · Jones · 2018 [cited by examiner]
US 20180325498A1 · Bongiorno et al. · 2018 [cited by applicant]
US 20190253700A1 · Tornéus et al. · 2019 [cited by applicant]
US 20190254754A1 · Johnson · 2019 [cited by examiner]
US 20190370532A1 · Soni · 2019 [cited by applicant]
US 20200117025A1 · Sauer · 2020 [cited by examiner]
US 20200330179A1 · Ton · 2020 [cited by examiner]
US 20210067764A1 · Shau · 2021 [cited by examiner]
US 20210117680A1 · Chaudhri et al. · 2021 [cited by applicant]
US 20210169417A1 · Burton · 2021 [cited by applicant]
US 20230280821A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230280866A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230281254A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230281256A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230282214A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230283705A1 · Chaudhri et al. · 2023 [cited by applicant]
US 20230283885A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230283886A1 · Kocienda et al. · 2023 [cited by applicant]
US 20230327497A1 · Chaudhri et al. · 2023 [cited by applicant]
US 20240126363A1 · Kocienda et al. · 2024 [cited by applicant]
US 20240155194A1 · Kocienda et al. · 2024 [cited by applicant]
US 20240242721A1 · Kocienda et al. · 2024 [cited by applicant]
CA 3223178A1 · 2022 [cited by applicant]
WO 2020257329A1 · 2020 [cited by applicant]
WO 2023168001A1 · 2023 [cited by applicant]
WO 2023168071A1 · 2023 [cited by applicant]
WO 2023168073A1 · 2023 [cited by applicant]
WO 2024118974A1 · 2024 [cited by applicant]
Md Messal Monem Miah et al., “Multimodal Contextual Dialogue Breakdown Detection for Conversational AI Models”, NAACL 2024 Industry Track, arXiv:2404.08156v1, Apr. 11, 2024, <https://arxiv.org/abs/2404.08156v1>, retriev… [cited by applicant]
Donggang Jia et al., “Voice: Visual Oracle For Interaction, Conversation, and Explanation”, arXiv:2304.04083v2, Jan. 22, 2024, <https://arxiv.org/abs/2304.04083>, pp. 1-21, retrieved on Jul. 24, 2024 (21 pages). [cited by applicant]
Ambuj Mehrish et al., “A Review of Deep Learning Techniques for Speech Processing”, arXiv:2305.00359v3, May 30, 2023, <https://arxiv.org/pdf/2305.00359>, pp. 1-111, retrieved on Jul. 24, 2024 (111 pages). [cited by applicant]
Giuseppe Attanasio et al., “Twists, Humps, and Pebbles: Multilingual Speech Recognition Models Exhibit Gender Performance Gaps”, arXiv:2402.17954v2, Jun. 19, 2024, <https://arxiv.org/pdf/2402.17954>, retrieved on Jul. 2… [cited by applicant]
Konstantinos Tsiakas et al., “Unpacking Human-AI interactions: From interaction primitives to a design space”, arXiv:2401.05115v1, Jan. 10, 2024, < https://arxiv.org/abs/2401.05115>, pp. 1-46, retrieved on Jul. 24, 2024… [cited by applicant]
Pabbathi Sri Charan et al., “Effective Gesture Based Framework for Capturing User Input”, arXiv:2208.00913, Aug. 1, 2022, <https://arxiv.org/ftp/arxiv/papers/2208/2208.00913.pdf>, pp. 1-10, retrieved on Jul. 24, 2024 (1… [cited by applicant]
Chao Chen et al., “Simple calibration method for dual-camera structured light system”, Journal of the European Optical Society-Rapid Publications, 14, Article No. 23 (2018), Oct. 26, 2018, <https://doi.org/10.1186/s4147… [cited by applicant]
David Pierce, “Limitless is a new AI tool for your meetings—and an all-hearing wearable gadget”, The Verge, Apr. 15, 2024, <https://www.theverge.com/2024/4/15/24130832/limitless-ai-pendant-wearable-meetings> retrieved o… [cited by applicant]