IP Library › Granted Patent US 10,838,206
Granted Patent B2
US 10,838,206 · App. 16/063,004 · Granted Nov 17, 2020

Head-mounted display for virtual and mixed reality with inside-out positional, user body and environment tracking

Inventors: Simon Fortin-Deschênes (Cupertino, CA); Vincent Chapdelaine-Couture (Cupertino, CA); Yan Côté (Cupertino, CA); Anthony Ghannoum (Cupertino, CA)
Assignee: Apple Inc.
G02B27/017G02B27/0093G06T19/006G02B2027/014G02B2027/0134G02B2027/0138G02B2027/0187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,838,206
App. No.
16/063,004
Filed
Sep 25, 2018
Granted
Nov 17, 2020
Kind
B2
Art Unit
2621
USPC
345/8
Abstract

A Head-Mounted Display system together with associated techniques for performing accurate and automatic inside-out positional, user body and environment tracking for virtual or mixed reality are disclosed. The system uses computer vision methods and data fusion from multiple sensors to achieve real-time tracking. High frame rate and low latency is achieved by performing part of the processing on the HMD itself.

Claims (41)

1. A method comprising:

at a head-mounted device including an image sensor, a display, a communications interface, and one or more processors:

capturing, via the image sensor, an image of a scene;

obtaining, via the communications interface, an image of content to be displayed in association with the image of the scene;

obtaining a depth map for at least a portion of the scene;

performing, using the one or more processors, one or more image signal processing functions on the image of the scene;

determining an alpha mask for the image of content by comparing depth values for each pixel in the image of content to depth vales for pixels in the processed image of the scene based on the depth map for at least the portion of the scene;

generating a display image by combining, using the one or more processors, the image of the content with the processed image of the scene based at least in part on the alpha mask; and

displaying, on the display, the display image.

2. The method of claim 1 , wherein displaying the display image is performed within 20 milliseconds of capturing the image of the scene.

3. The method of claim 1 , wherein the image of the scene is not transmitted via the communications interface.

4. The method of claim 1 , wherein a representation of the image of the scene is transmitted via the communications interface to enable remote computer vision processing.

5. The method of claim 1 , wherein neither the image of the scene, the processed image of the scene, nor the display image is received via the communications interface.

6. The method of claim 1 , wherein the head-mounted device includes one or more positional tracking sensors generating positional tracking data, further comprising, transmitting, via the communications interface, the positional tracking data, wherein the image of the content is based on the positional tracking data.

7. The method of claim 1 , wherein the one or more image signal processing functions include correction of image distortion for the display.

8. The method of claim 1 , wherein the one or more image signal processing functions include one or more of debayering, color correction, or noise reduction.

9. A head-mounted device comprising:

an image sensor to capture an image of a scene;

a depth sensor to obtain a depth map for at least a portion of the scene;

a communications interface to obtain an image of content to be displayed in association with the image of the scene; and

one or more processors to:

perform one or more image signal processing functions on the image of the scene; and

determine an alpha mask for the image of content by comparing depth values for each pixel in the image of content to depth vales for pixels in the processed image of the scene based on the depth map for at least the portion of the scene

generate a display image by combining the image of the content with the processed image of the scene based at least in part on the alpha mask; and

a display to display the display image.

10. The head-mounted device of claim 9 , wherein the display is to display the display image within 20 milliseconds of the image sensor capturing the image of the scene.

11. The head-mounted device of claim 9 , wherein the image of the scene is not transmitted via the communications interface.

12. The head-mounted device of claim 9 , wherein neither the image of the scene, the processed image of the scene, nor the display image is received via the communications interface.

13. The head-mounted device of claim 9 , wherein the one or more image signal processing functions include correction of image distortion for the display.

14. The head-mounted device of claim 9 , wherein the one or more image signal processing functions include one or more of debayering, color correction, or noise reduction.

15. A non-transitory computer-readable medium having instructions encoded thereon which, when executed by one or more processors of a head-mounted device including an image sensor, a communications interface, and a display caused the head-mounted device to:

capture, via the image sensor, an image of a scene;

obtain, via the communications interface, an image of content to be displayed in association with the image of the scene;

obtain a depth map for at least a portion of the scene;

perform, using the one or more processors, one or more image signal processing functions on the image of the scene;

determine an alpha mask for the image of content by comparing depth values for each pixel in the image of content to depth vales for pixels in the processed image of the scene based on the depth map for at least the portion of the scene;

generate a display image by combining, using the one or more processors, the image of the content with the processed image of the scene based at least in part on the alpha mask; and

display, on the display, the display image.

16. The non-transitory computer-readable medium of claim 15 , wherein displaying the display image is performed within 20 milliseconds of capturing the image of the scene.

17. The non-transitory computer-readable medium of claim 15 , wherein the image of the scene is not transmitted via the communications interface, and neither the image of the scene, the processed image of the scene, nor the display image is received via the communications interface.

18. The non-transitory computer-readable medium of claim 15 , wherein the image of the scene is not transmitted via the communications interface.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2018
From: FORTIN-DESCHENES, SIMON; CHAPDELAINE-COUTURE, VINCENT; COTE, YAN; GHANNOUM, ANTHONY
To: APPLE INC.
Reel/Frame 046591/0345 →
Continuity (2)
Provisional Application 62296829 · Feb 18, 2016
Related Publication 20190258058A1 · Aug 22, 2019
Cited By (27)
US 12,186,028 US 12,201,384 US 12,206,837 US 12,226,188 US 12,239,385 US 12,268,475 US 12,282,608 US 12,290,416 US 12,330,064 US 12,354,227 US 12,361,632 US 12,383,369 US 12,393,266 US 12,412,346 US 12,417,595 US 12,426,788 US 12,458,411 US 12,461,375 US 12,475,662 US 12,484,787 US 12,491,044 US 12,502,080 US 12,502,163 US 12,521,201 US 12,588,820 US 12,599,305 US 12,721,527