IP Library Granted Patent US 10,659,908
Granted Patent B2
US 10,659,908 · App. 16/542,930 · Granted May 19, 2020

System and method to capture image of pinna and characterize human auditory anatomy using image of pinna

Inventor: Kapil Jain (Redwood City, CA)
Assignee: EmbodyVR, Inc.
H04S7/304A61B5/121G01B7/14G06K9/00362G06K9/66G06T1/0007G06T7/73G06T7/80H04R1/005H04R1/1008H04R1/1058H04R1/22H04R1/32H04R1/323H04R3/04H04R5/04H04S3/008H04S7/301H04S7/306A61B5/4005A61B6/5217H04R1/1016H04R5/033H04R11/02H04R2201/029H04R2225/77H04S2400/11H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,659,908
App. No.
16/542,930
Granted
May 19, 2020
Kind
B2
Abstract

An image of a pinna is captured. Based on the image of the pinna, a non-linear transfer function is determined which characterizes how sound is transformed at the pinna. A signal is output indicative of one or more audio cues to facilitate spatial localization of sound via the pinna, where the one or more audio cues is based on the non-linear transfer function.

Claims (34)

1. A method comprising:

capturing, by an image sensor, one or more images of a reference pinna;

based on the one or more images of the reference pinna, determining a head related transfer function (HRTF) which characterizes how sound is transformed by the reference pinna, wherein determining the HRTF comprises inputting the one or more images of the reference pinna captured by the image sensor into a neural network model which outputs the HRTF which characterizes how sound is transformed by the reference pinna; and

outputting, by a transducer, a signal indicative of one or more audio cues to facilitate spatial localization of sound via the reference pinna, wherein the one or more audio cues is based on the HRTF.

2. The method of claim 1 , wherein capturing the one or more images of the reference pinna comprises determining, by a proximity sensor, a distance between a mobile device and the reference pinna; and capturing, by the image sensor of the mobile device, the one or more images of the reference pinna when a proximity sensor indicates that the mobile device is at a given linear distance from the reference pinna.

3. The method of claim 1 , wherein capturing the one or more images of the reference pinna comprises positioning a mobile device having a camera configured with the image sensor proximate to the reference pinna to capture a video of the reference pinna.

4. The method of claim 1 , wherein determining the HRTF comprises extracting characteristics of one or more features of the reference pinna based on the one or more images of the reference pinna.

5. The method of claim 1 , wherein capturing the one or more images of the reference pinna comprises capturing, via a mobile device configured with the image sensor, a plurality of images of the reference pinna at a plurality of orientations, wherein the plurality of images of the reference pinna at the plurality of orientations indicates at least a relative distance between features of the reference pinna and a depth of features of the reference pinna.

6. The method of claim 5 , wherein capturing the one or more images of the reference pinna comprises combining the plurality of images of the reference pinna to form a multi-dimensional image of the reference pinna; and wherein determining the HRTF comprises inputting the multi-dimensional image into the neural network model.

7. The method of claim 1 , wherein the one or more images of the reference pinna is associated with a first individual and the neural network model is fit to a plurality of HRTFs and associated images of pinna associated with second individuals other than the first individual.

8. One or more non-transitory machine-readable media comprising computer instructions stored in memory and executable by a processor, the computer instructions to:

capture, by an image sensor, one or more images of a reference pinna;

based on the one or more images of the reference pinna, determine a head related transfer function (HRTF) which characterizes how sound is transformed by the reference pinna, wherein determining the HRTF comprises inputting the one or more images of the reference pinna captured by the image sensor into a neural network model which outputs the HRTF which characterizes how sound is transformed by the reference pinna; and

output, by a transducer, a signal indicative of one or more audio cues to facilitate spatial localization of sound via the reference pinna, wherein the one or more audio cues is based on the HRTF.

9. The one or more non-transitory machine-readable media of claim 8 , wherein the computer instructions to capture the one or more images of the reference pinna comprises computer instructions to determine, by a proximity sensor, a distance between a mobile device and the reference pinna; and capture, by the image sensor of the mobile device, the one or more images of the reference pinna when a proximity sensor indicates that the mobile device is at a given linear distance from the reference pinna.

10. The one or more non-transitory machine-readable media of claim 8 , wherein the computer instructions to capture the one or more images of the reference pinna comprises receiving the one or more images of the reference pinna from a mobile device having a camera configured with the image sensor, the image sensor positioned proximate to the reference pinna to capture a video of the reference pinna.

11. The one or more non-transitory machine-readable media of claim 8 , wherein the computer instructions to determine the HRTF comprises computer instructions to extract characteristics of one or more features of the reference pinna based on the one or more images of the reference pinna.

12. The one or more non-transitory machine-readable media of claim 8 , wherein the computer instructions to capture the one or more images of the reference pinna comprises computer instructions to capture, via a mobile device configured with the image sensor, a plurality of images of the reference pinna at a plurality of orientations, wherein the plurality of images of the reference pinna at the plurality of orientations indicates at least a relative distance between features of the reference pinna and a depth of features of the reference pinna.

13. The one or more non-transitory machine-readable media of claim 12 , wherein the computer instructions to capture the one or more images of the reference pinna comprises computer instructions to combine the plurality of images of the reference pinna to form a multi-dimensional image of the reference pinna; and wherein the computer instructions to determine the HRTF comprises computer instructions to input the multi-dimensional image into the neural network model.

14. The one or more non-transitory machine-readable media of claim 8 , wherein the one or more images of the reference pinna is associated with a first individual and the neural network model is fit to a plurality of HRTFs and associated images of pinna associated with second individuals other than the first individual.

15. A system comprising:

a mobile device having a camera, the camera configured with an image sensor;

computer instructions stored in memory and executable by a processor to perform functions of:

capturing, by the image sensor of the camera, one or more images of a reference pinna;

based on the one or more images of the reference pinna, determining a head related transfer function (HRTF) which characterizes how sound is transformed by the reference pinna, wherein determining the HRTF comprises inputting the one or more images of the reference pinna captured by the image sensor into a neural network model which outputs the HRTF which characterizes how sound is transformed by the reference pinna; and

outputting a signal indicative of one or more audio cues to facilitate spatial localization of sound via the reference pinna, wherein the one or more audio cues is based on the HRTF.

16. The system of claim 15 , wherein the computer instructions to capture the one or more images of the reference pinna comprises receiving the one or more images of the reference pinna from the mobile device having the camera, the image sensor of the camera positioned proximate to the reference pinna to capture a video of the reference pinna.

17. The system of claim 15 , wherein the computer instructions to determine the HRTF comprises computer instructions to extract characteristics of one or more features of the reference pinna based on the one or more images of the reference pinna.

18. The system of claim 15 , wherein the computer instructions to capture the one or more images of the reference pinna comprises computer instructions to capture a plurality of images of the reference pinna at a plurality of orientations, wherein the plurality of images of the reference pinna at the plurality of orientations indicates at least a relative distance between features of the reference pinna and a depth of features of the reference pinna.

19. The system of claim 18 , wherein the computer instructions to capture the one or more images of the reference pinna comprises computer instructions to combine the plurality of images of the reference pinna to form a multi-dimensional image of the reference pinna; and wherein the computer instructions to determine the HRTF comprises computer instructions to input the multi-dimensional image into the neural network model.

20. The system of claim 15 , wherein the one or more images of the reference pinna is associated with a first individual and the neural network model is fit to a plurality of HRTFs and associated images of pinna associated with second individuals other than the first individual.

21. The method of claim 1 , wherein inputting the one or more images of the reference pinna captured by the image sensor into the neural network model which outputs the HRTF comprises the neural network model predicting the HRTF based on the one or more images of the reference pinna captured by the image sensor.

22. The one or more non-transitory machine-readable media of claim 8 , wherein the computer instructions to input the one or more images of the reference pinna captured by the image sensor into the neural network model which outputs the HRTF comprises computer instructions for the neural network model to predict the HRTF based on the one or more images of the reference pinna captured by the image sensor.

23. The system of claim 15 , wherein the computer instructions to input the one or more images of the reference pinna captured by the image sensor into the neural network model which outputs the HRTF comprises computer instructions for the neural network model to predict the HRTF based on the one or more images of the reference pinna captured by the image sensor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2019
From: JAIN, KAPIL
To: EMBODYVR, INC
Reel/Frame 050076/0311 →
Continuity (7)
Continuation 15811441 · Nov 13, 2017
Provisional Application 62421285 · Nov 13, 2016
Provisional Application 62421380 · Nov 14, 2016
Provisional Application 62424512 · Nov 20, 2016
Provisional Application 62466268 · Mar 2, 2017
Provisional Application 62468933 · Mar 8, 2017
Related Publication 20190379993A1 · Dec 12, 2019
Cited By (1)
US 12,225,371