IP Library Granted Patent US 12,586,235
Granted Patent B2
US 12,586,235 · App. 18/021,164 · Granted Mar 24, 2026

Systems and methods for head related transfer function personalization

Inventors: Ramani Duraiswami (Highland, MD); Bowen Zhi (College Park, MD); Dmitry Zotkin (College Park, MD)
Assignee: CEVA Technologies, Inc.
G06T7/73H04S7/301G06T2207/30204H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,235
App. No.
18/021,164
Granted
Mar 24, 2026
Kind
B2
Abstract

A head-related transfer function (HRTF) generation system includes one or more processors configured to retrieve first image data of a first ear of a subject, compare the first image data with second image data of a plurality of second ears to identify a particular second ear of the plurality of second ears matching the first ear, identify a template HRTF associated with the particular second ear, and assign an HRTF to the subject based on the template HRTF.

Claims (54)

1 . A method for generating an individualized head related transfer function (HRTF), comprising:

retrieving, by one or more processors, first image data of a first ear of a subject;

rotating, by the one or more processors, the first image data so that an orientation of the first image data matches an orientation of second image data of a plurality of second ears;

comparing, by the one or more processors, the first image data with the second image data of the plurality of second ears to identify a particular second ear of the plurality of second ears matching the first ear;

identifying, by the one or more processors, a template HRTF associated with the particular second ear; and

outputting, by the one or more processors, the individualized HRTF for the subject based on the template HRTF.

2 . The method of claim 1 , further comprising:

identifying, by the one or more processors, whether the first ear is a left ear or a right ear;

if the plurality of second ears are left ears, upon identifying that the first ear is a right ear, flipping, by the one or more processors, the first image data; and

if the plurality of second ears are right ears, upon identifying that the first ear is a left ear, flipping, by the one or more processors, the first image data.

3 . The method of claim 1 , wherein the first image data comprises a plurality of markers of the first ear, each marker of the plurality of markers representing a location of an anatomical landmark of the first ear.

4 . The method of claim 3 , wherein the first image data is an image of the first ear, the method further comprising generating, by the one or more processors, the plurality of markers from the first image.

5 . The method of claim 1 , comprising:

maintaining the second image data to have a non-dimensional scale based on a particular distance of the second image data; and

modifying the first image data to have the non-dimensional scale.

6 . The method of claim 1 , wherein comparing the first image data with the second image data comprises:

modifying, by the one or more processors, the first image data to have at least one of a non-dimensional scale or a scale of the second image data; and

comparing, by the one or more processors, the modified first image data with the second image data.

7 . The method of claim 1 , wherein determining the particular at least one second ear matching the first ear comprises:

determining, by the one or more processors, a plurality of first distances between a plurality of first markers of the first ear; and

selecting, by the one or more processors, the particular second ear responsive to comparing the plurality of first distances with a plurality of second distances between a plurality of second markers of each second ear of the plurality of second ears.

8 . The method of claim 7 , wherein selecting the particular second ear comprises determining, using the plurality of first distances and the plurality of second distances, a weighted minimization of differences between the plurality of first distances and the plurality of second distances.

9 . The method of claim 1 , further comprising:

training, by the one or more processors, a machine learning model to generate markers of ears using training data comprising second markers assigned to each ear of the plurality of second ears; and

generating, by the one or more processors, a plurality of first markers of the first ear by applying the first image data as input to the machine learning model.

10 . The method of claim 1 , further comprising applying a head and torso (HAT) model to the template HRTF to generate the individualized HRTF.

11 . The method of claim 1 , further comprising:

identifying a first rotational orientation of the first image data and a second rotational orientation of the second image data associated with the particular second ear; and

applying a rotation to the template HRTF based on the first rotational orientation and the second rotational orientation.

12 . The method of claim 1 , further comprising generating, by an audio output device, audio output data by applying audio data and a source location of the audio data as input to the individualized HRTF.

13 . A system, comprising:

one or more processors configured to:

retrieve first image data of a first ear of a subject;

rotate the first image data so that an orientation of the first image data matches an orientation of second image data of a plurality of second ears;

compare the first image data with the second image data of the plurality of second ears to identify a particular second ear of the plurality of second ears matching the first ear;

identify a template HRTF associated with the particular second ear; and

output an individualized HRTF for the subject based on the template HRTF.

14 . The system of claim 13 , wherein the one or more processors are further configured to:

identify whether the first ear is a left ear or a right ear

if the plurality of second ears are left ears, upon identifying that the first ear is a right ear, flipping the first image data; and

if the plurality of second ears are right ears, upon identifying that the first ear is a left ear, flipping the first image data.

15 . The system of claim 13 , wherein the one or more processors are configured to compare the first image data with the second image data by:

modifying the first image data to have at least one of a non-dimensional scale or a scale of the second image data; and

comparing the modified first image data with the second image data.

16 . The system of claim 13 , wherein the one or more processors are configured to:

train a machine learning model to generate markers of ears using training data comprising second markers assigned to each ear of the plurality of second ears; and

generate a plurality of first markers of the first ear by applying the first image data as input to the machine learning model.

17 . A method for generating an individualized head related transfer function HRTF), comprising:

retrieving, by one or more processors, first image data of a first ear of a subject;

comparing, by the one or more processors, the first image data with second image data of a plurality of second ears to identify a particular second ear of the plurality of second ears matching the first ear;

identifying, by the one or more processors, a template HRTF associated with the particular second ear;

identifying a first rotational orientation of the first image data and a second rotational orientation of the second image data associated with the particular second ear; and

applying a rotation operation to the template HRTF based on the first rotational orientation and the second rotational orientation; and

outputting, by the one or more processors, the individualized HRTF for the subject based on the template HRTF.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2023
From: DURAISWAMI, RAMANI; ZHI, BOWEN; ZOTKIN, DMITRY
To: VISISONICS CORPORATION
Reel/Frame 065441/0598 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2023
From: VISISONICS CORPORATION
To: CEVA TECHNOLOGIES, INC.
Reel/Frame 065441/0627 →
Continuity (2)
Provisional Application 63065660 · Aug 14, 2020
Related Publication 20230222687A1 · Jul 13, 2023
References Cited (17)
US 9900722B2 · Bilinski et al. · 2018 [cited by applicant]
US 20130169779A1 · Pedersen · 2013 [cited by examiner]
US 20180204341A1 · Kaneko · 2018 [cited by applicant]
US 20180373957A1 · Lee · 2018 [cited by examiner]
US 20190082283A1 · Riggs · 2019 [cited by examiner]
US 20210211825A1 · Joyner · 2021 [cited by examiner]
US 20210385600A1 · Fukuda · 2021 [cited by examiner]
International Written Opinion and Search Report in International Patent Application No. PCT/US2021/045971, issued Nov. 17, 2021, 9 pages. [cited by applicant]
Satarzadeh et al., “Physical and filter pinna models based on anthropometry,” Audio Engineering Society Convention 122, AES 122nd Convention, Vienna, Austria, May 5-8, 2007 (20 pages). [cited by applicant]
Algazi et al., “Structural composition and decomposition of hrtfs.” In Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No.01TH8575), Oct. 2001 (pp. 103-106). [cited by applicant]
Algazi et al., “The cipic hrtf database.” In Proceedings of the 2001 IEEE Workshop on the Applications of Signal Processing to Audio and Acoustics (Cat. No.01TH8575), Oct. 2001 (pp. 99-102). [cited by applicant]
Algazi et al., “The use of head-and-torso models for improved spatial sound synthesis,” AES 113th Convention, Los Angeles, CA, USA, Oct. 5-8, 2002 (18 pages). [cited by applicant]
Kumar et al., “Kepler: Keypoint and pose estimation of unconstrained faces by learning efficient h-cnn regressors.” In 2017 12th IEEE International Conference on Automatic Face Gesture Recognition (FG 2017) May 2017 (pp… [cited by applicant]
Raykar et al., “Extracting the frequencies of the pinna spectral notches in measured head related impulse responses.” The Journal of the Acoustical Society of America, Jul. 2005, vol. 118, No. 1 (pp. 364-374). [cited by applicant]
Ronneberger et al., “U-net: Convolutional networks for biomedical image segmentation,” CoRR, abs/1505.04597, May 18, 2015 (8 pages). [cited by applicant]
Zotkin et al., “Fast head-related transfer function measurement via reciprocity,” The Journal of the Acoustical Society of America, Oct. 2006, vol. 120, No. 4 (pp. 2202-2215). [cited by applicant]
Zotkin et al., “Hrtf personalization using anthropometric measurements,” 2003 IEEE Workshop on Applications of Signal Processing to Audio and Acoustics (IEEE Cat. No. 03TH8684), Oct. 2003 (pp. 157-160). [cited by applicant]