IP Library Granted Patent US 12,273,704
Granted Patent B2
US 12,273,704 · App. 18/171,096 · Granted Apr 8, 2025

Information processing device, information processing method, and information processing program

Inventors: Kazumi Fukuda (Tokyo, JP); Tetsu Magariyachi (Kanagawa, JP)
Assignee: Sony Group Corporation
H04S7/304G06V10/454G06V10/764G06V10/7715G06V10/82G06V20/647G06V40/10H04R5/033H04S3/004H04S7/301H04S2420/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,273,704
App. No.
18/171,096
Granted
Apr 8, 2025
Kind
B2
Abstract

An information processing device ( 100 ) according to the present disclosure includes: an acquisition unit ( 141 ) configured to acquire a first image including a content image of an ear of a user; and a calculation unit ( 142 ) configured to calculate, based on the first image acquired by the acquisition unit ( 141 ), a head-related transfer function corresponding to the user by using a learned model having learned to output a head-related transfer function corresponding to an ear when an image including a content image of the ear is input.

Claims (26)

1. An information processing device comprising:

circuitry configured to:

acquire a first image including a content image of an ear of a user;

acquire a first ear parameter from the first image by using a learned model having learned a relation between a second image including a content image of the ear and a second ear parameter; and

acquire a head-related transfer function based on input of the first ear parameter acquired by the first image by using a learned model having learned a relation between the second ear parameter and the head-related transfer function, wherein the circuitry is configured to acquire the first ear parameter of the ear included in the first image using an ear parameter estimation model having learned to output an ear parameter corresponding to an ear when an image including a content image of the ear is input.

2. The information processing device according to claim 1 ,

wherein the first ear parameter comprises a variable representing a characteristic of the ear included in the first image.

3. The information processing device according to claim 1 , wherein the ear parameter estimation model is generated by learning a relation between an image including a content image of the ear and an ear parameter of the ear.

4. The information processing device according to claim 3 , wherein the ear parameter estimation model is generated by learning a relation between the ear parameter and an ear image obtained by rendering three-dimensional data of the ear generated based on the ear parameter.

5. The information processing device according to claim 4 , wherein the ear parameter estimation model is generated by learning the relation between a plurality of ear images obtained by changing a camera angle in rendering.

6. The information processing device according to claim 5 , wherein the circuitry is configured to perform acoustic simulation for three-dimensional data obtained by synthesizing three-dimensional data of the ear generated based on the ear parameter and three-dimensional data of a head, and to generate a learned model by learning a relation between a head-related transfer function obtained through the acoustic simulation and the ear parameter.

7. An information processing method by which a computer performs:

acquiring a first image including a content image of an ear of a user;

acquiring a first ear parameter from the first image by using a learned model having learned a relation between a second image including a content image of the ear and a second ear parameter; and

acquiring a head-related transfer function based on input of the first ear parameter acquired by the first image by using a learned model having learned a relation between the second ear parameter and the head-related transfer function, wherein acquiring the first ear parameter of the ear included in the first image includes using an ear parameter estimation model having learned to output an ear parameter corresponding to an ear when an image including a content image of the ear is input.

8. The information processing method according to claim 7 , wherein the first ear parameter comprises a variable representing a characteristic of the ear included in the first image.

9. The information processing method according to claim 7 , wherein the ear parameter estimation model is generated by learning a relation between an image including a content image of the ear and an ear parameter of the ear.

10. The information processing method according to claim 9 , wherein the ear parameter estimation model is generated by learning a relation between the ear parameter and an ear image obtained by rendering three-dimensional data of the ear generated based on the ear parameter.

11. The information processing method according to claim 10 , wherein the ear parameter estimation model is generated by learning the relation between a plurality of ear images obtained by changing a camera angle in rendering.

12. The information processing method according to claim 11 , further comprising performing acoustic simulation for three-dimensional data obtained by synthesizing three-dimensional data of the ear generated based on the ear parameter and three-dimensional data of a head, and generating a learned model by learning a relation between a head-related transfer function obtained through the acoustic simulation and the ear parameter.

13. A non-transitory computer-readable storage medium encoded with executable instructions that, when executed by at least one processor, cause the at least one processor to perform:

acquiring a first image including a content image of an ear of a user;

acquiring a first ear parameter from the first image by using a learned model having learned a relation between a second image including a content image of the ear and a second ear parameter; and

acquiring a head-related transfer function based on input of the first ear parameter acquired by the first image by using a learned model having learned a relation between the second ear parameter and the head-related transfer function, wherein acquiring the first ear parameter of the ear included in the first image includes using an ear parameter estimation model having learned to output an ear parameter corresponding to an ear when an image including a content image of the ear is input.

14. The non-transitory computer-readable storage medium according to claim 13 , wherein the first ear parameter comprises a variable representing a characteristic of the ear included in the first image.

15. The non-transitory computer-readable storage medium according to claim 13 , wherein the ear parameter estimation model is generated by learning a relation between an image including a content image of the ear and an ear parameter of the ear.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 15, 2023
From: FUKUDA, KAZUMI; MAGARIYACHI, TETSU
To: SONY CORPORATION
Reel/Frame 063642/0187 →
CHANGE OF NAME Recorded May 15, 2023
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 063642/0919 →
Priority Claims (1)
JP 2018-191513 · Oct 10, 2018 · national
Continuity (2)
Continuation 17282705
Related Publication 20230283979A1 · Sep 7, 2023
References Cited (33)
US 6996244B1 · Slaney et al. · 2006 [cited by applicant]
US 9544706B1 · Hirst · 2017 [cited by applicant]
US 10038966B1 · Mehra · 2018 [cited by applicant]
US 10341803B1 · Mehra · 2019 [cited by applicant]
US 11595772B2 · Fukuda · 2023 [cited by examiner]
US 20060067548A1 · Slaney · 2006 [cited by examiner]
US 20070019812A1 · Kim · 2007 [cited by applicant]
US 20130169779A1 · Pedersen · 2013 [cited by applicant]
US 20180132764A1 · Jain · 2018 [cited by applicant]
US 20180373957A1 · Lee et al. · 2018 [cited by applicant]
US 20190014431A1 · Lee et al. · 2019 [cited by applicant]
US 20210385600A1 · Fukuda et al. · 2021 [cited by applicant]
US 20220027033A1 · Okimoto et al. · 2022 [cited by applicant]
CN 1901761A · 2007 [cited by applicant]
CN 103139677A · 2013 [cited by applicant]
EP 2611216A1 · 2013 [cited by applicant]
EP 3351172A1 · 2018 [cited by applicant]
JP H1083190A · 1998 [cited by applicant]
JP 2004314915A · 2004 [cited by applicant]
JP 2013168924A · 2013 [cited by applicant]
JP 2017216660A · 2017 [cited by applicant]
WO WO2017047309A1 · 2017 [cited by applicant]
WO WO2017116308A1 · 2017 [cited by applicant]
Kaneko, Shoken (The Acoustical Society of Japan, Meeting Report (Spring, 2017), “Estimating ear solid shape from ear photographs by statistical ear shape modeling and deep learning and HRTF personalization”), pp. 449-45… [cited by examiner]
International Search Report and English translation thereof mailed Nov. 26, 2019 in connection with International Application No. PCT/JP2019/039103. [cited by applicant]
International Written Opinion and English translation thereof mailed Nov. 26, 2019 in connection with International Application No. PCT/JP2019/039103. [cited by applicant]
International Preliminary Report on Patentability and English translation thereof mailed Apr. 22, 2021 in connection with International Application No. PCT/JP2019/039103. [cited by applicant]
Kaneko et al., Ear Three-Dimensional Shape Estimation and HRTF Individuation from Ear Photographs using Statistical Ear Shape Modeling and Deep Learning. Reports of the 2017 Spring Meeting the Acoustical Society of Japa… [cited by applicant]
CHUN et al., Deep Neural Network Based HRTF Personalization Using Anthropometric Measurements. AES 143RD Convention, Oct. 18-21, 2017. New York, NY, USA. 5 pages. [cited by applicant]
Kaneko et al., DeepEarNet: Individualizing Spatial Audio with Photography, Ear Shape Modeling, and Neural Networks. Conference Paper: 2016 AES International Conference on Audio for Virtual and Augmented Reality. Sep. 20… [cited by applicant]
Torres-Gallegos et al., Personalization of head-related transfer functions (HRTF) based on automatic photo-anthropometry and inference from a database. Applied Acoustics, Elsevier Publishing, GB. Oct. 2015, v97; pp. 84-… [cited by applicant]
Yao et al., Head-Related Transfer Function Selection Using Neural Networks. Archives of Acoustics. Apr. 2017, v 42(3); pp. 365-373. [cited by applicant]
“Statistical ear shape estimation and the estimation of the ear stereoscopic shape by using a deep layer learning”, Japan Audio Society, Two Thousand Seventeenth Division, Research Institute, vol. 2017, No. 03, pp. 449 … [cited by applicant]