IP Library › Granted Patent US 12,634,651
Granted Patent B2
US 12,634,651 · App. 18/261,557 · Granted May 19, 2026

Processing of audio data

Inventors: Jussi Artturi Leppänen (Tampere, FI); Arto Juhani Lehtiniemi (Tampere, FI); Lasse Juhani Laaksonen (Tampere, FI); Miikka Tapani Vilermo (Tampere, FI)
Assignee: Nokia Technologies Oy
H04S7/303G06F3/017
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,634,651
App. No.
18/261,557
Granted
May 19, 2026
Kind
B2
Abstract

An apparatus, method and computer program is disclosed. The apparatus may comprise means for providing audio data for output to a user device, the audio data representing a virtual space comprising a plurality of sounds located at respective spatial locations within the virtual space, the plurality of sounds being respectively associated with a plurality of sound sources. The apparatus may also comprise means for detecting a predetermined gesture associated with a user identifying one of the plurality of sound sources to be a sound source of interest, determining a directional vector between a position of the user at a time of detecting the predetermined gesture and a position of the sound source of interest in the virtual space, and processing the audio data such that sounds at least in the direction of the directional vector are modified when output to the user device.

Claims (51)

1 . An apparatus comprising:

at least one processor; and

at least one memory storing instructions that, when executed with the at least one processor, cause the apparatus at least to:

provide audio data for output to a user device, the audio data representing a virtual space comprising a plurality of sounds located at respective spatial locations within the virtual space, the plurality of sounds being respectively associated with a plurality of sound sources;

detect a predetermined gesture associated with a user identifying one of the plurality of sound sources to be a sound source of interest;

determine a directional vector between a position of the user, at a time of detecting the predetermined gesture, and a position of the sound source of interest in the virtual space;

generate a directional beam pattern in the direction of the directional vector, wherein the generated directional beam pattern is configurated to have a beam width that is dependent, at least partially, on a duration for which the predetermined gesture is detected; and

process the audio data such that sounds at least in a direction of the directional vector are modified when output to the user device based, at least partially, on the duration for which the predetermined gesture is detected, wherein processing the audio data comprises causing the apparatus to:

modify at least part of the audio data corresponding to the generated directional beam pattern.

2 . The apparatus of claim 1 , wherein processing the audio data comprises the instructions, when executed with the at least one processor, cause the apparatus to:

modify the sounds as the position of the user changes.

3 . The apparatus of claim 1 , wherein the instructions, when executed with the at least one processor, cause the apparatus to:

determine a distance between the position of the user in the virtual space and the position of the sound source of interest, and wherein the beam width of the generated directional beam pattern is further dependent on the determined distance.

4 . The apparatus of claim 3 , wherein the beam width of the generated directional beam pattern becomes greater as the determined distance becomes smaller.

5 . The apparatus of claim 1 , wherein processing the audio data comprises the instructions, when executed with the at least one processor, cause the apparatus to:

modify the audio data such that a gain associated with the sounds in the direction of the directional vector is or are increased.

6 . The apparatus of claim 5 , wherein an amount of modification of the audio data is further dependent on the duration for which the predetermined gesture is detected.

7 . The apparatus of claim 1 , wherein detecting the predetermined gesture comprises the instructions, when executed with the at least one processor, cause the apparatus to:

detect the predetermined gesture responsive to detecting a change in orientation of at least part of the user's body.

8 . The apparatus of claim 7 , wherein the instructions, when executed with the at least one processor, cause the apparatus to:

determine an orientation of the user's head,

wherein detecting the predetermined gesture comprises the instructions, when executed with the at least one processor, cause the apparatus to:

detect the predetermined gesture responsive to detecting a change in the orientation of the user's head above a predetermined first angular threshold.

9 . The apparatus of claim 8 , wherein the instructions, when executed with the at least one processor, cause the apparatus to:

determine an orientation of the user's upper body with respect to the user's lower body,

wherein detecting the predetermined gesture comprises the instructions, when executed with the at least one processor, cause the apparatus to:

detect the predetermined gesture responsive to further detecting a change in the orientation of the upper body with respect to the lower body above a predetermined second angular threshold indicative of a downwards leaning motion.

10 . The apparatus of claim 7 , wherein detecting the predetermined gesture comprises the instructions, when executed with the at least one processor, cause the apparatus to:

detect the predetermined gesture responsive to detecting that the changed orientation of the at least part of the user's body is maintained for at least a predetermined period of time.

11 . The apparatus of claim 1 , wherein the directional vector is determined based on a position of a body part of the user after detection of the predetermined gesture.

12 . The apparatus of claim 11 , wherein the instructions, when executed with the at least one processor, cause the apparatus to:

determine a position of the user's ear and wherein the directional vector is determined as extending outwards from the user's ear.

13 . A method comprising:

providing audio data for output to a user device, the audio data representing a virtual space comprising a plurality of sounds located at respective spatial locations within the virtual space, the plurality of sounds being respectively associated with a plurality of sound sources;

detecting a predetermined gesture associated with a user identifying one of the plurality of sound sources to be a sound source of interest;

determining a directional vector between a position of the user, at a time of detecting the predetermined gesture, and a position of the sound source of interest in the virtual space;

generating a directional beam pattern in the direction of the directional vector, wherein the generated directional beam pattern is configurated to have a beam width that is dependent, at least partially, on a duration for which the predetermined gesture is detected; and

processing the audio data such that sounds at least in a direction of the directional vector are modified when output to the user device based, at least partially, on the duration for which the predetermined gesture is detected, wherein processing the audio data comprises:

modifying at least part of the audio data corresponding to the generated directional beam pattern.

14 . The method of claim 13 , wherein the processing further comprises:

modifying the sounds as the position of the user changes.

15 . The method of claim 13 , further comprising:

determining a distance between the position of the user in the virtual space and the position of the sound source of interest, and wherein beam width of the generated directional beam pattern is further dependent on the determined distance.

16 . The method of claim 15 , wherein the beam width of the generated directional beam pattern becomes greater as the determined distance becomes smaller.

17 . A non-transitory computer readable medium comprising instructions stored thereon for performing at least the following:

providing audio data for output to a user device, the audio data representing a virtual space comprising a plurality of sounds located at respective spatial locations within the virtual space, the plurality of sounds being respectively associated with a plurality of sound sources;

detecting a predetermined gesture associated with a user identifying one of the plurality of sound sources to be a sound source of interest;

determining a directional vector between a position of the user, at a time of detecting the predetermined gesture, and a position of the sound source of interest in the virtual space;

generate a directional beam pattern in the direction of the directional vector, wherein the generated directional beam pattern is configurated to have a beam width that is dependent, at least partially, on a duration for which the predetermined gesture is detected; and

processing the audio data such that sounds at least in a direction of the directional vector are modified when output to the user device based, at least partially, on the duration for which the predetermined gesture is detected, wherein processing the audio data comprises:

modifying at least part of the audio data corresponding to the generated directional beam pattern.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 29, 2023
From: ARTTURI LEPPÄNEN, JUSSI; JUHANI LEHTINIEMI, ARTO; JUHANI LAAKSONEN, LASSE; TAPANI VILERMO, MIIKKA
To: NOKIA TECHNOLOGIES OY
Reel/Frame 064731/0762 →
Priority Claims (1)
EP 21153938 · Jan 28, 2021 · regional
Continuity (1)
Related Publication 20240089688A1 · Mar 14, 2024
References Cited (15)
US 11290837B1 · Brimijoin, II · 2022 [cited by examiner]
US 20120257036A1 · Stenberg et al. · 2012 [cited by applicant]
US 20140129937A1 · Jarvinen et al. · 2014 [cited by applicant]
US 20200265860A1 · Mouncer · 2020 [cited by examiner]
JP H0990963A · 1997 [cited by applicant]
JP 7114531B2 · 2022 [cited by examiner]
KR 20180118034A · 2018 [cited by applicant]
Lee et al, “JP7114531B2 Ear Set Control Method and System”. English translation by EPO. 12 pages. (Year: 2019). [cited by examiner]
Office action received for corresponding European Patent Application No. 21153938.2, dated Apr. 5, 2024, 6 pages. [cited by applicant]
“Ambisonics”, Wikipedia, Retrieved on Jul. 20, 2023, Webpage available at : https://en.wikipedia.org/wiki/Ambisonics#Virtual_microphones. [cited by applicant]
Politis et al., “Parametric Spatial Audio Effects”, Proceedings of the 15th International Conference on Digital Audio Effects (DAFx12), Sep. 17-21, 2012, pp. 1-8. [cited by applicant]
Extended European Search Report received for corresponding European Patent Application No. 21153938.2, dated Jun. 9, 2021, 8 pages. [cited by applicant]
Schultz-Amling et al., “Acoustical Zooming Based on a Parametric Sound Field Representation”, Audio Engineering Society, May 2010, pp. 1310-1318. [cited by applicant]
Kallinger et al., “Spatial filtering using directional audio coding parameters”, IEEE International Conference on Acoustics, Speech and Signal Processing, Apr. 19-24, 2009, pp. 217-220. [cited by applicant]
International Search Report and Written Opinion received for corresponding Patent CooperationPCT/EP2022/051202, dated May 4, 2022, 12 pages. [cited by applicant]