IP Library Granted Patent US 12,051,429
Granted Patent B2
US 12,051,429 · App. 18/138,684 · Granted Jul 30, 2024

Transform ambisonic coefficients using an adaptive network for preserving spatial direction

Inventors: Lae-Hoon Kim (San Diego, CA); Shankar Thagadur Shivappa (San Diego, CA); S M Akramus Salehin (San Diego, CA); Shuhua Zhang (San Diego, CA); Erik Visser (San Diego, CA)
Assignee: QUALCOMM Incorporated
G10L19/038G10L19/002H04R5/00G10L19/008H04R2430/21H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,051,429
App. No.
18/138,684
Granted
Jul 30, 2024
Kind
B2
Abstract

A device includes a memory configured to store untransformed ambisonic coefficients at different time segments. The device includes one or more processors configured to obtain the untransformed ambisonic coefficients at the different time segments, where the untransformed ambisonic coefficients at the different time segments represent a soundfield at the different time segments. The one or more processors are configured to apply one adaptive network, based on a constraint that includes preservation of a spatial direction of one or more audio sources in the soundfield at the different time segments, to the untransformed ambisonic coefficients at the different time segments to generate transformed ambisonic coefficients at the different time segments, wherein the transformed ambisonic coefficients at the different time segments represent a modified soundfield at the different time segments, that was modified based on the constraint. The one or more processors are also configured to apply an additional adaptive network.

Claims (36)

1. A device comprising:

a memory configured to store untransformed ambisonic coefficients at different time segments;

one or more processors configured to:

obtain the untransformed ambisonic coefficients at the different time segments, where the untransformed ambisonic coefficients at the different time segments represent a soundfield at the different time segments; and

apply one adaptive network, based on a constraint that includes preservation of a spatial direction of one or more audio sources in the soundfield at the different time segments, to the untransformed ambisonic coefficients at the different time segments to generate transformed ambisonic coefficients at the different time segments, wherein the transformed ambisonic coefficients at the different time segments represent a modified soundfield at the different time segments that was modified based on the constraint;

apply an additional adaptive network, and an additional constraint input into the additional adaptive network configured to output additional transformed ambisonic coefficients, based on the additional constraint, wherein the additional constraint includes preservation of a different spatial direction than the spatial direction preserved by the constraint; and

a renderer, configured to render the transformed ambisonic coefficients in a first spatial direction, and render the additional transformed ambisonic coefficients in a different spatial direction.

2. The device of claim 1 , further comprising an encoder configured to compress the transformed ambisonic coefficients, and further comprising a transmitter, configured to transmit the compressed transformed ambisonic coefficients over a transmit link.

3. The device of claim 1 , further comprising a receiver configured to receive compressed transformed ambisonic coefficients.

4. The device of claim 3 , further comprising a decoder configured to uncompress the compressed transformed ambisonic coefficients.

5. The device of claim 1 , further comprising a microphone array, configured to capture microphone signals that are converted to the untransformed ambisonic coefficients, and the constraint that includes the preservation of the spatial direction of one or more audio sources in the soundfield comes from a speaker zone in a vehicle.

6. The device of claim 1 , further comprising a combiner, wherein the combiner is configured to linearly add the additional transformed ambisonic coefficients and the transformed ambisonic coefficients.

7. The device of claim 1 wherein the transformed ambisonic coefficients in the first spatial direction are rendered to produce sound in a privacy zone.

8. The device of claim 7 , wherein the additional transformed ambisonic coefficients, in the different spatial direction, represent a masking signal, and are rendered to produce sound outside of the privacy zone.

9. The device of claim 7 , wherein the sound in the privacy zone is louder than sound produced outside of the privacy zone.

10. The device of claim 7 , wherein a privacy zone mode is activated in response to an incoming or an outgoing telephone call.

11. The device of claim 1 , wherein the constraint includes scaling the soundfield, at the different time segments by a scaling factor, wherein application of the scaling factor amplifies at least a first audio source in the soundfield represented by the untransformed ambisonic coefficients at the different time segments, wherein the transformed ambisonic coefficients, at the different time segments, represent a modified soundfield at the different time segments, that includes the at least first audio source that is amplified.

12. The device of claim 1 , wherein the constraint includes scaling the soundfield, at the different time segments by a scaling factor, wherein application of the scaling factor attenuates at least a first audio source in the soundfield represented by the untransformed ambisonic coefficients at the different time segments.

13. The device of claim 12 , wherein the transformed ambisonic coefficients at the different time segments, represent a modified soundfield at the different time segments, that includes the at least first audio source that is attenuated.

14. The device of claim 1 , the one or more processors convert microphone signals output captured at different microphone positions of a non-ideal microphone array into untransformed ambisonic coefficients based on performing a directivity adjustment.

15. The device of claim 14 , wherein the constraint includes correcting a biasing error introduced by the directivity adjustment, and the transformed ambisonic coefficients output by the adaptive network represent the audio source without the biasing error.

16. The device of claim 14 , wherein the untransformed ambisonic coefficients are transformed into transformed ambisonic coefficients based on the constraint of adjusting the microphone signals captured by a non-ideal microphone array as if the microphone signals had been captured by microphones at different positions of an ideal microphone array.

17. The device of claim 16 , wherein the ideal microphone array includes four microphones or thirty-two microphones.

18. The device of claim 1 , wherein the constraint includes target order of transformed ambisonic coefficients.

19. The device of claim 1 , wherein the constraint includes microphone positions for a form factor.

20. The device of claim 19 , wherein the form factor is a handset, glasses, VR headset, AR headset, another device integrated into a vehicle, or audio headset.

21. The device of claim 1 , wherein the transformed ambisonic coefficients are used by a first audio application that includes instructions that are executed by the one or more processors.

22. The device of claim 21 , wherein the first audio application includes compressing the transformed ambisonic coefficients at the different time segments and storing them in the memory.

23. The device of claim 22 , wherein compressed transformed ambisonic coefficients at the different time segments are transmitted over the air using a wireless link between the device and a remote device.

24. The device of claim 21 , wherein the first audio application further includes decompressing the compressed transformed ambisonic coefficients at the different time segments.

25. The device of claim 21 , wherein the first audio application includes renderer that is configured to render the transformed ambisonic coefficients at the different time segments.

26. The device of claim 21 , wherein the first audio application further includes a keyword detector, coupled to a device controller that is configured to control the device based on the constraint.

27. The device of claim 21 , wherein the first audio application further includes a direction detector, coupled to a device controller that is configured to control the device based on the constraint.

28. The device of claim 1 further comprising one or more loudspeakers configured to play the transformed ambisonic coefficients at the different time segments that were rendered by the renderer.

29. The device of claim 1 , wherein the device further comprises a microphone array configured to capture one or more audio sources that are represented by the untransformed ambisonic coefficients.

30. The device of claim 1 , wherein transformed ambisonic coefficients are stored in the memory, and the device further comprises a decoder configured to decode the transformed ambisonic coefficients based on the constraint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: KIM, LAE-HOON; THAGADUR SHIVAPPA, SHANKAR; SALEHIN, S M AKRAMUS; ZHANG, SHUHUA; VISSER, ERIK
To: QUALCOMM INCORPORATED
Reel/Frame 064198/0629 →
Continuity (4)
Continuation 17210357 · Mar 23, 2021
Provisional Application 62994147 · Mar 24, 2020
Provisional Application 62994158 · Mar 24, 2020
Related Publication 20230260525A1 · Aug 17, 2023