IP Library › Granted Patent US 11,962,990
Granted Patent B2
US 11,962,990 · App. 17/498,707 · Granted Apr 16, 2024

Reordering of foreground audio objects in the ambisonics domain

Inventors: Dipanjan Sen (San Diego, CA); Sang-Uk Ryu (San Diego, CA)
Assignee: QUALCOMM Incorporated
H04S5/005G06F17/16G10L19/002G10L19/008G10L19/0204G10L19/038G10L19/06G10L19/167G10L19/20G10L25/18H04S7/30H04S7/304H04S7/40G10L2019/0001G10L2019/0005H04R2205/021H04S2400/01H04S2400/15H04S2420/01H04S2420/03H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,962,990
App. No.
17/498,707
Granted
Apr 16, 2024
Kind
B2
Abstract

In general, disclosed is a device that includes one or more processors, coupled to the memory, configured to perform an energy analysis with respect to one or more audio objects, in the ambisonics domain, in the first time segment. The one or more processors are also configured to perform a similarity measure between the one or more audio objects, in the ambisonics domain, in the first time segment, and the one or more audio objects, in the ambisonics domain, in the second time segment. In addition, the one or more processors are configured to perform a reorder of the one or more audio objects, in the ambisonics domain, in the first time segment with the one or more audio objects, in the ambisonics domain, in the second time segment, to generate one or more reordered audio objects in the first time segment.

Claims (36)

1. A device configured to encode foreground audio objects comprising:

a memory configured to store one or more foreground audio objects, in an ambisonics domain, in a first time segment and one or more foreground audio objects, in an ambisonics domain, in a second time segment; and

one or more processors, coupled to the memory, configured to:

perform an energy analysis with respect to the one, or more foreground audio objects, in the ambisonics domain, in the first time segment;

perform a similarity measure between the one or more foreground audio objects, in the ambisonics domain, in the first time segment, and the one or more foreground audio objects, in the ambisonics domain, in the second time segment;

perform a reorder of the one or more foreground audio objects, in the ambisonics domain, in the first time segment with the one or more foreground audio objects, in the ambisonics domain, in the second time segment, to generate, one or more reordered foreground audio objects, in the ambisonics domain, in the first time segment;

perform a reorder of one or more spatial vectors, in the ambisonics domain in the first time segment, corresponding to the one or more foreground audio objects, in the first time segment, based on a swap of spatial positions of at least two of the one or more foreground audio objects in a soundfield such that a first spatial position in the ambisonics domain is associated with a second of the at least two of the one or more foreground audio objects in the soundfield after the reorder of the one or more spatial vectors, and that a second spatial position in the ambisonics domain is associated with a first of the at least two of the one or more foreground audio objects in the soundfield after the reorder of the one or more spatial vectors; and

encode (a) the reordered foreground audio objects, in the ambisonics domain, in the first time segment, and (b) the reordered spatial vectors, in the ambisonics domain, in the first time segment.

2. The device of 1 , wherein the one or more processors are configured to:

determine directional property parameters of (i) the one or more foreground audio objects in the ambisonics domain in the first time segment, (ii) the one or more foreground audio objects in the ambisonics domain in the first time segment, or both (i) the one or more foreground audio objects, in the ambisonics domain, in the first time segment and (ii) the one or more foreground audio objects, in the ambisonics domain, in the second time segment.

3. The device of claim 2 , wherein the directional property parameters provide an indication of movement and location of the one more foreground audio objects, in the ambisonics domain, in the first time segment.

4. The device of claim 2 , wherein the directional property parameters provide an indication of movement and location of the one more foreground audio objects, in the ambisonics domain, in the second time segment.

5. The device of claim 1 , wherein the first time segment is an audio frame, and the second time segment is an audio frame.

6. The device of claim 1 , wherein the one or more processors are configured to perform the reorder of the one or more foreground audio objects, in the ambisonic domain, in the first time segment with the one or more foreground audio objects, in the ambisonic domain, in the second time segment comprises:

an energy comparison of one of the one or more foreground audio objects in the first time segment with more than one of the one or more foreground audio objects in the second time segment.

7. The device of claim 6 , wherein the one or more processors are configured to discard at least one of the one or more foreground audio objects in the second time segment, as reorder candidates, with the one of the one or more foreground audio objects in the first time segment.

8. The device of claim 1 , wherein the one or more processors are configured to sequentially perform (a) the reorder of the one or more spatial vectors, in the ambisonics domain, in the first time segment, corresponding to the one or more foreground audio objects in the first time segment with (b) the reorder of the one or more foreground audio objects, in the ambisonic domain, in the first time segment.

9. The device of claim 1 , wherein the one or more processors are configured to: concurrently perform (a) the reorder of the one or more spatial vectors, in the ambisonics domain, corresponding to the one or more foreground audio objects in the first time segment with (b) the reorder of the one or more foreground audio objects, in the ambisonics domain, in the first time segment.

10. The device of claim 1 , wherein the one or more processors are configured to generated separate syntax elements that to indicate the reorder of the one or more spatial vectors in the first time segment and the reorder of the one or more foreground audio objects in the first time segment.

11. The device of claim 1 , wherein the one or more processors are configured to differently reorder (a) the one or more spatial vectors in the first time segment than (b) the one or more foreground audio objects in the first time segment.

12. The device of claim 11 , wherein the one or more processors are configured to differently reorder based on the swap of spatial positions of the at least two of the one or more foreground audio objects in the soundfield.

13. Device of claim 1 , wherein the similarity measure is based on a correlation operation.

14. A device configured to decode a bitstream comprising:

a memory configured to store the bitstream;

one or more processors, coupled to the memory, configured to:

receive the bitstream that includes reorder information to determine how one or more reordered foreground audio objects, in an ambisonics domain, in a first time segment were reordered;

decompress the received bitstream to:

generate the foreground audio objects, in the ambisonics domain, in the first time segment based on the reorder information; and

generate one or more spatial vectors, in the ambisonics domain, wherein the one or more spatial vectors corresponds to the one or more foreground audio objects in the first time segment, wherein the reorder information is based on a swap of spatial positions of at least two of the one or more foreground audio objects in a soundfield such that a first spatial position in the ambisonics domain is associated with a second of the at least two of the one or more foreground audio objects in the soundfield after the reorder of the one or more spatial vectors, and that a second spatial position in the ambisonics domain is associated with a first of the at least two of the one or more foreground audio objects in the soundfield after the reorder of the one or more spatial vectors.

15. The device of claim 14 , wherein the one or more processors are configured to generate un-reordered one or more foreground audio objects, in the ambisonics domain, in the first time segment.

16. The device of claim 15 , wherein the one or more processors are configured to sequentially perform (a) the un-reorder of the one or more spatial vectors, in the ambisonics domain, corresponding to the one or more foreground audio objects in the first time segment with (b) the un-reorder of the one or more foreground audio objects, in the ambisonics domain, in the first time segment.

17. The device of claim 15 , wherein the one or more processors are configured to concurrently perform (a) the un-reorder of the one or more spatial vectors, in the ambisonics domain, corresponding to the one or more foreground audio objects in the first time segment with (b) the un-reorder of the one or more foreground audio objects, in the ambisonics domain, in the first time segment.

18. The device of claim 15 , wherein the one or more processors are configured to differently un-reorder (a) the one or more spatial vectors in the first time segment than (b) the one or more foreground audio objects in the first time segment.

19. The device of claim 18 , wherein the one or more processors are configured to differently un-reorder based on the swap of spatial positions of the at least two of the one or more foreground audio objects in the soundfield.

20. The device of claim 14 , wherein the first time segment is an audio frame and the second time frame is an audio frame.

21. The device of claim 15 , wherein the one or more processors are configured to receive syntax elements separately, a first syntax element for the one or more foreground audio objects in the first time segment and a second syntax element for the one or more spatial vectors, in the ambisonics domain, in the first time segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 1, 2022
From: SEN, DIPANJAN; RYU, SANG-UK
To: QUALCOMM INCORPORATED
Reel/Frame 059135/0783 →
Continuity (20)
Continuation 14289522 · May 28, 2014
Provisional Application 61828445 · May 29, 2013
Provisional Application 61828615 · May 29, 2013
Provisional Application 61829791 · May 31, 2013
Provisional Application 61899034 · Nov 1, 2013
Provisional Application 61899041 · Nov 1, 2013
Provisional Application 61829182 · May 30, 2013
Provisional Application 61829174 · May 30, 2013
Provisional Application 61829155 · May 30, 2013
Provisional Application 61933706 · Jan 30, 2014
Provisional Application 61829846 · May 31, 2013
Provisional Application 61886605 · Oct 3, 2013
Provisional Application 61886617 · Oct 3, 2013
Provisional Application 61925158 · Jan 8, 2014
Provisional Application 61933721 · Jan 30, 2014
Provisional Application 61925074 · Jan 8, 2014
Provisional Application 61925112 · Jan 8, 2014
Provisional Application 61925126 · Jan 8, 2014
Provisional Application 62003515 · May 27, 2014
Related Publication 20220030372A1 · Jan 27, 2022
Cited By (1)
US 12,586,591