IP Library › Granted Patent US 12,676,156
Granted Patent B2
US 12,676,156 · App. 18/493,363 · Granted Jul 7, 2026

Encoding device and encoding method, decoding device and decoding method, and program

Inventors: Toru Chinen (Kanagawa, JP); Masayuki Nishiguchi (Kanagawa, JP); Runyu Shi (Kanagawa, JP); Mitsuyuki Hatanaka (Kanagawa, JP); Yuki Yamamoto (Tokyo, JP)
Assignee: Sony Group Corporation
G10L19/008G10L19/032G10L19/06G10L19/22G10L19/24G10L21/038H03M7/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,676,156
App. No.
18/493,363
Filed
Oct 24, 2023
Granted
Jul 7, 2026
Kind
B2
Art Unit
2654
USPC
704/500
Abstract

There is provided a decoding device including at least one circuit configured to acquire one or more encoded audio signals including a plurality of channels and/or a plurality of objects and priority information for each of the plurality of channels and/or the plurality of objects, and to decode the one or more encoded audio signals according to the priority information.

Claims (70)

1 . A decoding device comprising:

circuitry configured to:

acquire one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;

acquire meta-data for each of the plurality of audio objects;

decode the one or more encoded audio signals according to the meta-data; and

render the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:

the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,

a value range of the priority information is from 0 to 7,

the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:

a high-frequency component of a non-encoded audio signal is generated from an audio signal of:

a low frequency component generated from performing overlapping addition, and

a high frequency power value,

the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,

a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,

an audio signal including the high frequency component is generated, and

the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.

2 . The decoding device according to claim 1 , wherein

a lowest priority degree is 0 and a highest priority degree is 7.

3 . The decoding device according to claim 1 , wherein

the meta-data is position information indicating each of position of the plurality of audio objects.

4 . The decoding device according to claim 3 , wherein

the position information is expressed by a horizontal angle from a predetermined reference position, a vertical angle from the predetermined reference position, and a distance from the predetermined reference position to a predetermined audio object.

5 . The decoding device according to claim 1 , wherein

the meta-data is a gain of the audio object.

6 . The decoding device according to claim 1 , wherein

a parameter called num_objects indicates a number of objects representing the plurality of audio objects.

7 . The decoding device according to claim 1 , wherein

the circuitry is further configured to perform IMDCT (inverse modified discrete cosine transform) to generate an audio signal.

8 . The decoding device according to claim 1 , wherein

a sum of audio signals from each of the plurality of channels is multiplied by a VBAP gain of corresponding channels in the plurality of channels, producing an audio signal for each channel.

9 . The decoding device according to claim 8 , wherein

gain adjustment of the encoded audio signals of each object is performed through the VBAP gain.

10 . The decoding device according to claim 1 , wherein

a speaker index, s, acts to specify a speaker corresponding to a predetermined channel.

11 . The decoding device according to claim 1 , wherein

the processing comprises SBR (spectral band) processing is performed with respect to and the audio signal comprises an audio signal obtained by overlappingly adding IMDCT signals output from an IMCDT unit.

12 . The decoding device according to claim 1 , wherein

a signal of each of the sub-bands of high frequency is generated based on a signal of the plurality of low frequency sub-bands and a power value of each of the sub-bands of high frequency.

13 . The decoding device according to claim 1 , wherein

a target high frequency sub-band signal is generated by adjusting power of a low frequency sub-band signal of a predetermined sub-band by a power of a target sub-band of high frequency or by shifting the frequency thereof.

14 . A method executed by at least one processing circuit of a decoding device, the method comprising:

acquiring one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;

acquiring meta-data for each of the plurality of audio objects;

decoding the one or more encoded audio signals according to the meta-data; and

rendering the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:

the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,

a value range of the priority information is from 0 to 7,

the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:

a high-frequency component of a non-encoded audio signal is generated from an audio signal of:

a low frequency component generated from performing overlapping addition, and

a high frequency power value,

the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,

a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,

an audio signal including the high frequency component is generated, and

the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.

15 . A non-transitory computer readable medium storing instructions that, when executed by at least one processing circuit of a decoding device, causes the at least one processing circuit to perform a method comprising:

acquiring one or more encoded audio signals including at least one of a plurality of channels and a plurality of audio objects;

acquiring meta-data for each of the plurality of audio objects;

decoding the one or more encoded audio signals according to the meta-data; and

rendering the plurality of audio objects based on the meta-data by using VBAP (Vector Base Amplitude Panning) wherein:

the meta-data is priority information for each of the plurality of channels and/or the plurality of audio objects,

a value range of the priority information is from 0 to 7,

the circuitry is further configured to, prior to the decoding of the one or more encoded audio signals, perform processing with respect to an obtained audio signal such that in the processing:

a high-frequency component of a non-encoded audio signal is generated from an audio signal of:

a low frequency component generated from performing overlapping addition, and

a high frequency power value,

the obtained audio signal of one frame is divided into sections called time slots, and audio of each of the time slots is band-divided into a signal of a plurality of low frequency sub-bands,

a high-frequency sub-band signal and a low-frequency sub-band signal are synthesized,

an audio signal including the high frequency component is generated, and

the audio signal including the high frequency component generated for each of the time slots are combined, resulting in an audio signal of one time frame including the high frequency component.

Assignments (2)
CHANGE OF NAME Recorded Jan 10, 2024
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 066261/0464 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2024
From: CHINEN, TORU; NISHIGUCHI, MASAYUKI; SHI, RUNYU; HATANAKA, MITSUYUKI; YAMAMOTO, YUKI
To: SONY CORPORATION
Reel/Frame 066265/0814 →
Priority Claims (2)
JP 2014-060486 · Mar 24, 2014 · national
JP 2014-136633 · Jul 2, 2014 · national
Continuity (4)
Continuation 17464594 · Sep 1, 2021
Continuation 16726755 · Dec 24, 2019
Continuation 15127182 · Mar 16, 2015
Related Publication 20240055007A1 · Feb 15, 2024
References Cited (102)
US 5020135A · Kasparian · 1991 [cited by examiner]
US 7328162B2 · Liljeryd · 2008 [cited by examiner]
US 7974422B1 · Ho · 2011 [cited by examiner]
US 8396576B2 · Kraemer · 2013 [cited by examiner]
US 8724830B1 · Fedigan · 2014 [cited by examiner]
US 10748547B2 · Mehta · 2020 [cited by examiner]
US 11188666B2 · Coburn, IV · 2021 [cited by examiner]
US 11900956B2 · Yamamoto · 2024 [cited by examiner]
US 20020031086A1 · Welin · 2002 [cited by examiner]
US 20020057705A1 · Hagai · 2002 [cited by examiner]
US 20050234714A1 · Takagi · 2005 [cited by examiner]
US 20060045295A1 · Kim · 2006 [cited by applicant]
US 20060059535A1 · D'Avello · 2006 [cited by examiner]
US 20070291951A1 · Faller · 2007 [cited by examiner]
US 20080008323A1 · Hilpert · 2008 [cited by examiner]
US 20080082321A1 · Ide · 2008 [cited by examiner]
US 20090024397A1 · Ryu · 2009 [cited by examiner]
US 20100241434A1 · Ono · 2010 [cited by examiner]
US 20110040395A1 · Kraemer · 2011 [cited by examiner]
US 20110040396A1 · Kraemer et al. · 2011 [cited by applicant]
US 20110040397A1 · Kraemer et al. · 2011 [cited by applicant]
US 20110051940A1 · Ishikawa · 2011 [cited by examiner]
US 20110182432A1 · Ishikawa · 2011 [cited by examiner]
US 20110286593A1 · Ho · 2011 [cited by examiner]
US 20120005347A1 · Chen et al. · 2012 [cited by applicant]
US 20120005457A1 · Chen et al. · 2012 [cited by applicant]
US 20120243526A1 · Yamamoto · 2012 [cited by examiner]
US 20130108054A1 · Groh · 2013 [cited by examiner]
US 20130202129A1 · Kraemer · 2013 [cited by examiner]
US 20130223456A1 · Kim · 2013 [cited by examiner]
US 20130329922A1 · Lemieux · 2013 [cited by examiner]
US 20140036999A1 · Ryu · 2014 [cited by examiner]
US 20140056461A1 · Afshar · 2014 [cited by examiner]
US 20140112140A1 · Chan · 2014 [cited by examiner]
US 20140133682A1 · Chabanne · 2014 [cited by examiner]
US 20140350944A1 · Jot et al. · 2014 [cited by applicant]
US 20150201041A1 · Wang · 2015 [cited by examiner]
US 20150213790A1 · Oh · 2015 [cited by examiner]
US 20150332680A1 · Crockett · 2015 [cited by examiner]
US 20150358754A1 · Koppens · 2015 [cited by examiner]
US 20160029138A1 · France · 2016 [cited by examiner]
US 20160050508A1 · Redmann · 2016 [cited by examiner]
US 20160057556A1 · Boehm · 2016 [cited by examiner]
US 20160093289A1 · Pollet · 2016 [cited by examiner]
US 20160133267A1 · Adami · 2016 [cited by examiner]
US 20160174008A1 · Boehm · 2016 [cited by examiner]
US 20160330560A1 · Chon · 2016 [cited by examiner]
US 20170229139A1 · Yamamoto · 2017 [cited by examiner]
US 20180033440A1 · Chinen · 2018 [cited by examiner]
US 20190115037A1 · Choo · 2019 [cited by examiner]
US 20200135216A1 · Chinen et al. · 2020 [cited by applicant]
US 20210360694A1 · Pandian · 2021 [cited by examiner]
US 20210398546A1 · Chinen et al. · 2021 [cited by applicant]
US 20240055007A1 · Chinen · 2024 [cited by examiner]
CN 1272259A · 2000 [cited by applicant]
CN 101529504A · 2009 [cited by applicant]
CN 102549655A · 2012 [cited by applicant]
CN 103649706A · 2014 [cited by applicant]
EP 2465114A1 · 2012 [cited by applicant]
EP 2465259A1 · 2012 [cited by applicant]
EP 2686654A1 · 2014 [cited by applicant]
EP 3059732A1 · 2016 [cited by applicant]
JP 2002516421A · 2002 [cited by applicant]
JP 2002229593A · 2002 [cited by applicant]
JP 2003066994A · 2003 [cited by applicant]
JP 2005031289A · 2005 [cited by applicant]
JP 2007225762A · 2007 [cited by applicant]
JP 2011164638A · 2011 [cited by applicant]
JP 2012133366A · 2012 [cited by applicant]
JP 2012521012A · 2012 [cited by applicant]
JP 2013502184A · 2013 [cited by applicant]
JP 2013511752A · 2013 [cited by applicant]
JP 2013257569A · 2013 [cited by applicant]
JP 2014525048A · 2014 [cited by applicant]
JP 2014526168A · 2014 [cited by applicant]
KR 20120084314A · 2012 [cited by applicant]
WO WO9857436A2 · 1998 [cited by applicant]
WO WO2010109918A1 · 2010 [cited by applicant]
WO WO2011020065A1 · 2011 [cited by applicant]
WO WO2011020067A1 · 2011 [cited by applicant]
WO WO2012125855A1 · 2012 [cited by applicant]
WO WO2013006342A1 · 2013 [cited by applicant]
WO WO2013181272A2 · 2013 [cited by applicant]
International Search Report and Written Opinion thereof mailed Jun. 29, 2015 in connection with International Application No. PCT/JP2015/001432. [cited by applicant]
International Preliminary Report on Patentability mailed Oct. 6, 2016 in connection with International Application No. PCT/JP2015/001432. [cited by applicant]
Chinese Office Action issued Mar. 15, 2019 in connection with Chinese Application No. 201580014248.6, and English translation thereof. [cited by applicant]
Japanese Office Action mailed Jan. 16, 2020 in connection with Japanese Application No. 2018-217178 and English translation thereof. [cited by applicant]
Brazilian Search Report and Written Opinion dated Jun. 30, 2020 in connection with Brazilian Application No. BR112016021407, and partial English translation thereof. [cited by applicant]
Japanese Office Action dated Jun. 24, 2020 in connection with Japanese Application No. 2018-217178, and English translation thereof. [cited by applicant]
[No Author Listed], Information technology—Coding of audio-visual objects—Part 3: Audio, International Standard, ISO/IEC 14496-3, Fourth Edition, Sep. 1, 2009, 1416 pages. [cited by applicant]
[No Author Listed], Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio. ISO/IEC CD 23008-3. Apr. 4, 2014. 339 pages. [cited by applicant]
[No Author Listed], Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding, International Standard, ISO/IEC 23003-3, First Edition, Apr. 1, 2012, 286 pages. [cited by applicant]
Benyassine et al., ITU-T Recommendation G. 719 Annex B: A Silence Compression Scheme for Use with G.729 Optimized for V.70 Digital Simultaneous Voice and Data Applications. IEEE Communications Magazine. Sep. 1997, pp. 6… [cited by applicant]
Herre et al., New concepts in parametric coding of spatial audio: From SAC to SAOC. 2007 IEEE International Conference on Multimedia and Expo Jul. 2, 2007:1894-97. [cited by applicant]
Yamamoto et al., Proposal on Complexity Reduction of the MPEG-H 3D Audio CO Object Renderer, MPEG Meeting, Valencia, Spain, Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11, No. M33137, XP030061589, Mar. 2014. [cited by applicant]
[No Author Listed], “Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”, ISO/IEC DIS 23008-3, ISO/IEC JTC 1/SC 29/WG 11, Jul. 25, 2014. [cited by applicant]
[No Author Listed], “Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”, ISO/IEC DIS 23008-3, ISO/IEC 23008-3, Aug. 5, 2014. [cited by applicant]
[No Author Listed], “Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio”,ISO/IEC WD1 23008-3, ISO/IEC 23008-3, Jan. 24, 2014. [cited by applicant]
Takeshi Norimatsu, “The audible signal coding which unified a sound and musical tone” Journal of the Acoustical Society of Japan, Mar. 2012, the 68th volume, No. 3, pp. 123-128. [cited by applicant]
(No Author Listed), “Information technology—MPEG audio technologies—Part 3: Unified speech and audio coding”. ISO/IEC JTC 1/SC 29/WG 11, ISO/IEC FDIS 23003-3:2011(E), Final Draft International Standard, Sep. 20, 2011. [cited by applicant]
Kazuya Iwata, et al, “Sound Field for Movie Surround Sound Field Generation Control Technology”, Panasonic Technical Journal, vol. 56, No. 4, Jan. 2011, pp. 27-29. [cited by applicant]
Ville Pulkki, “Compensating Displacement of Amplitude-Panned Virtual Sources”, AES 22nd International Conference on Virtual, Synthetic and Entertainment Audio, Nov. 5, 2025 (Nov. 5, 2025), pp. 1-10, XP055119863. [cited by applicant]