IP Library Granted Patent US 12,256,169
Granted Patent B2
US 12,256,169 · App. 18/407,825 · Granted Mar 18, 2025

Apparatus and method for video-audio processing, and program for separating an object sound corresponding to a selected video object

Inventors: Hiroyuki Honma (Chiba, JP); Yuki Yamamoto (Tokyo, JP)
Assignee: Sony Group Corporation
H04N5/9202G06V20/46G06V40/16G06V40/161G10L19/00G10L19/008G10L21/0272G11B27/3081H04N9/802H04N19/46H04R1/40H04R3/00G06F2218/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,256,169
App. No.
18/407,825
Granted
Mar 18, 2025
Kind
B2
Abstract

The present technique relates to an apparatus and a method for video-audio processing, and a program each of which enables a desired object sound to be more simply and accurately separated. A video-audio processing apparatus includes a display control portion configured to cause a video object based on a video signal to be displayed; an object selecting portion configured to select the predetermined video object from the one video object or among a plurality of the video objects; and an extraction portion configured to extract an audio signal of the video object selected by the object selecting portion as an audio object signal. The present technique can be applied to a video-audio processing apparatus.

Claims (41)

1. A video-audio processing apparatus, comprising:

processing circuitry and a memory containing instructions that, when executed by the processing circuitry, are configured to:

cause one or more video objects, based on a video signal, to be displayed in an image;

select a video object from the one or more video objects;

extract an audio object signal of the selected video object from an audio signal; and

produce metadata of the selected video object, the metadata including spread information indicating a spatial spread of an area of the selected video object, wherein the audio object signal of the selected video object is reproduced based on the spatial spread of the selected video object,

wherein the spread information is produced based on a frame image surrounding the selected video object, and

wherein the spread information is produced based on an angle between a first vector from an origin to a center of the frame image and a second vector from the origin to a side of the frame image.

2. The video-audio processing apparatus according to claim 1 , wherein

the processing circuitry is configured to extract a signal other than the audio object signal of the selected video object as a background sound signal from the audio signal.

3. The video-audio processing apparatus according to claim 1 , wherein

the processing circuitry is configured to produce object position information indicating a position, on a space, of the selected video object, and to extract the audio object signal based on the object position information.

4. The video-audio processing apparatus according to claim 3 , wherein

the processing circuitry is configured to extract the audio object signal through sound source separation using the object position information.

5. The video-audio processing apparatus according to claim 4 , wherein

the processing circuitry is configured to carry out fixed beam forming as the sound source separation.

6. The video-audio processing apparatus according to claim 1 , wherein the processing circuitry is further configured to recognize the video object based on the video signal, and to cause an image based on a recognition result of the video object to be displayed together with the video object.

7. The video-audio processing apparatus according to claim 6 , wherein

the processing circuitry is configured to recognize the video object by face recognition.

8. The video-audio processing apparatus according to claim 1 , wherein

the processing circuitry is configured to include object position information indicating a position, on a space, of the selected video object in the metadata.

9. The video-audio processing apparatus according to claim 1 , wherein

the processing circuitry is configured to include a processing priority of the selected video object in the metadata.

10. The video-audio processing apparatus according to claim 1 , wherein the processing circuitry is further configured to encode the audio object signal and the metadata.

11. The video-audio processing apparatus according to claim 10 , wherein the processing circuitry is further configured to encode the video signal, and to multiplex a video bit stream obtained by encoding the video signal, and an audio bit stream obtained by encoding the audio object signal and the metadata.

12. The video-audio processing apparatus according to claim 1 , further comprising an image pickup device configured to obtain the video signal by photographing.

13. The video-audio processing apparatus according to claim 1 , further comprising a sound acquisition device configured to obtain the audio signal by sound acquisition.

14. A video-audio processing method executed by processing circuitry, the method comprising:

causing one or more video objects, based on a video signal, to be displayed in an image;

selecting a video object from the one or more video objects;

extracting an audio object signal of the selected video object from an audio signal; and

producing metadata of the selected video object, the metadata including spread information indicating a spatial spread of an area of the selected video object, wherein the audio object signal of the selected video object is reproduced based on the spatial spread of the selected video object,

wherein the spread information is produced based on a frame image surrounding the selected video object, and

wherein the spread information is produced based on an angle between a first vector from an origin to a center of the frame image and a second vector from the origin to a side of the frame image.

15. A non-transitory computer readable medium containing instructions that, when executed by processing circuitry, perform a video-audio processing method comprising:

causing one or more video objects, based on a video signal, to be displayed in an image;

selecting a video object from the one or more video objects;

extracting an audio object signal of the selected video object from an audio signal; and

producing metadata of the selected video object, the metadata including spread information indicating a spatial spread of an area of the selected video object, wherein the audio object signal of the selected video object is reproduced based on the spatial spread of the selected video object,

wherein the spread information is produced based on a frame image surrounding the selected video object, and

wherein the spread information is produced based on an angle between a first vector from an origin to a center of the frame image and a second vector from the origin to a side of the frame image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2024
From: HONMA, HIROYUKI; YAMAMOTO, YUKI
To: SONY CORPORATION
Reel/Frame 066914/0776 →
CHANGE OF NAME Recorded Mar 27, 2024
From: SONY CORPORATION
To: SONY GROUP CORPORATION
Reel/Frame 066915/0044 →
Priority Claims (1)
JP 2016-107042 · May 30, 2016 · national
Continuity (3)
Continuation 17527578 · Nov 16, 2021
Continuation 16303331
Related Publication 20240146867A1 · May 2, 2024
References Cited (116)
US 4653102A · Hansen · 1987 [cited by applicant]
US 5793875A · Lehr et al. · 1998 [cited by applicant]
US 8363848B2 · Huang et al. · 2013 [cited by applicant]
US 8547416B2 · Ozawa · 2013 [cited by applicant]
US 9008320B2 · Sakagami · 2015 [cited by applicant]
US 9247192B2 · Lee et al. · 2016 [cited by applicant]
US 9330673B2 · Cho et al. · 2016 [cited by applicant]
US 9621919B2 · Wang et al. · 2017 [cited by applicant]
US 9729994B1 · Eddins et al. · 2017 [cited by applicant]
US 9749738B1 · Adsumilli et al. · 2017 [cited by applicant]
US 9756421B2 · Hsu et al. · 2017 [cited by applicant]
US 9826211B2 · Sawa et al. · 2017 [cited by applicant]
US 10134414B1 · Feng et al. · 2018 [cited by applicant]
US 10200804B2 · Chen et al. · 2019 [cited by applicant]
US 10206030B2 · Matsumoto et al. · 2019 [cited by applicant]
US 10225650B2 · Tatematsu et al. · 2019 [cited by applicant]
US 10635383B2 · Makinen et al. · 2020 [cited by applicant]
US 11184579B2 · Honma · 2021 [cited by examiner]
US 11902704B2 · Honma · 2024 [cited by examiner]
US 20010055318A1 · Obata et al. · 2001 [cited by applicant]
US 20030053680A1 · Lin et al. · 2003 [cited by applicant]
US 20040013038A1 · Kajala et al. · 2004 [cited by applicant]
US 20050140810A1 · Ozawa · 2005 [cited by applicant]
US 20050228665A1 · Kobayashi et al. · 2005 [cited by applicant]
US 20060193350A1 · Chen · 2006 [cited by applicant]
US 20090066798A1 · Oku et al. · 2009 [cited by applicant]
US 20090154499A1 · Yamakage et al. · 2009 [cited by applicant]
US 20090174805A1 · Alberth, Jr. et al. · 2009 [cited by applicant]
US 20100123785A1 · Chen · 2010 [cited by examiner]
US 20100182501A1 · Sato et al. · 2010 [cited by applicant]
US 20100254543A1 · Kjolerbakken · 2010 [cited by applicant]
US 20100302401A1 · Oku et al. · 2010 [cited by applicant]
US 20100315528A1 · Goh et al. · 2010 [cited by applicant]
US 20110085061A1 · Kim · 2011 [cited by applicant]
US 20110317023A1 · Tsuda et al. · 2011 [cited by applicant]
US 20120076304A1 · Suzuki · 2012 [cited by applicant]
US 20120124603A1 · Amada · 2012 [cited by applicant]
US 20120155703A1 · Hernandez-Abrego et al. · 2012 [cited by applicant]
US 20120163610A1 · Sakagami · 2012 [cited by applicant]
US 20120218377A1 · Oku · 2012 [cited by applicant]
US 20120263315A1 · Hiroe · 2012 [cited by applicant]
US 20120327115A1 · Chhetri et al. · 2012 [cited by applicant]
US 20130021502A1 · Oku et al. · 2013 [cited by applicant]
US 20130151249A1 · Nakadai et al. · 2013 [cited by applicant]
US 20130218570A1 · Imoto et al. · 2013 [cited by applicant]
US 20130272548A1 · Visser et al. · 2013 [cited by applicant]
US 20130342731A1 · Lee et al. · 2013 [cited by applicant]
US 20140085538A1 · Kaine et al. · 2014 [cited by applicant]
US 20140086551A1 · Kaneko · 2014 [cited by examiner]
US 20140211969A1 · Kim et al. · 2014 [cited by applicant]
US 20140233917A1 · Xiang · 2014 [cited by examiner]
US 20140244880A1 · Soffer · 2014 [cited by applicant]
US 20140362253A1 · Kim et al. · 2014 [cited by applicant]
US 20150016641A1 · Ugur et al. · 2015 [cited by applicant]
US 20150054943A1 · Zad Issa et al. · 2015 [cited by applicant]
US 20150162019A1 · An et al. · 2015 [cited by applicant]
US 20150237455A1 · Mitra et al. · 2015 [cited by applicant]
US 20150281832A1 · Kishimoto et al. · 2015 [cited by applicant]
US 20150281833A1 · Shigenaga et al. · 2015 [cited by applicant]
US 20150296317A1 · Park et al. · 2015 [cited by applicant]
US 20150312662A1 · Kishimoto et al. · 2015 [cited by applicant]
US 20150341735A1 · Kitazawa · 2015 [cited by applicant]
US 20160064000A1 · Mizumoto et al. · 2016 [cited by applicant]
US 20160105478A1 · Oyman · 2016 [cited by applicant]
US 20160142620A1 · Sawa et al. · 2016 [cited by applicant]
US 20160234593A1 · Matsumoto et al. · 2016 [cited by applicant]
US 20160249134A1 · Wang et al. · 2016 [cited by applicant]
US 20170019744A1 · Matsumoto et al. · 2017 [cited by applicant]
US 20170180882A1 · Lunner et al. · 2017 [cited by applicant]
US 20170215005A1 · Hsu et al. · 2017 [cited by applicant]
US 20170264999A1 · Fukuda et al. · 2017 [cited by applicant]
US 20170364752A1 · Zhou et al. · 2017 [cited by applicant]
US 20180084365A1 · Ugur et al. · 2018 [cited by applicant]
US 20180158446A1 · Miyamoto et al. · 2018 [cited by applicant]
US 20180270571A1 · Di Censo et al. · 2018 [cited by applicant]
US 20180341455A1 · Ivanov et al. · 2018 [cited by applicant]
US 20190037283A1 · Krauss · 2019 [cited by examiner]
US 20190222798A1 · Honma et al. · 2019 [cited by applicant]
US 20220078371A1 · Honma et al. · 2022 [cited by applicant]
CN 1323139A · 2001 [cited by applicant]
CN 101727908A · 2010 [cited by applicant]
CN 102549655A · 2012 [cited by applicant]
CN 106463128A · 2017 [cited by applicant]
JP H11242499A · 1999 [cited by applicant]
JP 2000357000A · 2000 [cited by applicant]
JP 2002185573A · 2002 [cited by applicant]
JP 2008193196A · 2008 [cited by applicant]
JP 2008271157A · 2008 [cited by applicant]
JP 2009156888A · 2009 [cited by applicant]
JP 2010193476A · 2010 [cited by applicant]
JP 2010233173A · 2010 [cited by applicant]
JP 2011069948A · 2011 [cited by applicant]
JP 2012015651A · 2012 [cited by applicant]
JP 2013106298A · 2013 [cited by applicant]
JP 2013183315A · 2013 [cited by applicant]
JP 2013254433A · 2013 [cited by applicant]
JP 2014086551A · 2014 [cited by applicant]
JP 2014207589A · 2014 [cited by applicant]
JP 2015226104A · 2015 [cited by applicant]
JP 2016051081A · 2016 [cited by applicant]
KR 20040037437A · 2004 [cited by applicant]
KR 20080013827A · 2008 [cited by applicant]
KR 20110121304A · 2011 [cited by applicant]
KR 20150117693A · 2015 [cited by applicant]
WO WO2014159272A1 · 2014 [cited by applicant]
WO WO2016138168A1 · 2016 [cited by applicant]
†U.S. Appl. No. 16/303,331, filed May 17, 2017, Honma et al. [cited by applicant]
\U.S. Appl. No. 17/527,578, filed Nov. 20, 2018, Honma et al. [cited by applicant]
International Search Report and English translation thereof mailed Aug. 1, 2017 in connection with International Application No. PCT/JP2017/018499. [cited by applicant]
Written Opinion and English translation thereof mailed Aug. 1, 2017 in connection with International Application No. PCT/JP2017/018499. [cited by applicant]
International Preliminary Report on Patentability and English translation thereof mailed Dec. 13, 2018 in connection with International Application No. PCT/JP2017/018499. [cited by applicant]
Partial Supplementary European Search Report dated May 2, 2019 in connection with European Application No. 17806378.0. [cited by applicant]
Extended European Search Report issued Aug. 23, 2019 in connection with European Application No. 17806378.0. [cited by applicant]
No Author Listed, Information technology—High efficiency coding and media delivery in heterogeneous environments—Part 3: 3D audio, International Standard, ISO/IEC 23008-3, First edition Oct. 15, 2015, Corrected version … [cited by applicant]
Jurgen Herre et al., MPEG-H 30 Audio—The New Standard for Coding of Immersive Spatial Audio. IEEE Journal of Selected Topics in Signal Processing. Aug. 1, 2015;9(5):770-9. Doi: 10.1109/JSTSP.2015.2411578. [cited by applicant]
Zhao J, Digital Television Technology. Jan. 31, 2016, University of Electronic Science and Technology, Xidian University Press, 27 pages. [cited by applicant]