IP Library Granted Patent US 11,184,579
Granted Patent B2
US 11,184,579 · App. 16/303,331 · Granted Nov 23, 2021

Apparatus and method for video-audio processing, and program for separating an object sound corresponding to a selected video object

Inventors: Hiroyuki Honma (Chiba, JP); Yuki Yamamoto (Tokyo, JP)
Assignee: Sony Corporation
H04N5/9202G06K9/0057G06K9/00221G06K9/00228G06K9/00744G10L19/00G10L19/008G10L21/0272G11B27/3081H04N9/802H04N19/46H04R1/40H04R3/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,184,579
App. No.
16/303,331
Granted
Nov 23, 2021
Kind
B2
Abstract

The present technique relates to an apparatus and a method for video-audio processing, and a program each of which enables a desired object sound to be more simply and accurately separated. A video-audio processing apparatus includes a display control portion configured to cause a video object based on a video signal to be displayed; an object selecting portion configured to select the predetermined video object from the one video object or among a plurality of the video objects; and an extraction portion configured to extract an audio signal of the video object selected by the object selecting portion as an audio object signal.

Claims (49)

1. A video-audio processing apparatus, comprising:

processing circuitry and a memory containing instructions that, when executed by the processing circuitry, are configured to:

cause one or more video objects, based on a video signal, to be displayed in an image;

select a video object from the one or more video objects; and

extract an audio object signal of the selected video object from an audio signal,

wherein a signal other than the audio object signal of the selected video object is extracted as a background sound signal from the audio signal,

wherein the audio object signal and the background sound signal are encoded independently of each other, and

wherein the audio object signal and the background sound signal are multiplexed to produce the audio signal.

2. The video-audio processing apparatus according to claim 1 , wherein the instructions are configured to

produce object position information indicating the position, in a space, of the selected video object, and

extract the audio object signal based on the object position information.

3. The video-audio processing apparatus according to claim 2 , wherein the instructions are configured to

extract the audio object signal through sound source separation using the object position information.

4. The video-audio processing apparatus according to claim 3 , wherein the instructions are configured to

carry out fixed beam forming as the sound source separation.

5. The video-audio processing apparatus according to claim 1 , wherein the instructions are further configured to recognize the video object based on the video signal, and

to cause an image based on a recognition result of the video object to be displayed together with the video object.

6. The video-audio processing apparatus according to claim 5 , wherein the instructions are configured to

recognize the video object from face recognition.

7. The video-audio processing apparatus according to claim 5 , wherein the instructions are configured to

cause a frame to be displayed as the image in an area of the video object.

8. The video-audio processing apparatus according to claim 1 , wherein the instructions are configured to

produce metadata of the selected video object.

9. The video-audio processing apparatus according to claim 8 , wherein the instructions are configured to

produce object position information indicating the position, in a space, of the selected video object as the metadata.

10. The video-audio processing apparatus according to claim 8 , wherein the instructions are configured to

produce processing priority of the selected video object as the metadata.

11. The video-audio processing apparatus according to claim 8 , wherein the instructions are configured to

produce spread information indicating a spread state of an area of the selected video object as the metadata.

12. The video-audio processing apparatus according to claim 8 , wherein the instructions are further configured to encode the audio object signal and the metadata.

13. The video-audio processing apparatus according to claim 12 , wherein the instructions are further configured to:

encode the video signal; and

multiplex a video bit stream obtained by encoding the video signal, and an audio bit stream obtained by encoding the audio object signal and the metadata.

14. The video-audio processing apparatus according to claim 1 , further comprising an image pickup device configured to obtain the video signal by carrying out photographing.

15. The video-audio processing apparatus according to claim 1 , further comprising a sound acquisition device configured to obtain the audio signal by carrying out sound acquisition.

16. A video-audio processing method, comprising:

causing one or more video objects, based on a video signal, to be displayed in an image;

selecting a video object from the one or more video objects; and

extracting an audio object signal of the selected video object from an audio signal,

wherein a signal other than the audio object signal of the selected video object is extracted as a background sound signal from the audio signal,

wherein the audio object signal and the background sound signal are encoded independently of each other; and

wherein the audio object signal and the background sound signal are multiplexed to produce the audio signal.

17. A non-transitory computer readable medium containing instructions that, when executed by a processing device, perform a video-audio processing method comprising:

causing one or more video objects, based on a video signal, to be displayed in an image;

selecting a video object from the one or more video objects; and

extracting an audio object signal of the selected video object from an audio signal,

wherein a signal other than the audio object signal of the selected video object is extracted as a background sound signal from the audio signal,

wherein the audio object signal and the background sound signal are encoded independently of each other, and

wherein the audio object signal and the background sound signal are multiplexed to produce the audio signal.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2018
From: HONMA, HIROYUKI; YAMAMOTO, YUKI
To: SONY CORPORATION
Reel/Frame 047669/0129 →
Priority Claims (1)
JP JP2016-107042 · May 30, 2016 · national
Continuity (1)
Related Publication 20190222798A1 · Jul 18, 2019
Cited By (3)
US 12,256,169 US 12,646,322 US 12,670,922