IP Library Granted Patent US 12,425,792
Granted Patent B2
US 12,425,792 · App. 18/126,794 · Granted Sep 23, 2025

Video processing device and method

Inventors: Woohyun Nam (Suwon-si, KR); Yoonjae Son (Suwon-si, KR); Hyunkwon Chung (Suwon-si, KR); Sunghee Hwang (Suwon-si, KR)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
H04S7/302G06T7/246G11B27/10H04S3/008H04S7/307G06T2207/10016G06T2207/20081G06T2207/20084H04S2400/01H04S2400/11H04S2400/15H04S2420/11
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,425,792
App. No.
18/126,794
Granted
Sep 23, 2025
Kind
B2
Abstract

A video processing apparatus includes a memory storing instructions, and at least one processor configured to execute the instructions to generate a plurality of feature information by analyzing a video signal comprising a plurality of images based on a first DNN, extract a first altitude component and a first planar component corresponding to a movement of an object in a video from the video signal based on a second DNN, extract a second planar component corresponding to a movement of a sound source in audio from a first audio signal based on a third DNN, generate a second altitude component based on the first altitude component, the first planar component, and the second planar component, output a second audio signal comprising the second altitude component based on the feature information, and synchronize the second audio signal with the video signal and output the synchronized second audio signal and video signal.

Claims (86)

1. A video processing apparatus comprising:

a memory storing at least one instruction; and

at least one processor configured to execute the at least one instruction to:

generate a plurality of feature information for time and frequency by analyzing a video signal comprising a plurality of images, based on a first deep neural network (DNN);

extract a first altitude component corresponding to a movement of an object in a vertical direction in a video and a first planar component corresponding to the movement of the object in a horizontal direction in the video from the video signal, based on a second DNN;

extract a second planar component corresponding to a movement of a sound source in the horizontal direction in audio from a first audio signal, based on a third DNN;

generate a second altitude component corresponding to the movement of the sound source in the vertical direction in the audio based on the first altitude component, the first planar component, and the second planar component;

output a second audio signal comprising the second altitude component, based on the plurality of feature information; and

synchronize the second audio signal with the video signal and output the synchronized second audio signal and video signal.

2. The video processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the at least one instruction to:

synchronize the video signal with the first audio signal;

generate M pieces of one-dimensional image feature map information corresponding to the movement of the object in the video from the video signal by using the first DNN, M being an integer greater than or equal to 1; and

generate the plurality of feature information for time and frequency by performing tiling related to frequency on the M pieces of one-dimensional image feature map information, the plurality of feature information including the M pieces of one-dimensional image feature map information for time and frequency.

3. The video processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the at least one instruction to:

synchronize the video signal with the first audio signal:

extract N+M pieces of feature map information corresponding to the movement of the object in the horizontal direction in the video with respect to time from the video signal by using a (2-1st) DNN, N and M being integers greater than or equal to 1;

extract N+M pieces of feature map information corresponding to the movement of the object in the vertical direction in the video with respect to time from the video signal by using a (2-2nd) DNN, wherein the (2-1st) DNN and the (2-2nd) DNN are included in the second DNN and are different from each other;

extract N+M pieces of feature map information corresponding to the movement of the sound source in the horizontal direction in the audio from the first audio signal by using the third DNN;

generate N+M pieces of correction map information with respect to time corresponding to the second altitude component based on the N+M pieces of feature map information corresponding to the movement of the object in the horizontal direction in the video, the N+M pieces of feature map information corresponding to the movement of the object in the vertical direction in the video, and the N+M pieces of feature map information corresponding to the movement of the sound source in the horizontal direction in the audio; and

generate N+M pieces of correction map information with respect to time and frequency corresponding to the second altitude component by performing tiling related to frequency on the N+M pieces of correction map information with respect to time.

4. The video processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the at least one instruction to:

generate time and frequency information for a 2-channel by performing frequency conversion operation on the first audio signal;

generate N pieces of audio feature map information with respect to time and frequency from the time and frequency information for the 2-channel by using a (4-1st) DNN, N being an integer greater than or equal to 1;

generate N+M pieces of audio and image integrated feature map information based on M pieces of image feature map information with respect to time and frequency included in the plurality of feature information for time and frequency and the N pieces of audio feature map information with respect to time and frequency;

generate a frequency domain second audio signal for n-channel (where, n is an integer greater than 2) from the N+M pieces of audio and image integrated feature map information by using a (4-2nd) DNN;

generate an audio correction map information for the n-channel from N+M pieces of correction map information with respect to time and frequency corresponding to the N+M pieces of audio/image integrated feature map information and the second altitude component by using a (4-3rd) DNN;

generate a corrected frequency domain second audio signal for the n-channel by performing correction on the frequency domain second audio signal for the n-channel based on the audio correction map information for the n-channel; and

output the second audio signal for the n-channel by inversely frequency converting the corrected frequency domain second audio signal for the n-channel, and

wherein the (4-1st) DNN, the (4-2nd) DNN and the (4-3rd) DNN are included in a fourth DNN for outputting the second audio signal and are different from each other.

5. The video processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the at least one instruction to output the second audio signal based on a fourth DNN for outputting the second audio signal,

wherein the first DNN is a DNN for generating the plurality of feature information for time and frequency, the second DNN is a DNN for extracting the first altitude component and the first planar component, the third DNN is a DNN for extracting the second planar component, and

wherein the at least one processor is further configured to execute the at least one instruction to train the first DNN, the second DNN, the third DNN and the fourth DNN according to a result of comparison of a first frequency domain training reconstruction three-dimensional audio signal reconstructed based on a first training two-dimensional audio signal and a first training image signal with a first frequency domain training three-dimensional audio signal obtained by frequency converting a first training three-dimensional audio signal.

6. The video processing apparatus of claim 5 , wherein the at least one processor is further configured to execute the at least one instruction to:

determine generation loss information by comparing the first frequency domain training reconstruction three-dimensional audio signal with the first frequency domain training three-dimensional audio signal, and

update parameters of the first DNN, the second DNN, the third DNN and the fourth DNN based on the generation loss information.

7. The video processing apparatus of claim 1 , wherein the at least one processor is further configured to execute the at least one instruction to output the second audio signal based on a fourth DNN for outputting the second audio signal,

wherein the first DNN is a DNN for generating the plurality of feature information for time and frequency, the second DNN is a DNN for extracting the first altitude component and the first planar component, and the third DNN is a DNN for extracting the second planar component, and

wherein the at least one processor is further configured to execute the at least one instruction to train the first DNN, the second DNN, the third DNN, and the fourth DNN according to a result of comparison of a frequency domain training reconstruction three-dimensional audio signal reconstructed based on a first training two-dimensional audio signal, a first training image signal and a user input parameter information with a first frequency domain training three-dimensional audio signal obtained by frequency converting a first training three-dimensional audio signal.

8. The video processing apparatus of claim 7 , wherein the at least one processor is further configured to execute the at least one instruction to:

determine generation loss information by comparing the frequency domain training reconstruction three-dimensional audio signal with the first frequency domain training three-dimensional audio signal, and

update parameters of the first DNN, the second DNN, the third DNN, and the fourth DNN based on the generation loss information.

9. The video processing apparatus of claim 5 , wherein the first training two-dimensional audio signal and the first training image signal are obtained from a portable terminal, and

wherein the first training three-dimensional audio signal is obtained from an ambisonic microphone of the portable terminal.

10. The video processing apparatus of claim 5 , wherein parameter information of the first DNN, the second DNN, the third DNN, and the fourth DNN obtained as a result of training of the first DNN, the second DNN, the third DNN, and the fourth DNN is stored in the video processing apparatus or is received from a terminal connected to the video processing apparatus.

11. A video processing method of a video processing apparatus, the video processing method comprising:

generating a plurality of feature information for time and frequency by analyzing a video signal comprising a plurality of images based on a first deep neural network (DNN);

extracting a first altitude component corresponding to a movement of an object in a vertical direction in a video and a first planar component corresponding to the movement of the object in a horizontal direction in the video, from the video signal based on a second DNN;

extracting a second planar component corresponding a movement of a sound source in the horizontal direction in an audio from a first audio signal, based on a third DNN;

generating a second altitude component corresponding to the movement of the sound source in the vertical direction in the audio based on the first altitude component, the first planar component, and the second planar component;

outputting a second audio signal comprising the second altitude component, based on the plurality of feature information; and

synchronizing the second audio signal with the video signal and outputting the synchronized second audio signal and video signal.

12. The video processing method of claim 11 , wherein the generating the plurality of feature information for time and frequency comprises:

synchronizing the video signal with the first audio signal;

generating M pieces of one-dimensional image feature map information corresponding to the movement of the object in the video from the video signal by using the first DNN, M being an integer greater than or equal to 1; and

generating the plurality of feature information for time and frequency by performing tiling related to frequency on the M pieces of one-dimensional image feature map information, the plurality of feature information including M pieces of image feature map information for time and frequency.

13. The video processing method of claim 11 , wherein the extracting the first altitude component and the first planar component based on the second DNN and the extracting of the second planar component based on the third DNN comprise:

synchronizing the video signal with the first audio signal;

extracting N+M pieces of feature map information corresponding to the movement of the object in the horizontal direction in the video with respect to time from the video signal by using a (2-1st) DNN, N and M being integers greater than or equal to 1;

extracting N+M pieces of feature map information corresponding to the movement of the object in the vertical direction in the video with respect to time from the video signal by using a (2-2nd) DNN, wherein the (2-1st) DNN and the (2-2nd) DNN are included in the second DNN and are different from each other;

extracting N+M pieces of feature map information corresponding to the movement of the sound source in the horizontal direction in the audio from the first audio signal by using the third DNN, and

wherein the generating of the second altitude component based on the first altitude component, the first planar component, and the second planar component comprises:

generating N+M pieces of correction map information with respect to time corresponding to the second altitude component based on the N+M pieces of feature map information corresponding to the movement of the object in the horizontal direction in the video, the N+M pieces of feature map information corresponding to the movement of the object in the vertical direction, and the N+M pieces of feature map information corresponding to the movement of the sound source in the horizontal direction in the audio; and

generating N+M pieces of correction map information with respect to time and frequency corresponding to the second altitude component by performing tiling related to a frequency on the N+M pieces of correction map information with respect to time.

14. The video processing method of claim 11 , wherein the outputting the second audio signal comprising the second altitude component based on the plurality of feature information comprises:

obtaining time and frequency information for a 2-channel by performing frequency conversion operation on the first audio signal;

generating, from the time and frequency information for the 2-channel, N pieces of audio feature map information with respect to time and frequency by using a (4-1st) DNN, N being an integer greater than or equal to 1;

generating N+M pieces of audio and image integrated feature map information based on M pieces of image feature map information with respect to time and frequency included in the plurality of feature information for time and frequency and the N pieces of audio feature map information with respect to time and frequency;

generating a frequency domain second audio signal for n-channel (wherein, n is an integer greater than 2) from the N+M pieces of audio and image integrated feature map information by using a (4-2nd) DNN;

generating a audio correction map information with respect to the n-channel corresponding to the second altitude component from the N+M pieces of audio/image integrated feature map information by using a (4-3rd) DNN;

generating a corrected frequency domain second audio signal for the n-channel by performing correction on the frequency domain second audio signal for the n-channel based on the audio correction map information for the n-channel; and

outputting the second audio signal for the n-channel by inversely frequency converting the corrected frequency domain second audio signal, and

wherein the (4-1st) DNN, the (4-2nd) DNN and the (4-3rd) DNN are included in a fourth DNN for outputting the second audio signal and are different from each other.

15. The video processing method of claim 11 , wherein the outputting of the second audio signal comprises outputting the second audio signal based on a fourth DNN for outputting the second audio signal,

wherein the first DNN is a DNN for generating the plurality of feature information for each time and frequency, the second DNN is a DNN for extracting the first altitude component and the first planar component, and the third DNN is a DNN for extracting the second planar component, and

wherein the video processing method further comprises training the first DNN, the second DNN, the third DNN, and the fourth DNN according to a result of comparison of a first frequency domain training reconstruction three-dimensional audio signal reconstructed based on a first training two-dimensional audio signal and a first training image signal with a first frequency domain training three-dimensional audio signal obtained by frequency converting a first training three-dimensional audio signal.

16. The video processing method of claim 15 , further comprising:

determining generation loss information by comparing the first frequency domain training reconstruction three-dimensional audio signal with the first frequency domain training three-dimensional audio signal, and

updating parameters of the first DNN, the second DNN, the third DNN, and the fourth DNN based on the generation loss information.

17. The video processing apparatus of claim 15 , wherein parameter information of the first DNN, the second DNN, the third DNN, and the fourth DNN obtained as a result of training of the first DNN, the second DNN, the third DNN, and the fourth DNN is stored in the video processing apparatus or is received from a terminal connected to the video processing apparatus.

18. The video processing method of claim 11 , wherein the outputting the second audio signal comprises outputting the second audio signal based on a fourth DNN for outputting the second audio signal,

wherein the first DNN is a DNN for generating the plurality of feature information for each time and frequency, the second DNN is a DNN for extracting the first altitude component and the first planar component, and the third DNN is a DNN for extracting the second planar component, and

wherein the method further comprises training the first DNN, the second DNN, the third DNN, and the fourth DNN according to a result of comparison of a first frequency domain training reconstruction three-dimensional audio signal reconstructed based on a first training two-dimensional audio signal, a first training image signal and user input information with a first frequency domain training three-dimensional audio signal obtained by frequency converting a first training three-dimensional audio signal.

19. The video processing method of claim 18 , further comprising:

determining generation loss information by comparing the first frequency domain training reconstruction three-dimensional audio signal with the first frequency domain training three-dimensional audio signal, and

updating parameters of the first to fourth DNNs based on the generation loss information.

20. A non-transitory computer-readable recording medium having recorded thereon a program that is executable by a processor to perform the method of claim 11 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2023
From: NAM, WOOHYUN; SON, YOONJAE; CHUNG, HYUNKWON; HWANG, SUNGHEE
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 063132/0024 →
Priority Claims (2)
KR 10-2020-0126361 · Sep 28, 2020 · national
KR 10-2021-0007681 · Jan 19, 2021 · national
Continuity (2)
Continuation PCTKR2021013231 · Sep 28, 2021
Related Publication 20230239643A1 · Jul 27, 2023
References Cited (27)
US 7590249B2 · Jang et al. · 2009 [cited by applicant]
US 9473870B2 · Sen · 2016 [cited by applicant]
US 9723287B2 · Jeong et al. · 2017 [cited by applicant]
US 9888333B2 · Zurek et al. · 2018 [cited by applicant]
US 10419867B2 · Seo et al. · 2019 [cited by applicant]
US 20160140980A1 · Disch · 2016 [cited by examiner]
US 20180054689A1 · Chen et al. · 2018 [cited by applicant]
US 20190306451A1 · Wang · 2019 [cited by examiner]
US 20200288255A1 · Jung et al. · 2020 [cited by applicant]
US 20200288256A1 · Jung et al. · 2020 [cited by applicant]
JP 201171685A · 2011 [cited by applicant]
JP 2020144574A · 2020 [cited by applicant]
JP 7116424A · 2022 [cited by applicant]
KR 100542129B1 · 2006 [cited by applicant]
KR 1020150032253A · 2015 [cited by applicant]
KR 101516644B1 · 2015 [cited by applicant]
KR 1020200107757A · 2020 [cited by applicant]
WO 2017126895A1 · 2017 [cited by applicant]
Yu Jin Lee et al., “A Personal Video Event Classification Method based on Multi-Modalities by DNN-Learning”, The Korean Institute of Information Scientists and Engineers, pp. 1281-1297, 2016. [cited by applicant]
“Development of sound source object separation/location estimation and 3D rendering software technology for converting 2D stereo content into 3D stereo sound content”, Apr. 14, 2016, (102 pages). [cited by applicant]
Chai-Jong Song et al., “Sound source object position estimation technology for converting 2D stereo content into 3D stereo sound content”, The Korean Institute of Electrical Engineers, 2014, (4 pages). [cited by applicant]
Communication dated Feb. 28, 2022, issued by the Korean Intellectual Property Office in Korean Patent Application No. 10-2021-0007681. [cited by applicant]
Ruohan Gao et al., “2.5D Visual Sound”, pp. 324-333, 2019. [cited by applicant]
International Search Report (PCT/ISA/210) issued by the International Searching Authority on Jan. 5, 2022 in International Application No. PCT/KR2021/013231. [cited by applicant]
Written Opinion (PCT/ISA/237) issued by the International Searching Authority on Jan. 5, 2022 in International Application No. PCT/KR2021/013231. [cited by applicant]
Communication dated Feb. 8, 2024 issued by the European Patent Office in European Application No. 21873009.1. [cited by applicant]
Senocak et al., “Learning to Localize Sound Source in Visual Scenes”, 2018 IEEE/CVF Conference On Computer Vision and Pattern Recognition, IEEE, Jun. 18, 2018, pp. 4358-4366 (9 pages total). [cited by applicant]