IP Library Granted Patent US 12,646,322
Granted Patent B2
US 12,646,322 · App. 17/888,215 · Granted Jun 2, 2026

Video-driven spatial audio enhancement

Inventor: Qireng Liang (Shenzhen, CN)
Assignee: Tencent Technology (Shenzhen) Company Limited
G06V20/46G06T7/194G06T7/246G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,322
App. No.
17/888,215
Granted
Jun 2, 2026
Kind
B2
Abstract

A data processing method includes acquiring video frame data including one or more video frames and audio data of a video, and determining position attribute information of a target object in the acquired one or more video frames, the target object being associated with the audio data. The method also includes acquiring a channel encoding parameter associated with the position attribute information, and performing azimuth enhancement processing on the audio data according to the channel encoding parameter to obtain enhanced audio data. Apparatus and non-transitory computer-readable storage medium counterpart embodiments are also contemplated.

Claims (101)

1 . A data processing method, comprising:

acquiring video frame data including one or more video frames and audio data of a video;

determining, by processing circuitry of a data processing apparatus, position attribute information of a target object in the acquired one or more video frames, the target object being associated with the audio data; and

acquiring, from a parameter mapping table that stores a mapping relationship between predefined channel encoding parameters associated with ear positions of a user and relative positions of the target object with respect to the ear positions of the user, a set of channel encoding parameters including a first channel encoding parameter associated with a first one of the ear positions of the user and a second channel encoding parameter associated with a second one of the ear positions of the user for different audio channels corresponding to the position attribute information;

performing azimuth enhancement processing on the audio data based on the set of channel encoding parameters, the azimuth enhancement processing comprising:

processing the audio data based on the first channel encoding parameter to obtain first enhanced audio data;

processing the audio data based on the second channel encoding parameter to obtain second enhanced audio data; and

generating enhanced audio data from the first enhanced audio data and the second enhanced audio data, the first enhanced audio data being louder than the second enhanced audio data to the user based on the position attribute information of the target object indicating that the target object is closer to the first one of the ear positions of the user than the second one of the ear positions of the user.

2 . The method according to claim 1 , wherein the acquiring the video frame data and the audio data comprises:

acquiring the video, inputting the video to a video decapsulation component, and decapsulating the video through the video decapsulation component to obtain video stream data and audio stream data; and

decoding the video stream data and the audio stream data respectively in the video decapsulation component to obtain the video frame data and the audio data.

3 . The method according to claim 1 , wherein the target object is an object in a rest state; and

the determining comprises:

inputting the video frame data to an object identification model, and acquiring N continuous video frames from the object identification model, the N continuous video frames being video frames with continuous timestamps, each of the N continuous video frames comprising the target object, N being a positive integer less than or equal to M, M being a total quantity of video frames in the video frame data, M being an integer greater than 1;

identifying, in the N continuous video frames, video frames in which a sound emitting part of the target object changes, and setting the video frames in which the sound emitting part of the target object changes as changed video frames;

determining a position coordinate of the target object in the changed video frames; and

determining the position attribute information of the target object in the video according to the position coordinate.

4 . The method according to claim 1 , wherein the target object is an object in a motion state; and

the determining comprises:

inputting the video frame data to an object identification model, and identifying a background image in the video frame data through the object identification model;

acquiring a background pixel value of the background image, and acquiring a video frame pixel value corresponding to the video frame data;

determining a difference pixel value between the background pixel value and the video frame pixel value, and determining a region in the one or more video frames where the difference pixel value is located as a position coordinate of the target object; and

determining the position attribute information of the target object in the video according to the position coordinate.

5 . The method according to claim 3 , wherein the determining the position attribute information of the target object in the video according to the position coordinate comprises:

acquiring central position information of a virtual camera, the virtual camera simulating shooting of the target object;

determining a depth of field distance between the target object and the central position information according to the position coordinate;

determining a position offset angle between the target object and the virtual camera; and

determining the depth of field distance and the position offset angle as the position attribute information of the target object.

6 . The method according to claim 1 , wherein

the processing the audio data based on the first channel encoding parameter comprises:

performing convolution processing on the audio data according to the first channel encoding parameter to obtain the first enhanced audio data; and

the processing the audio data based on the second channel encoding parameter comprises:

performing the convolution processing on the audio data according to the second channel encoding parameter to obtain the second enhanced audio data.

7 . The method according to claim 1 , wherein

the processing the audio data based on the first channel encoding parameter to obtain the first enhanced audio data and the processing the audio data based on the second channel encoding parameter to obtain the second enhanced audio data comprises:

performing frequency-domain conversion on the audio data to obtain frequency-domain audio data;

performing the frequency-domain conversion on the first channel encoding parameter and the second channel encoding parameter respectively to obtain a first channel frequency-domain encoding parameter and a second channel frequency-domain encoding parameter;

multiplying the first channel frequency-domain encoding parameter by the frequency-domain audio data to obtain first enhanced frequency-domain audio data;

multiplying the second channel frequency-domain encoding parameter by the frequency-domain audio data to obtain second enhanced frequency-domain audio data; and

the generating the enhanced audio data comprises:

determining the enhanced audio data according to the first enhanced frequency-domain audio data and the second enhanced frequency-domain audio data.

8 . The method according to claim 7 , wherein the determining the enhanced audio data according to the first enhanced frequency-domain audio data and the second enhanced frequency-domain audio data comprises:

performing time-domain conversion on the first enhanced frequency-domain audio data to obtain the first enhanced audio data;

performing time-domain conversion on the second enhanced frequency-domain audio data to obtain the second enhanced audio data; and

determining audio data formed by the first enhanced audio data and the second enhanced audio data as the enhanced audio data.

9 . The method according to claim 6 , wherein the method further comprises:

storing the video frame data in association with the enhanced audio data in a cache server;

acquiring the video frame data and the enhanced audio data from the cache server in response to a video playback operation for the video; and

outputting the video frame data and the enhanced audio data.

10 . The method according to claim 9 , wherein the outputting comprises:

outputting the video frame data;

outputting the first enhanced audio data through a first sound output channel of a user terminal; and

outputting the second enhanced audio data through a second sound output channel of the user terminal.

11 . The method according to claim 1 , wherein

the determining comprises:

inputting the video frame data to an object identification model, and receiving a target object category of the target object and the position attribute information of the target object in the video as output of the object identification model; and

the method further comprises:

inputting the audio data to an audio identification model, and determining, through the audio identification model, a sound-emitting object category to which the audio data belongs;

matching the target object category with the sound-emitting object category to obtain a matching result; and

in response to the matching result indicating that the target object category matches the sound-emitting object category, performing the acquiring the set of channel encoding parameters associated with the position attribute information, and performing the azimuth enhancement processing on the audio data to obtain the enhanced audio data.

12 . A data processing apparatus, comprising:

processing circuitry configured to:

acquire video frame data including one or more video frames and audio data of a video;

determine position attribute information of a target object in the acquired one or more video frames, the target object being associated with the audio data; and

acquire, from a parameter mapping table that stores a mapping relationship between predefined channel encoding parameters associated with ear positions of a user and relative positions of the target object with respect to the ear positions of the user, a set of channel encoding parameters including a first channel encoding parameter associated with a first one of the ear positions of the user and a second channel encoding parameter associated with a second one of the ear positions of the user for different audio channels corresponding to the position attribute information;

perform azimuth enhancement processing on the audio data based on the set of channel encoding parameters, the azimuth enhancement processing comprising:

processing the audio data based on the first channel encoding parameter to obtain first enhanced audio data;

processing the audio data based on the second channel encoding parameter to obtain second enhanced audio data; and

generating enhanced audio data from the first enhanced audio data and the second enhanced audio data, the first enhanced audio data being louder than the second enhanced audio data to the user based on the position attribute information of the target object indicating that the target object is closer to the first one of the ear positions of the user than the second one of the ear positions of the user.

13 . The apparatus according to claim 12 , wherein the processing circuitry is further configured to:

acquire the video, input the video to a video decapsulation component, and decapsulate the video through the video decapsulation component to obtain video stream data and audio stream data; and

decode the video stream data and the audio stream data respectively in the video decapsulation component to obtain the video frame data and the audio data.

14 . The apparatus according to claim 12 , wherein the target object is an object in a rest state; and

the processing circuitry is further configured to:

input the video frame data to an object identification model, and acquire N continuous video frames from the object identification model, the N continuous video frames being video frames with continuous timestamps, each of the N continuous video frames comprising the target object, N being a positive integer less than or equal to M, M being a total quantity of video frames in the video frame data, M being an integer greater than 1;

identify, in the N continuous video frames, video frames in which a sound emitting part of the target object changes, and set the video frames in which the sound emitting part of the target object changes as changed video frames;

determine a position coordinate of the target object in the changed video frames; and

determine the position attribute information of the target object in the video according to the position coordinate.

15 . The apparatus according to claim 12 , wherein the target object is an object in a motion state; and

the processing circuitry is further configured to:

input the video frame data to an object identification model, and identify a background image in the video frame data through the object identification model;

acquire a background pixel value of the background image, and acquire a video frame pixel value corresponding to the video frame data;

determine a difference pixel value between the background pixel value and the video frame pixel value, and determine a region in the one or more video frames where the difference pixel value is located as a position coordinate of the target object; and

determine the position attribute information of the target object in the video according to the position coordinate.

16 . The apparatus according to claim 14 , wherein the processing circuitry is further configured to:

acquire central position information of a virtual camera, the virtual camera simulating shooting of the target object;

determine a depth of field distance between the target object and the central position information according to the position coordinate;

determine a position offset angle between the target object and the virtual camera; and

determine the depth of field distance and the position offset angle as the position attribute information of the target object.

17 . The apparatus according to claim 12 , wherein

the processing circuitry is further configured to:

perform convolution processing on the audio data according to the first channel encoding parameter to obtain the first enhanced audio data; and

perform the convolution processing on the audio data according to the second channel encoding parameter to obtain the second enhanced audio data.

18 . A non-transitory computer-readable storage medium storing computer-readable instructions thereon, which, when executed by a computer, cause the computer to perform a data processing method comprising:

acquiring video frame data including one or more video frames and audio data of a video;

determining position attribute information of a target object in the acquired one or more video frames, the target object being associated with the audio data; and

acquiring, from a parameter mapping table that stores a mapping relationship between predefined channel encoding parameters associated with ear positions of a user and relative positions of the target object with respect to the ear positions of the user, a set of channel encoding parameters including a first channel encoding parameter associated with a first one of the ear positions of the user and a second channel encoding parameter associated with a second one of the ear positions of the user for different audio channels corresponding to the position attribute information;

performing azimuth enhancement processing on the audio data based on the set of channel encoding parameters, the azimuth enhancement processing comprising:

processing the audio data based on the first channel encoding parameter to obtain first enhanced audio data;

processing the audio data based on the second channel encoding parameter to obtain second enhanced audio data; and

generating enhanced audio data from the first enhanced audio data and the second enhanced audio data, the first enhanced audio data being louder than the second enhanced audio data to the user based on the position attribute information of the target object indicating that the target object is closer to the first one of the ear positions of the user than the second one of the ear positions of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: LIANG, QIRENG
To: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
Reel/Frame 060811/0375 →
Priority Claims (1)
CN 202010724466.1 · Jul 24, 2020 · national
Continuity (2)
Continuation PCTCN2021100306 · Jun 16, 2021
Related Publication 20220392224A1 · Dec 8, 2022
References Cited (34)
US 6829018B2 · Lin · 2004 [cited by examiner]
US 9113280B2 · Cho · 2015 [cited by examiner]
US 10672408B2 · Breebaart · 2020 [cited by examiner]
US 11184579B2 · Honma · 2021 [cited by examiner]
US 20030053680A1 · Lin et al. · 2003 [cited by applicant]
US 20140032987A1 · Nagaraj · 2014 [cited by examiner]
US 20140180684A1 · Strub · 2014 [cited by examiner]
US 20160125888A1 · Purnhagen · 2016 [cited by examiner]
US 20160142847A1 · Disch · 2016 [cited by examiner]
US 20170092280A1 · Hirabayashi · 2017 [cited by examiner]
US 20170265016A1 · Oh · 2017 [cited by examiner]
US 20170364752A1 · Zhou · 2017 [cited by examiner]
US 20180233154A1 · Vaillancourt · 2018 [cited by examiner]
US 20180374233A1 · Zhou · 2018 [cited by examiner]
US 20190069110A1 · Gorzel · 2019 [cited by examiner]
US 20190313200A1 · Stein · 2019 [cited by examiner]
US 20200221230A1 · Fuchs · 2020 [cited by examiner]
US 20200412772A1 · Nesta · 2020 [cited by examiner]
US 20210021949A1 · Sridharan · 2021 [cited by examiner]
US 20220279299A1 · Vasilache · 2022 [cited by examiner]
CN 101390443A · 2009 [cited by applicant]
CN 103702180A · 2014 [cited by applicant]
CN 108847248A · 2018 [cited by applicant]
CN 109313904A · 2019 [cited by applicant]
CN 109640112A · 2019 [cited by applicant]
CN 110168638A · 2019 [cited by applicant]
CN 111050269A · 2020 [cited by applicant]
CN 111669696A · 2020 [cited by applicant]
CN 111885414A · 2020 [cited by applicant]
JP 2014195267A · 2014 [cited by applicant]
International Search Report and Written Opinion issued Sep. 15, 2021 in International Application No. PCT/CN2021/100306 with English Translation (9 pages). [cited by applicant]
Chinese Office Action issued Mar. 1, 2021 in International Application No. 202010724466.1 with Concise English Translation (9 pages). [cited by applicant]
Chinese Office Action issued Sep. 1, 2021 in International Application No. 202010724466.1 with Concise English Translation (10 pages). [cited by applicant]
Supplementary European Search Report issued Apr. 17, 2023 in Application No. 21846310.7 (8 pages). [cited by applicant]