Audio processing method and apparatus
Embodiments of this application provide an audio processing method and apparatus. The method includes: A terminal device displays a first interface; when the terminal device receives an operation on a control for enabling recording, the terminal device displays a second interface and obtains a first audio signal; the terminal device performs sound source separation on the first audio signal to obtain N channels of audio signals, where N is an integer greater than or equal to 2; and the terminal device generates a first video and a second video, where when the N channels of audio signals satisfy a preset condition, the second video is obtained based on the N channels of audio signals and a second picture, and a target audio signal is an audio signal of a target object.
1 . An audio processing method, comprising:
displaying, by a terminal device, a first interface, wherein the first interface comprises: a control for enabling recording;
when the terminal device receives an operation on the control for enabling recording, displaying, by the terminal device, a second interface and obtaining a first audio signal, wherein the second interface comprises a first picture and a second picture, the second picture is overlaid on the first picture, and the first picture comprises content in the second picture, and the second picture comprises a target object;
performing, by the terminal device, sound source separation on the first audio signal to obtain N channels of audio signals, wherein Nis an integer greater than or equal to 2; and
generating, by the terminal device, a first video and a second video, wherein
the first video is obtained based on the first audio signal and the first picture, when the N channels of audio signals satisfy a preset condition, the second video is obtained based on the N channels of audio signals and the second picture, a second audio signal corresponding to the second video is obtained by processing a target audio signal in the N channels of audio signals and/or a signal other than the target audio signal in the N channels of audio signals, and the target audio signal is an audio signal of the target object,
wherein when the N channels of audio signals do not satisfy the preset condition, the second video is obtained based on the first audio signal and the second picture, and
wherein that the N channels of audio signals do not satisfy the preset condition comprises:
that energy of any one of the N channels of audio signals is greater than an energy threshold, and angle variance corresponding to an angle of the any audio signal within a time threshold is greater than a variance threshold; and/or
that the energy of the any audio signal is greater than the energy threshold and correlation of the any audio signal with another audio signal in the N channels of audio signals is greater than or equal to a correlation threshold.
2 . The method according to claim 1 , wherein the angle of the any audio signal is obtained based on column data corresponding to the any audio signal in a demixing matrix and a transfer function of the terminal device at each preset angle, wherein the demixing matrix is obtained by the terminal device performing sound source separation on the first audio signal.
3 . The method according to claim 2 , wherein when a quantity of microphones in the terminal device is 2, a range of the preset angle is: 0°-180° or 180°-360°.
4 . The method according to claim 3 , wherein the when the terminal device receives an operation on the control for enabling recording, displaying, by the terminal device, a second interface and obtaining a first audio signal comprises:
when the terminal device receives the operation on the control for enabling recording, displaying, by the terminal device, a third interface, wherein the third interface comprises the first picture, and the first picture comprises the target object; and
when the terminal device receives an operation on the target object, displaying, by the terminal device, the second interface.
5 . The method according to claim 4 , wherein the second interface comprises: a control for ending recording, wherein the performing, by the terminal device, sound source separation on the first audio signal to obtain N channels of audio signals comprises:
when the terminal device receives an operation on the control for ending recording, performing, by the terminal device, sound source separation on the first audio signal to obtain the N channels of audio signals.
6 . The method according to claim 1 , wherein that the N channels of audio signals satisfy a preset condition comprises: that the energy of the any audio signal is greater than the energy threshold, and the angle variance corresponding to the angle of the any audio signal within the time threshold is less than or equal to the variance threshold; and/or
that the energy of the any audio signal is greater than the energy threshold and the correlation of the any audio signal with the another audio signal is less than the correlation threshold.
7 . A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, enables the terminal device to perform the following steps:
displaying, a first interface, wherein the first interface comprises: a control for enabling recording;
when the terminal device receives an operation on the control for enabling recording, displaying, a second interface and obtaining a first audio signal, wherein the second interface comprises a first picture and a second picture, the second picture is overlaid on the first picture, and the first picture comprises content in the second picture, and the second picture comprises a target object;
performing, sound source separation on the first audio signal to obtain N channels of audio signals, wherein N is an integer greater than or equal to 2; and
generating, a first video and a second video, wherein
the first video is obtained based on the first audio signal and the first picture, when the N channels of audio signals satisfy a preset condition, the second video is obtained based on the N channels of audio signals and the second picture, a second audio signal corresponding to the second video is obtained by processing a target audio signal in the N channels of audio signals and/or a signal other than the target audio signal in the N channels of audio signals, and the target audio signal is an audio signal of the target object,
wherein when the N channels of audio signals do not satisfy the preset condition, the second video is obtained based on the first audio signal and the second picture, and
wherein that the N channels of audio signals do not satisfy the preset condition comprises:
that energy of any one of the N channels of audio signals is greater than an energy threshold, and angle variance corresponding to an angle of the any audio signal within a time threshold is greater than a variance threshold; and/or
that the energy of the any audio signal is greater than the energy threshold and correlation of the any audio signal with another audio signal in the N channels of audio signals is greater than or equal to a correlation threshold.
8 . The terminal device according to claim 7 , wherein the angle of the any audio signal is obtained based on column data corresponding to the any audio signal in a demixing matrix and a transfer function of the terminal device at each preset angle, wherein the demixing matrix is obtained by the terminal device performing sound source separation on the first audio signal.
9 . The terminal device according to claim 8 , wherein when a quantity of microphones in the terminal device is 2, a range of the preset angle is: 0°-180° or 180°-360°.
10 . The terminal device according to claim 9 , wherein the when the terminal device receives an operation on the control for enabling recording, displaying, a second interface and obtaining a first audio signal comprises:
when the terminal device receives the operation on the control for enabling recording, displaying, a third interface, wherein the third interface comprises the first picture, and the first picture comprises the target object; and
when the terminal device receives an operation on the target object, displaying, the second interface.
11 . The terminal device according to claim 10 , wherein the second interface comprises: a control for ending recording, wherein the performing, by the terminal device, sound source separation on the first audio signal to obtain N channels of audio signals comprises:
when the terminal device receives an operation on the control for ending recording, performing, sound source separation on the first audio signal to obtain the N channels of audio signals.
12 . The terminal device according to claim 7 , wherein that the N channels of audio signals satisfy a preset condition comprises: that the energy of the any audio signal is greater than the energy threshold, and the angle variance corresponding to the angle of the any audio signal within the time threshold is less than or equal to the variance threshold; and/or
that the energy of the any audio signal is greater than the energy threshold and the correlation of the any audio signal with the another audio signal is less than the correlation threshold.
13 . A non-transitory computer-readable storage medium, comprising instructions, wherein when the instructions are run on a terminal device, the terminal device is enabled to perform the following steps:
displaying, a first interface, wherein the first interface comprises: a control for enabling recording;
when the terminal device receives an operation on the control for enabling recording, displaying, a second interface and obtaining a first audio signal, wherein the second interface comprises a first picture and a second picture, the second picture is overlaid on the first picture, and the first picture comprises content in the second picture, and the second picture comprises a target object;
performing, sound source separation on the first audio signal to obtain N channels of audio signals, wherein N is an integer greater than or equal to 2; and
generating, a first video and a second video, wherein
the first video is obtained based on the first audio signal and the first picture, when the N channels of audio signals satisfy a preset condition, the second video is obtained based on the N channels of audio signals and the second picture, a second audio signal corresponding to the second video is obtained by processing a target audio signal in the N channels of audio signals and/or a signal other than the target audio signal in the N channels of audio signals, and the target audio signal is an audio signal of the target object,
wherein when the N channels of audio signals do not satisfy the preset condition, the second video is obtained based on the first audio signal and the second picture, and
wherein that the N channels of audio signals do not satisfy the preset condition comprises:
that energy of any one of the N channels of audio signals is greater than an energy threshold, and angle variance corresponding to an angle of the any audio signal within a time threshold is greater than a variance threshold; and/or
that the energy of the any audio signal is greater than the energy threshold and correlation of the any audio signal with another audio signal in the N channels of audio signals is greater than or equal to a correlation threshold.
14 . The non-transitory computer-readable storage medium according to claim 13 , wherein the angle of the any audio signal is obtained based on column data corresponding to the any audio signal in a demixing matrix and a transfer function of the terminal device at each preset angle, wherein the demixing matrix is obtained by the terminal device performing sound source separation on the first audio signal.