IP Library › Granted Patent US 12,267,591
Granted Patent B2
US 12,267,591 · App. 18/024,869 · Granted Apr 1, 2025

Video processing method and related electronic device

Inventors: Zhenyi Liu (Beijing, CN); Jianyong Xuan (Beijing, CN); Haikuan Gao (Beijing, CN)
Assignee: BEIJING HONOR DEVICE CO., LTD.
H04N23/69G10L21/0208H04R3/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,267,591
App. No.
18/024,869
Granted
Apr 1, 2025
Kind
B2
Abstract

This application provides a video processing method and a related electronic device. The video processing method includes: When generating a video, the electronic device may perform image zooming based on a change in a zoom ratio, or may perform audio zooming on an audio based on a change in a zoom ratio. That the electronic device performs audio zooming on the audio includes: When the zoom ratio increases and an angle of view decreases, suppressing a sound of an object outside a photographing range and enhancing a sound of a photographed object within the photographing range; when the zoom ratio decreases and the angle of view increases, suppressing a sound of an object outside the photographing range and weakening a sound of a photographed object within the photographing range.

Claims (190)

1. A video processing method, wherein the method is applied to an electronic device, and comprises:

starting, by the electronic device, a camera;

displaying a preview interface, wherein the preview interface comprises a first control;

detecting a first operation with respect to the first control;

starting photographing in response to the first operation;

displaying a photographing interface, wherein the photographing interface comprises a second control and the second control is used to adjust a zoom ratio;

displaying a first photographed image at a first moment when the zoom ratio is a first zoom ratio;

collecting, by a microphone, a first audio at the first moment;

detecting a third operation with respect to a third control; and

stopping photographing and saving a first video in response to the third operation; and

the method further comprises: processing the first audio to obtain a first left channel output audio and a first right channel output audio, wherein

the processing the first audio to obtain a first left channel output audio and a first right channel output audio comprises:

performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio;

performing second processing on the first audio to obtain M channels of first sound source audios, wherein M represents a quantity of microphones of the electronic device;

fusing the first left channel input audio with a first target audio to obtain a first left channel audio, wherein the first target audio is a sound source audio having highest correlation with the first left channel input audio among the M channels of first sound source audios;

fusing the first right channel input audio with a second target audio to obtain a first right channel audio, wherein the second target audio is a sound source audio having highest correlation with the first right channel input audio among the M channels of first sound source audios; and

performing enhancement processing on the first left channel audio and the first right channel audio to obtain the first left channel output audio and the first right channel output audio.

2. The method according to claim 1 , wherein the first photographed image comprises a first target object and a second target object, and the method further comprises:

detecting a second operation with respect to the second control;

adjusting the zoom ratio to be a second zoom ratio in response to the second operation, wherein the second zoom ratio is greater than the first zoom ratio;

displaying a second photographed image at a second moment, wherein the second photographed image comprises the first target object and does not comprise the second target object;

collecting, by the microphone, a second audio at the second moment, wherein the second audio comprises a first sound corresponding to the first target object and a second sound corresponding to the second target object; and

processing the second audio to obtain a second left channel output audio and a second right channel output audio, wherein the second left channel output audio and the second right channel output audio comprise a third sound and a fourth sound, the third sound corresponds to the first target object, the fourth sound corresponds to the second target object, the third sound is enhanced with respect to the first sound, and the fourth sound is suppressed with respect to the second sound.

3. The method according to claim 2 , wherein the processing the second audio to obtain a second left channel output audio and a second right channel output audio comprises:

performing first processing on the second audio based on the second zoom ratio to obtain a second left channel input audio and a second right channel input audio;

performing second processing on the second audio to obtain M channels of second sound source audios, wherein M represents a quantity of microphones of the electronic device;

fusing the second left channel input audio with a third target audio to obtain a second left channel audio, wherein the third target audio is a sound source audio having highest correlation with the second left channel input audio among the M channels of second sound source audios;

fusing the second right channel input audio with a fourth target audio to obtain a second right channel audio, wherein the fourth target audio is a sound source audio having highest correlation with the second right channel input audio among the M channels of second sound source audios; and

performing enhancement processing on the second left channel audio and the second right channel audio to obtain the second left channel output audio and the second right channel output audio.

4. The method according to claim 1 , wherein the performing second processing on the first audio to obtain M channels of first sound source audios specifically comprises:

obtaining the M channels of first sound source audios through calculation according to the formula Y(ω)=Σ i=1 M W i (ω)x i (ω), wherein

x i (ω)represents an audio signal of a first audio collected by the i th microphone in frequency domain, W i (ω) represents a first non-negative matrix corresponding to the i th microphone, Y(ω) represents a first matrix whose size is M*L, and each row vector of the first matrix is one channel of first sound source audio.

5. The method according to claim 1 , wherein the performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio specifically comprises:

obtaining the first left channel audio according to the formula y i1 (ω)=α 1 *y 1 (ω)+(1−α 1 )*y 2 (ω); and

obtaining the first right channel audio according to the formula y r1 (ω)=α 1 *y 3 (ω) +(1−α 1 )*y 2 (ω), wherein

y l1 represents the first left channel input audio, y r1 (ω) represents the first right channel input audio, α 1 represents a fusion coefficient obtained based on the first zoom ratio, y 1 (ω) represents a first beam obtained based on the first audio and a first filter coefficient, y 2 (ω) represents a second beam obtained based on the first audio and a second filter coefficient, and y 3 (ω) represents a third beam obtained based on the first audio and a third filter coefficient.

6. The method according to claim 1 , before the performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio, further comprising:

obtaining a first beam, a second beam, and a third beam respectively according to the formula y 1 (ω)=Σ i=1 M w 1i (ω)x i1 (ω), the formula y 2 (ω)=Σ i=1 M w 2i (ω)x i1 (ω), and the formula y 3 (ω)=Σ i=1 M w 3i (ω)x i1 (ω), wherein

y 1 (ω) represents the first beam, y 2 (ω) represents the second beam, y 3 (ω) represents the third beam, w 1i (ω) represents a first filter coefficient corresponding to the i th microphone in a first direction, w 2i (ω) represents a second filter coefficient corresponding to the i th microphone in a second direction, w 3i (ω) represents a third filter coefficient corresponding to the i th microphone in a third direction, x i1 (ω) represents the first audio collected by the i th microphone, the first direction is any direction within a range of 10° counterclockwise from the front to 90° counterclockwise from the front of the electronic device, the second direction is any direction within a range of 10° counterclockwise from the front to 10° clockwise from the front of the electronic device, and the third direction is any direction within a range of 10° clockwise from the front to 90° clockwise from the front of the electronic device.

7. The method according to claim 1 ,

before the fusing the first left channel input audio with a first target sound source to obtain a first left channel audio, further comprising:

calculating a correlation value between the first left channel input audio and the M channels of first sound source audios according to the formula

γ

i

=

∅

li

∅

ll

⁢

_

⁢

1

⁢

∅

ii

,

Ø li represents E{y l1 (ω)Y i (ω)*}, Ø ll_1 represents E{y l1 (ω)y l1 (ω)*}, Ø ii represents E{Y i (ω)Y i (ω)*}, γ i represents a correlation value between the first left channel input audio and the i th channel of first sound source audio, y l1 (ω) represents the first left channel input audio, and Y i (ω) represents the i th channel of first sound source audio;

if there is only one maximum correlation value among the M correlation values, determining a first sound source audio having the maximum correlation value as the first target audio; and

if there are a plurality of maximum correlation values among the M correlation values, calculating an average value of first sound source audios corresponding to the plurality of maximum correlation values to obtain the first target audio;

before the fusing the first right channel input audio with a second target sound source to obtain a first right channel audio, further comprising:

calculating a correlation value between the first right channel input audio and the M channels of first sound source audios according to the formula

γ

j

=

∅

rj

∅

rr

⁢

_

⁢

1

⁢

∅

jj

,

wherein Ø rj represents E{y r1 (ω)Y j (ω)*}, Ø rr_1 represents E{y r1 (ω)y r1 (ω)*}, Ø jj represents E{Y j (ω)Y j (ω)*}, γ j represents a correlation value between the first right channel input audio and the j th channel of first sound source audio, y r1 (ω) represents the first right channel input audio, and Y j (ω) represents the j th channel of first sound source audio; and

determining a first sound source audio having a maximum correlation value among the M correlation values as the second target audio.

8. The method according to claim 1 , wherein the fusing the first left channel input audio with a first target sound source to obtain a first left channel audio specifically comprises:

obtaining a second left channel audio according to the formula y l1 ′(ω)=β 1 *y l1 (ω) +(1−β 1 )*Y t1 (ω), wherein

y l1 ′(ω) represents the first left channel audio, β 1 represents the first fusion coefficient, Y t1 (ω) represents the first target audio, and y l1 (ω) represents the first left channel input audio.

9. The method according to claim 1 , wherein the fusing the first right channel input audio with a second target sound source to obtain a first right channel audio specifically comprises:

obtaining the first right channel audio according to the formula y l1 ′(ω)=β 1 *y r1 (ω) +(1−β 1 )*Y t2 (ω), wherein

y r1 ′(ω) represents the first left channel audio, β 1 represents the first fusion coefficient, Y t2 (ω) represents the first target audio, and y r1 (ω) represents the first right channel input audio.

10. An electronic device, comprising a memory, a processor, and a touchscreen, wherein

the touchscreen is configured to display content;

the memory is configured to store a computer program, wherein the computer program comprises program instructions; and

the processor is configured to invoke the program instructions to enable the electronic device to perform:

starting, by the electronic device, a camera;

displaying a preview interface, wherein the preview interface comprises a first control;

detecting a first operation with respect to the first control;

starting photographing in response to the first operation;

displaying a photographing interface, wherein the photographing interface comprises a second control and the second control is used to adjust a zoom ratio;

displaying a first photographed image at a first moment when the zoom ratio is a first zoom ratio;

collecting, by a microphone, a first audio at the first moment;

detecting a third operation with respect to a third control; and

stopping photographing and saving a first video in response to the third operation; and

processing the first audio to obtain a first left channel output audio and a first right channel output audio, wherein

the processing the first audio to obtain a first left channel output audio and a first right channel output audio comprises:

performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio;

performing second processing on the first audio to obtain M channels of first sound source audios, wherein M represents a quantity of microphones of the electronic device;

fusing the first left channel input audio with a first target audio to obtain a first left channel audio, wherein the first target audio is a sound source audio having highest correlation with the first left channel input audio among the M channels of first sound source audios;

fusing the first right channel input audio with a second target audio to obtain a first right channel audio, wherein the second target audio is a sound source audio having highest correlation with the first right channel input audio among the M channels of first sound source audios; and

performing enhancement processing on the first left channel audio and the first right channel audio to obtain the first left channel output audio and the first right channel output audio.

11. The electronic device according to claim 10 , wherein the first photographed image comprises a first target object and a second target object, and the method further comprises:

detecting a second operation with respect to the second control;

adjusting the zoom ratio to be a second zoom ratio in response to the second operation, wherein the second zoom ratio is greater than the first zoom ratio;

displaying a second photographed image at a second moment, wherein the second photographed image comprises the first target object and does not comprise the second target object;

collecting, by the microphone, a second audio at the second moment, wherein the second audio comprises a first sound corresponding to the first target object and a second sound corresponding to the second target object; and

processing the second audio to obtain a second left channel output audio and a second right channel output audio, wherein the second left channel output audio and the second right channel output audio comprise a third sound and a fourth sound, the third sound corresponds to the first target object, the fourth sound corresponds to the second target object, the third sound is enhanced with respect to the first sound, and the fourth sound is suppressed with respect to the second sound.

12. The electronic device according to claim 11 , wherein the processing the second audio to obtain a second left channel output audio and a second right channel output audio comprises:

performing first processing on the second audio based on the second zoom ratio to obtain a second left channel input audio and a second right channel input audio;

performing second processing on the second audio to obtain M channels of second sound source audios, wherein M represents a quantity of microphones of the electronic device;

fusing the second left channel input audio with a third target audio to obtain a second left channel audio, wherein the third target audio is a sound source audio having highest correlation with the second left channel input audio among the M channels of second sound source audios;

fusing the second right channel input audio with a fourth target audio to obtain a second right channel audio, wherein the fourth target audio is a sound source audio having highest correlation with the second right channel input audio among the M channels of second sound source audios; and

performing enhancement processing on the second left channel audio and the second right channel audio to obtain the second left channel output audio and the second right channel output audio.

13. The electronic device according to claim 10 ,

wherein the performing second processing on the first audio to obtain M channels of first sound source audios specifically comprises:

obtaining the M channels of first sound source audios through calculation according to the formula Y(ω)=Σ i=1 M W i (ω)x i (ω), wherein

x i (ω) represents an audio signal of a first audio collected by the i th microphone in frequency domain, W i (ω) represents a first non-negative matrix corresponding to the i th microphone, Y(ω) represents a first matrix whose size is M*L, and each row vector of the first matrix is one channel of first sound source audio.

14. The electronic device according to claim 10 , wherein the performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio specifically comprises:

obtaining the first left channel audio according to the formula y l1 (ω)=α 1 *y 1 (ω)+(1−α 1 )*y 2 (ω); and

obtaining the first right channel audio according to the formula y r1 (ω)=α 1 *y 3 (ω) +(1−α 1 )*y 2 (ω), wherein

y l1 represents the first left channel input audio, y r1 (ω) represents the first right channel input audio, α 1 represents a fusion coefficient obtained based on the first zoom ratio, y 1 (ω) represents a first beam obtained based on the first audio and a first filter coefficient, y 2 (ω) represents a second beam obtained based on the first audio and a second filter coefficient, and y 3 (ω) represents a third beam obtained based on the first audio and a third filter coefficient.

15. The electronic device according to claim 10 , before the performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio, further comprising:

obtaining a first beam, a second beam, and a third beam respectively according to the formula y 1 (ω)=Σ i=1 M w 1i (ω)x i1 (ω), the formula y 2 (ω)=Σ i=1 M w 2i (ω)x i1 (ω), and the formula y 3 (ω)=Σ i=1 M w 3i (ω)x i1 (ω), wherein

y 1 (ω) represents the first beam, y 2 (ω) represents the second beam, y 3 (ω) represents the third beam, w 1i (ω) represents a first filter coefficient corresponding to the i th microphone in a first direction, w 2i (ω) represents a second filter coefficient corresponding to the i th microphone in a second direction, w 3i (ω) represents a third filter coefficient corresponding to the i th microphone in a third direction, x i1 (ω) represents the first audio collected by the i th microphone, the first direction is any direction within a range of 10° counterclockwise from the front to 90° counterclockwise from the front of the electronic device, the second direction is any direction within a range of 10° counterclockwise from the front to 10° clockwise from the front of the electronic device, and the third direction is any direction within a range of 10° clockwise from the front to 90° clockwise from the front of the electronic device.

16. The electronic device according to claim 10 , before the fusing the first left channel input audio with a first target sound source to obtain a first left channel audio, further comprising:

calculating a correlation value between the first left channel input audio and the M channels of first sound source audios according to the formula

γ

i

=

∅

li

∅

ll

⁢

_

⁢

1

⁢

∅

ii

,

Ø li represents E{y l1 (ω)Y i (ω)*}, Ø ll_1 represents E{y l1 (ω)y l1 (ω)*}, Ø ii represents E{Y 1 (ω)Y 1 (ω)*}, γ 1 represents a correlation value between the first left channel input audio and the i th channel of first sound source audio, y l1 (ω) represents the first left channel input audio, and Y i (ω) represents the i th channel of first sound source audio;

if there is only one maximum correlation value among the M correlation values, determining a first sound source audio having the maximum correlation value as the first target audio; and

if there are a plurality of maximum correlation values among the M correlation values, calculating an average value of first sound source audios corresponding to the plurality of maximum correlation values to obtain the first target audio.

17. The electronic device according to claim 10 , before the fusing the first right channel input audio with a second target sound source to obtain a first right channel audio, further comprising:

calculating a correlation value between the first right channel input audio and the M channels of first sound source audios according to the formula

γ

j

=

∅

rj

∅

rr

⁢

_

⁢

1

⁢

∅

jj

,

wherein Ø rj represents {y r1 (ω)Y j (ω)*}, Ø rr_1 represents E{y r1 (ω)y r1 (ω)*}, Ø jj represents E{Y j (ω)Y j (ω)*}, γ j represents a correlation value between the first right channel input audio and the j th channel of first sound source audio, y r1 (ω) represents the first right channel input audio, and Y j (ω) represents the j th channel of first sound source audio; and

determining a first sound source audio having a maximum correlation value among the M correlation values as the second target audio.

18. The electronic device according to claim 10 , wherein the fusing the first left channel input audio with a first target sound source to obtain a first left channel audio specifically comprises:

obtaining a second left channel audio according to the formula y l1 ′(ω)=β 1 *y l1 (ω) +(1−β 1 )*Y t1 (ω), wherein

y l1 ′(ω) represents the first left channel audio, β 1 represents the first fusion coefficient, Y t1 (ω) represents the first target audio y l1 (ω) represents the first left channel input audio.

19. The electronic device according to claim 10 , wherein the fusing the first right channel input audio with a second target sound source to obtain a first right channel audio specifically comprises:

obtaining the first right channel audio according to the formula y r1 ′(ω)=β 1 *y r1 (ω)+(1−β 1 )*Y t2 (ω), wherein

y r1 ′(ω) represents the first left channel audio, β 1 represents the first fusion coefficient, Y t2 (ω) represents the first target audio, and y r1 (ω) represents the first right channel input audio.

20. A non-transitory machine readable storage medium, wherein the non-transitory machine readable storage medium stores a computer program, and when the computer program is executed by a processor, an electronic device is caused to execute:

starting a camera;

displaying a preview interface, wherein the preview interface comprises a first control;

detecting a first operation with respect to the first control;

starting photographing in response to the first operation;

displaying a photographing interface, wherein the photographing interface comprises a second control and the second control is used to adjust a zoom ratio;

displaying a first photographed image at a first moment when the zoom ratio is a first zoom ratio;

collecting, by a microphone, a first audio at the first moment;

detecting a third operation with respect to a third control; and

stopping photographing and saving a first video in response to the third operation; and

processing the first audio to obtain a first left channel output audio and a first right channel output audio, wherein

the processing the first audio to obtain a first left channel output audio and a first right channel output audio comprises:

performing first processing on the first audio based on the first zoom ratio to obtain a first left channel input audio and a first right channel input audio;

performing second processing on the first audio to obtain M channels of first sound source audios, wherein M represents a quantity of microphones of the electronic device;

fusing the first left channel input audio with a first target audio to obtain a first left channel audio, wherein the first target audio is a sound source audio having highest correlation with the first left channel input audio among the M channels of first sound source audios;

fusing the first right channel input audio with a second target audio to obtain a first right channel audio, wherein the second target audio is a sound source audio having highest correlation with the first right channel input audio among the M channels of first sound source audios; and

performing enhancement processing on the first left channel audio and the first right channel audio to obtain the first left channel output audio and the first right channel output audio.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: HONOR DEVICE CO., LTD.
To: BEIJING HONOR DEVICE CO., LTD.
Reel/Frame 064483/0425 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: LIU, ZHENYI; XUAN, JIANYONG
To: HONOR DEVICE CO., LTD.
Reel/Frame 064455/0369 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: GAO, HAIKUAN
To: HONOR DEVICE CO., LTD.
Reel/Frame 064455/0608 →
Priority Claims (2)
CN 202111161876.0 · Sep 30, 2021 · national
CN 202111593768.0 · Dec 23, 2021 · national
Continuity (1)
Related Publication 20240305890A1 · Sep 12, 2024
References Cited (39)
US 8218033B2 · Oku · 2012 [cited by examiner]
US 8401364B2 · Oku · 2013 [cited by examiner]
US 8712231B2 · Ohtsuka · 2014 [cited by examiner]
US 9258644B2 · Maenpaa et al. · 2016 [cited by applicant]
US 10998870B2 · Aoyama · 2021 [cited by examiner]
US 11277686B2 · Kim et al. · 2022 [cited by applicant]
US 11425497B2 · Salehin · 2022 [cited by examiner]
US 11671752B2 · Kim · 2023 [cited by examiner]
US 20090066798A1 · Oku · 2009 [cited by examiner]
US 20110052139A1 · Oku · 2011 [cited by examiner]
US 20110085061A1 · Kim · 2011 [cited by examiner]
US 20120308220A1 · Ohtsuka · 2012 [cited by applicant]
US 20130021502A1 · Oku · 2013 [cited by examiner]
US 20140029761A1 · Maenpaa · 2014 [cited by examiner]
US 20200322540A1 · Tsujimoto · 2020 [cited by examiner]
US 20200329202A1 · Toriumi · 2020 [cited by examiner]
US 20200358415A1 · Aoyama et al. · 2020 [cited by applicant]
US 20210044896A1 · Kim et al. · 2021 [cited by applicant]
US 20220201395A1 · Salehin · 2022 [cited by examiner]
US 20220360891A1 · Kim · 2022 [cited by examiner]
US 20230018004A1 · Wen et al. · 2023 [cited by applicant]
US 20230041730A1 · Hu et al. · 2023 [cited by applicant]
US 20230115929A1 · Bian et al. · 2023 [cited by applicant]
US 20230116044A1 · Han et al. · 2023 [cited by applicant]
CN 107105183A · 2017 [cited by applicant]
CN 110827843A · 2020 [cited by applicant]
CN 112492380A · 2021 [cited by applicant]
CN 112599144A · 2021 [cited by applicant]
CN 112951257A · 2021 [cited by applicant]
CN 113365012A · 2021 [cited by applicant]
CN 113365013A · 2021 [cited by applicant]
CN 113452898A · 2021 [cited by applicant]
CN 113497882A · 2021 [cited by applicant]
CN 114363512A · 2022 [cited by applicant]
EP 2690886A1 · 2014 [cited by applicant]
KR 20210017229A · 2021 [cited by applicant]
WO 2021175165A1 · 2021 [cited by applicant]
WO 2021175197A1 · 2021 [cited by applicant]
Arun Asokan Nair et al: “Audiovisual Zooming: What You See Is What You Hear”, In Proceedings of the 27th ACM International Conference on Multimedia (MM '19), Association for Computing Machinery, New York, NY, USA, 1107-… [cited by applicant]
Cited By (1)
US 12,464,304