Voice masking method, apparatus, and system, and vehicle
A voice masking method, apparatus, and system are provided, which are applicable to the field of intelligent vehicles. The method includes: determining a voice source location and a target masking location; receiving a sound signal from the voice source location, and detecting whether a voice signal exists in the sound signal, to generate a detection result; and if the detection result is that the voice signal exists in the sound signal, generating a masking sound for the voice signal, and outputting the masking sound to a first speaker at the voice source location and a second speaker at the target masking location; or if the detection result is that no voice signal exists in the sound signal, skipping generating the masking sound. According to this application, a requirement of an occupant for private voice communication in vehicle cockpit space can be met, and information security can be improved.
1 . A voice masking method performed by a vehicle, comprising:
determining a voice source location and a target masking location;
receiving a sound signal from the voice source location;
detecting whether a voice signal exists in the sound signal to generate a detection result; and
if the detection result is that the voice signal exists in the sound signal:
generating a masking sound for the voice signal; and
outputting the masking sound to a second speaker at the target masking location;
wherein the generating a masking sound for the voice signal comprises: generating the masking sound by performing time domain inversion processing on the voice signal;
or
wherein the generating a masking sound for the voice signal comprises:
analyzing a voice feature of the voice signal, and obtaining the noise data, from the noise database, corresponding to the voice feature;
generating the masking sound based on the noise data, or the voice signal and the noise data.
2 . The method according to claim 1 , further comprising:
before the determining the voice source location and the target masking location, receiving a voice masking activation instruction; and
wherein the determining the voice source location and the target masking location comprises:
determining the voice source location based on a source location of the activation instruction; and
determining the target masking location based on occupant information obtained by a sensor in the vehicle.
3 . The method according to claim 2 , further comprising:
receiving a voice masking deactivation instruction; and
stopping, based on the deactivation instruction, receiving the sound signal at the voice source location.
4 . The method according to claim 1 , wherein the voice source location and the target masking location are input by an occupant in the vehicle.
5 . The method according to claim 1 , wherein after the receiving the sound signal from the voice source location, the method further comprises:
performing enhancement processing on the sound signal.
6 . The method according to claim 5 , wherein the enhancement processing comprises at least one of echo cancellation processing or adaptive voice noise reduction processing.
7 . The method according to claim 1 , further comprising:
generating a second masking sound and outputting the second masking sound to the second speaker at the target masking location;
or
generating a first masking sound and outputting the first masking sound to a first speaker at the voice source location;
or
generating a first masking sound and a second masking sound, and outputting the first masking sound to a first speaker at the voice source location and outputting the second masking sound to the second speaker at the target masking location, wherein a volume of the first masking sound is lower than the second masking sound;
or
generating a first masking sound and a second masking sound, and outputting the first masking sound to a first speaker at the voice source location and outputting the second masking sound to the second speaker at the target masking location, wherein the masking sound played by the first speaker at the voice source location is used to cancel the masking sound received at a target source location from the second speaker.
8 . The method according to claim 7 , wherein the outputting the first masking sound to the first speaker at the voice source location and the second masking sound to the second speaker at the target masking location comprises:
performing sound field control processing on the first masking sound and the second masking sound based on the voice source location and the target masking location to output the first masking sound and the second masking sound, wherein the sound field control processing comprises adjusting a phase or an amplitude of each frequency signal in the first masking sound and the second masking sound; and
outputting the first masking sound to the first speaker at the voice source location, and outputting the second masking sound to the second speaker at the target masking location.
9 . The method according to claim 8 , wherein the performing sound field control processing comprises:
performing sound field control processing on the first masking sound and the second masking sound based on the voice source location and the target masking location to output masking sounds of N channels, wherein the masking sounds of the N channels comprise the first masking sound and the second masking sound, and N is a quantity of speakers in a cockpit, wherein
a volume of a masking sound received by the first speaker is lower than a volume of a masking sound received by the second speaker.
10 . The method according to claim 8 , wherein when the target masking location comprises a driver seat, the second speaker comprises a speaker located at a headrest of the driver seat.
11 . The method according to claim 1 , wherein before the outputting the masking sound to the second speaker at the target masking location, the method further comprises: performing automatic gain control adjustment on the masking sound, so that a volume of the masking sound falls within a specified range.
12 . The method according to claim 1 , wherein when the target masking location comprises a driver seat, the method further comprises: increasing a volume of a vehicle safety alarm sound.
13 . The method according to claim 12 , further comprising:
obtaining an outside-vehicle sound, performing specific type identification on the outside-vehicle sound to identify a specific sound in the outside-vehicle sound; and
outputting the specific sound to a speaker in the driver seat.
14 . A voice masking apparatus, comprising at least one processor and a memory, wherein the memory stores instructions for execution by the at least one processor to perform operations comprising:
determining a voice source location and a target masking location;
receiving a sound signal from the voice source location;
detecting whether a voice signal exists in the sound signal to generate a detection result; and
if the detection result is that the voice signal exists in the sound signal;
generating a masking sound for the voice signal; and
outputting the masking sound to a second speaker at the target masking location;
wherein the generating a masking sound for the voice signal comprises: generating the masking sound by performing time domain inversion processing on the voice signal;
or
wherein the generating a masking sound for the voice signal comprises:
analyzing a voice feature of the voice signal, and obtaining the noise data, from the noise database, corresponding to the voice feature;
generating the masking sound based on the noise data, or the voice signal and the noise data.
15 . The voice masking apparatus according to claim 14 , the operations further comprising:
before the determining the voice source location and the target masking location, receiving a voice masking activation instruction; and
wherein the determining the voice source location and the target masking location comprises:
determining the voice source location based on a source location of the activation instruction; and
determining the target masking location based on occupant information obtained by a sensor in a vehicle.
16 . The voice masking apparatus according to claim 14 , wherein the voice source location and the target masking location are input by an occupant in a vehicle.
17 . The voice masking apparatus according to claim 14 , wherein after the receiving the sound signal from the voice source location, the operations further comprising:
performing enhancement processing on the sound signal.
18 . A non-transitory computer-readable storage medium, comprising instructions, wherein when the instructions are executed on a computer, the computer is enabled to perform operations comprising:
determining a voice source location and a target masking location;
receiving a sound signal from the voice source location;
detecting whether a voice signal exists in the sound signal to generate a detection result; and
if the detection result is that the voice signal exists in the sound signal:
generating a masking sound for the voice signal; and
outputting the masking sound to a second speaker at the target masking location;
wherein the generating a masking sound for the voice signal comprises: generating the masking sound by performing time domain inversion processing on the voice signal;
or
wherein the generating a masking sound for the voice signal comprises:
analyzing a voice feature of the voice signal, and obtaining the noise data, from the noise database, corresponding to the voice feature;
generating the masking sound based on the noise data, or the voice signal and the noise data.
19 . A voice masking apparatus, comprising at least one processor and a memory, wherein the memory stores instructions for execution by the at least one processor to perform operations comprising:
determining a voice source location and a target masking location;
receiving a sound signal from the voice source location;
detecting whether a voice signal exists in the sound signal to generate a detection result; and
if the detection result is that the voice signal exists in the sound signal;
generating a masking sound for the voice signal;
outputting the masking sound to a second speaker at the target masking location; and
increasing a volume of a vehicle safety alarm sound.
20 . A voice masking apparatus, comprising at least one processor and a memory, wherein the memory stores instructions for execution by the at least one processor to perform operations comprising:
determining a voice source location and a target masking location;
receiving a sound signal from the voice source location;
detecting whether a voice signal exists in the sound signal to generate a detection result; and
if the detection result is that the voice signal exists in the sound signal:
generating a masking sound for the voice signal; and
outputting the masking sound to a second speaker at the target masking location;
obtaining an outside-vehicle sound;
performing specific type identification on the outside-vehicle sound to identify a specific sound in the outside-vehicle sound; and
outputting the specific sound to the speaker in a driver seat.