Source speech modification based on an input speech characteristic mapped to multiple reference embeddings within a threshold similarity
A device includes one or more processors configured to process an input audio spectrum of input speech to detect a first characteristic associated with the input speech. The one or more processors are also configured to select, based at least in part on the first characteristic, one or more reference embeddings from among multiple reference embeddings. The one or more processors are further configured to process a representation of source speech, using the one or more reference embeddings, to generate an output audio spectrum of output speech.
1 . A device comprising:
a memory configured to store data; and
one or more processors coupled to the memory and configured to:
process an input audio spectrum of input speech to detect an input characteristic associated with the input speech;
determine a target characteristic that corresponds to the input characteristic;
based on a determination that the target characteristic fails to correspond to any reference embedding of multiple reference embeddings, select a plurality of reference embeddings from among the multiple reference embeddings, wherein the plurality of reference embeddings is selected from the multiple reference embeddings based on a determination that the target characteristic is within a threshold similarity of respective characteristics corresponding to the plurality of reference embeddings; and
process a representation of source speech based on the plurality of reference embeddings to generate an output audio spectrum of output speech.
2 . The device of claim 1 , wherein the one or more processors are further configured to:
process the input audio spectrum to detect a first emotion;
process image data to detect a second emotion; and
determine the target characteristic based on the first emotion and the second emotion.
3 . The device of claim 2 , wherein the one or more processors are further configured to perform face detection on the image data, and wherein the second emotion is detected at least partially based on an output of the face detection.
4 . The device of claim 2 , wherein the one or more processors are further configured to receive audio data from one or more microphones concurrently with receiving the image data from one or more image sensors, and wherein the audio data represents the input speech, the source speech, or both.
5 . The device of claim 4 , further comprising the one or more microphones and the one or more image sensors.
6 . The device of claim 1 , wherein the representation of the source speech includes encoded source speech, and wherein the one or more processors are further configured to:
generate a conversion embedding based on the plurality of reference embeddings;
apply the conversion embedding to the encoded source speech to generate converted encoded source speech; and
decode the converted encoded source speech to generate the output audio spectrum.
7 . The device of claim 6 , wherein the one or more processors are configured to combine the plurality of reference embeddings and a baseline embedding to generate the conversion embedding.
8 . The device of claim 6 , wherein the one or more processors are configured to combine the plurality of the reference embeddings to generate the conversion embedding.
9 . The device of claim 1 , wherein the one or more processors are configured to map the input characteristic to the target characteristic according to an operation mode.
10 . The device of claim 9 , wherein the operation mode is based on a user input, a configuration setting, default data, or a combination thereof.
11 . The device of claim 1 , wherein the one or more processors are configured to:
obtain a representation of the input speech;
process the representation of the input speech to generate the input audio spectrum; and
generate a representation of the output speech based on the output audio spectrum.
12 . The device of claim 11 , wherein the representation of the input speech includes first text, and wherein the representation of the output speech includes second text.
13 . The device of claim 1 , wherein:
the input characteristic includes an emotion of the input speech,
the target characteristic includes a target emotion,
and
a first reference embedding of the plurality of reference embeddings is selected based on a determination that the first reference embedding represents a particular emotion that is within the threshold similarity of the target emotion.
14 . The device of claim 1 , wherein:
the input characteristic includes a volume of the input speech,
the target characteristic includes a target volume,
and
a first reference embedding of the plurality of reference embeddings is selected based on a determination that the first reference embedding represents a particular volume that is within the threshold similarity of the target volume.
15 . The device of claim 1 , wherein:
the target characteristic includes a target pitch,
and
a first reference embedding of the plurality of reference embeddings is selected based on a determination that the first reference embedding represents a particular pitch that is within the threshold similarity of the target pitch.
16 . The device of claim 1 , wherein:
the input characteristic includes a speed of the input speech,
the target characteristic includes a target speed,
and
a first reference embedding of the plurality of reference embeddings is selected based on a determination that the first reference embedding represents a particular speed that is within the threshold similarity of the target speed.
17 . The device of claim 1 , wherein the one or more processors are further configured to determine a second target characteristic that corresponds to a second characteristic associated with the input speech, wherein a second reference embedding of the plurality of reference embeddings is selected based on a determination that the second reference embedding represents the second target characteristic.
18 . The device of claim 1 , wherein
a first reference embedding of the plurality of reference embeddings is a data structure that stores speech feature values that are indicative of a first characteristic that is within the threshold similarity of the target characteristic, and
the input speech is identical to the source speech.
19 . The device of claim 1 , wherein the one or more processors are further configured to:
receive an input speech representation of the input speech via a first source;
process the input speech representation to generate the input audio spectrum; and
receive the representation of the source speech via a second source that is different from the first source, wherein:
the second source includes a virtual assistant software application, and
the output speech corresponds to a social interaction response from the virtual assistant software application that is based on the target characteristic.
20 . The device of claim 1 , wherein a first speech characteristic of the output speech matches a second speech characteristic of the input speech.
21 . The device of claim 1 , wherein the representation of the source speech is based on at least one of source speech audio, source speech text, a source speech spectrum, linear predictive coding (LPC) coefficients, or mel-frequency cepstral coefficients (MFCCs), and wherein the one or more processors are further configured to:
process, using an encoder, a source audio spectrum of the source speech to generate a source speech embedding; and
process, using a fundamental frequency (F0) extractor, the source audio spectrum to generate a F0 embedding, wherein the representation of the source speech is based on the source speech embedding and the F0 embedding.
22 . The device of claim 1 , wherein the one or more processors are integrated into at least one of a vehicle, a communication device, a gaming device, an extended reality (XR) device, or a computing device.
23 . The device of claim 1 , wherein the one or more processors are further configured to:
detect a second characteristic and a third characteristic associated with the input speech; and
determine a second target characteristic that corresponds to the second characteristic and a third target characteristic that corresponds to the third characteristic
wherein a second reference embedding and a third reference embedding of one or more reference embeddings are selected from the multiple reference embeddings based on a determination that the second reference embedding represents the second target characteristic and that the third reference embedding represents the third target characteristic.
24 . A method comprising:
processing, at a device, an input audio spectrum of input speech to detect an input characteristic associated with the input speech;
determining, at the device, a target characteristic that corresponds to the input characteristic;
based on determining that the target characteristic fails to correspond to any reference embedding of multiple reference embeddings, selecting a plurality of reference embeddings from among the multiple reference embeddings, wherein the plurality of reference embeddings is selected from the multiple reference embeddings based on determining that the target characteristic is within a threshold similarity of respective characteristics corresponding to the plurality of reference embeddings; and
processing a representation of source speech based on the plurality of reference embeddings to generate an output audio spectrum of output speech.
25 . The method of claim 24 , further comprising:
generating, at the device, a conversion embedding based on the plurality of reference embeddings;
applying the conversion embedding to encoded source speech to generate converted encoded source speech, wherein the representation of the source speech includes encoded source speech; and
decoding, at the device, the converted encoded source speech to generate the output audio spectrum.
26 . The method of claim 25 , further comprising combining, at the device, the plurality of reference embeddings and a baseline embedding to generate the conversion embedding.
27 . The method of claim 24 , further comprising:
processing, using an encoder, a source audio spectrum of the source speech to generate a source speech embedding; and
processing, using a fundamental frequency (F0) extractor, the source audio spectrum to generate a F0 embedding, wherein the representation of the source speech is based on the source speech embedding and the F0 embedding.
28 . The method of claim 24 , further comprising:
receiving, at the device, an input speech representation of the input speech via a first source;
processing, at the device, the input speech representation to generate the input audio spectrum; and
receiving, at the device, the representation of the source speech via a second source that is different from the first source, wherein the second source includes a virtual assistant software application, and wherein the output speech corresponds to a social interaction response from the virtual assistant software application that is based on the target characteristic.
29 . A non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to:
process an input audio spectrum of input speech to detect an input characteristic associated with the input speech;
determine a target characteristic that corresponds to the input characteristic;
based on a determination that the target characteristic fails to correspond to any reference embedding of multiple reference embeddings, select a plurality of reference embeddings from among the multiple reference embeddings, wherein the plurality of reference embeddings is selected from the multiple reference embeddings based on a determination that the target characteristic is within a threshold similarity of respective characteristics corresponding to the plurality of reference embeddings; and
process a representation of source speech, based on the plurality of reference embeddings, to generate an output audio spectrum of output speech.
30 . An apparatus comprising:
means for processing an input audio spectrum of input speech to detect an input characteristic associated with the input speech;
means for determining a target characteristic that corresponds to the input characteristic;
means for selecting a plurality of reference embeddings from among multiple reference embeddings based on determining that the target characteristic fails to correspond to any reference embedding of the multiple reference embeddings, wherein the plurality of reference embeddings is selected from the multiple reference embeddings based on a determination that the target characteristic is within a threshold similarity of respective characteristics corresponding to the plurality of reference embeddings; and
means for processing a representation of source speech based on the plurality of reference embeddings to generate an output audio spectrum of output speech.