IP Library Granted Patent US 12670639
Granted Patent B2
US 12670639 · App. 18/407,309 · Granted Jun 30, 2026

Selective amplification of voice and interactive language simulator

Inventors: Shiraz Akmal (Playa Vista, CA); Aaron M. Burns (Sunnyvale, CA); Brad K. Herman (Culver City, CA)
Assignee: Apple Inc.
G06T11/60G10L15/005G10L15/187G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670639
App. No.
18/407,309
Granted
Jun 30, 2026
Kind
B2
Abstract

Systems and processes for operating a digital assistant are provided. An example method includes, at an electronic device having one or more processors and memory, receiving an audio input including an utterance, determining, based on a speaker profile, an identity of a speaker of the utterance, determining whether the identity of the speaker matches a predetermined identity, and in accordance with a determination that the identity of the speaker matches the predetermined identity selectively adjusting a volume of the utterance relative to a volume of other sound of the audio input and providing an output of the adjusted utterance.

Claims (115)

1 . A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors of an electronic device, the one or more programs including instructions for:

receiving an audio input including an utterance;

determining, based on a speaker profile, an identity of a speaker of the utterance;

determining whether the identity of the speaker matches a predetermined identity; and

in accordance with a determination that the identity of the speaker matches the predetermined identity:

selectively adjusting a volume of the utterance relative to a volume of other sound of the audio input based on a personal setting associated with the identity of the speaker, wherein the personal setting associated with the identity of the speaker is provided by a user of the electronic device other than the speaker and wherein the speaker of the utterance is different from the user of the electronic device; and

providing an output of the selectively adjusted audio input to the user of the electronic device.

2 . The non-transitory computer-readable storage medium of claim 1 , wherein determining, based on the speaker profile, the identity of the speaker of the utterance further comprises:

selecting a voiceprint of the speaker profile;

comparing a voiceprint derived from the utterance to the selected voiceprint of the speaker profile; and

determining whether the speaker of the utterance matches an identity associated with the speaker profile based on the comparison of the voiceprints.

3 . The non-transitory computer-readable storage medium of claim 2 , wherein the voiceprint selected from the speaker profile was derived from a previous utterance received from the speaker.

4 . The non-transitory computer-readable storage medium of claim 2 , wherein the speaker profile is received from a second electronic device prior to receipt of the audio input.

5 . The non-transitory computer-readable storage medium of claim 2 , wherein the speaker profile is received from a third electronic device associated with the speaker prior to receipt of the audio input.

6 . The non-transitory computer-readable storage medium of claim 2 , wherein the voiceprint of the speaker profile is a first voiceprint of a plurality of voiceprints stored in associated with the speaker profile.

7 . The non-transitory computer-readable storage medium of claim 2 , the one or more programs further including instructions for:

in accordance with a determination that the speaker of the utterance matches the speaker associated with the speaker profile, adding the voiceprint derived from the utterance to the speaker profile associated with the speaker.

8 . The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the identity of the speaker is included in a set of identities associated with amplification.

9 . The non-transitory computer-readable storage medium of claim 1 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the speaker profile associated with the speaker includes an amplification property.

10 . The non-transitory computer-readable storage medium of claim 1 , the one or more programs further including instructions for:

determining whether to selectively adjust the audio input based on the determined identity, whether the electronic device is providing a virtual reality output, and one or more speech characteristics of the audio input.

11 . The non-transitory computer-readable storage medium of claim 1 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

determining a current volume of the utterance;

determining an output volume of the utterance; and

adjusting one or more characteristics of the utterance to increase the current volume to the output volume.

12 . The non-transitory computer-readable storage medium of claim 1 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the utterance.

13 . The non-transitory computer-readable storage medium of claim 1 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the audio input.

14 . The non-transitory computer-readable storage medium of claim 1 , wherein the output of the adjusted utterance is provided at an accessory device communicatively coupled to the electronic device.

15 . The non-transitory computer-readable storage medium of claim 1 , wherein the utterance is a first utterance, the speaker profile is a first speaker profile, the speaker is a first speaker, the identity is a first identity, and wherein the received audio input further includes a second utterance, the one or more programs further including instructions for:

determining, based on a second speaker profile, a second identity of a second speaker of the second utterance;

determining whether the second identity of the second speaker matches a second predetermined identity; and

in accordance with a determination that the second identity of the second speaker matches the second predetermined identity:

selectively adjusting a volume of the second utterance relative to a volume of other sound of the audio input.

16 . The non-transitory computer-readable storage medium of claim 15 , wherein the first utterance and the second utterance are received concurrently from different speakers.

17 . The non-transitory computer-readable storage medium of claim 1 , the one or more programs further including instructions for:

increasing a volume of the audio input; and

maintaining a volume of an audio output associated with a provided virtual reality.

18 . The non-transitory computer-readable storage medium of claim 10 wherein the determination of whether to selectively adjust the audio input is further based on an application of the electronic device.

19 . The non-transitory computer-readable storage medium of claim 1 wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input is performed in accordance with a determination that the wearable electronic device is providing a virtual reality output when the utterance is received.

20 . An electronic device comprising:

one or more processors;

a memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

receiving an audio input including an utterance;

determining, based on a speaker profile, an identity of a speaker of the utterance;

determining whether the identity of the speaker matches a predetermined identity; and

in accordance with a determination that the identity of the speaker matches the predetermined identity:

selectively adjusting a volume of the utterance relative to a volume of other sound of the audio input based on a personal setting associated with the identity of the speaker, wherein the personal setting associated with the identity of the speaker is provided by a user of the electronic device other than the speaker and wherein the speaker of the utterance is different from the user of the electronic device; and

providing an output of the selectively adjusted audio input to the user of the electronic device.

21 . The electronic device of claim 20 , wherein determining, based on the speaker profile, the identity of the speaker of the utterance further comprises:

selecting a voiceprint of the speaker profile;

comparing a voiceprint derived from the utterance to the selected voiceprint of the speaker profile; and

determining whether the speaker of the utterance matches an identity associated with the speaker profile based on the comparison of the voiceprints.

22 . The electronic device of claim 21 , wherein the voiceprint selected from the speaker profile was derived from a previous utterance received from the speaker.

23 . The electronic device of claim 21 , wherein the speaker profile is received from a second electronic device prior to receipt of the audio input.

24 . The electronic device of claim 20 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the identity of the speaker is included in a set of identities associated with amplification.

25 . The electronic device of claim 20 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the speaker profile associated with the speaker includes an amplification property.

26 . The electronic device of claim 20 , the one or more programs further including instructions for:

determining whether to selectively adjust the audio input based on the determined identity, whether the electronic device is providing a virtual reality output, and one or more speech characteristics of the audio input.

27 . The electronic device of claim 20 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

determining a current volume of the utterance;

determining an output volume of the utterance; and

adjusting one or more characteristics of the utterance to increase the current volume to the output volume.

28 . The electronic device of claim 20 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the utterance.

29 . The electronic device of claim 20 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the audio input.

30 . The electronic device of claim 20 , wherein the output of the adjusted utterance is provided at an accessory device communicatively coupled to the electronic device.

31 . The electronic device of claim 20 , wherein the utterance is a first utterance, the speaker profile is a first speaker profile, the speaker is a first speaker, the identity is a first identity, and wherein the received audio input further includes a second utterance, the one or more programs further including instructions for:

determining, based on a second speaker profile, a second identity of a second speaker of the second utterance;

determining whether the second identity of the second speaker matches a second predetermined identity; and

in accordance with a determination that the second identity of the second speaker matches the second predetermined identity:

selectively adjusting a volume of the second utterance relative to a volume of other sound of the audio input.

32 . The electronic device of claim 20 , the one or more programs further including instructions for:

increasing a volume of the audio input; and

maintaining a volume of an audio output associated with a provided virtual reality.

33 . The electronic device of claim 20 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input is performed in accordance with a determination that the wearable electronic device is providing a virtual reality output when the utterance is received.

34 . A method, comprising:

at an electronic device with one or more processors and memory:

receiving an audio input including an utterance;

determining, based on a speaker profile, an identity of a speaker of the utterance;

determining whether the identity of the speaker matches a predetermined identity; and

in accordance with a determination that the identity of the speaker matches the predetermined identity:

selectively adjusting a volume of the utterance relative to a volume of other sound of the audio input based on a personal setting associated with the identity of the speaker, wherein the personal setting associated with the identity of the speaker is provided by a user of the electronic device other than the speaker and wherein the speaker of the utterance is different from the user of the electronic device; and

providing an output of the selectively adjusted audio input to the user of the electronic device.

35 . The method of claim 34 , wherein determining, based on the speaker profile, the identity of the speaker of the utterance further comprises:

selecting a voiceprint of the speaker profile;

comparing a voiceprint derived from the utterance to the selected voiceprint of the speaker profile; and

determining whether the speaker of the utterance matches an identity associated with the speaker profile based on the comparison of the voiceprints.

36 . The method of claim 35 , wherein the voiceprint selected from the speaker profile was derived from a previous utterance received from the speaker.

37 . The method of claim 35 , wherein the speaker profile is received from a second electronic device prior to receipt of the audio input.

38 . The method of claim 34 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the identity of the speaker is included in a set of identities associated with amplification.

39 . The method of claim 34 , wherein determining whether the identity of the speaker matches the predetermined identity further comprises determining whether the speaker profile associated with the speaker includes an amplification property.

40 . The method of claim 34 , further comprising:

determining whether to selectively adjust the audio input based on the determined identity, whether the electronic device is providing a virtual reality output, and one or more speech characteristics of the audio input.

41 . The method of claim 34 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

determining a current volume of the utterance;

determining an output volume of the utterance; and

adjusting one or more characteristics of the utterance to increase the current volume to the output volume.

42 . The method of claim 34 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the utterance.

43 . The method of claim 34 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input further comprises:

decreasing the volume of the audio input.

44 . The method of claim 34 , wherein the output of the adjusted utterance is provided at an accessory device communicatively coupled to the electronic device.

45 . The method of claim 20 , wherein the utterance is a first utterance, the speaker profile is a first speaker profile, the speaker is a first speaker, the identity is a first identity, and wherein the received audio input further includes a second utterance, the method further comprising:

determining, based on a second speaker profile, a second identity of a second speaker of the second utterance;

determining whether the second identity of the second speaker matches a second predetermined identity; and

in accordance with a determination that the second identity of the second speaker matches the second predetermined identity:

selectively adjusting a volume of the second utterance relative to a volume of other sound of the audio input.

46 . The method of claim 20 , further comprising:

increasing a volume of the audio input; and

maintaining a volume of an audio output associated with a provided virtual reality.

47 . The method of claim 34 , wherein selectively adjusting the volume of the utterance relative to the volume of other sound of the audio input is performed in accordance with a determination that the wearable electronic device is providing a virtual reality output when the utterance is received.