IP Library › Granted Patent US 12,380,912
Granted Patent B2
US 12,380,912 · App. 17/689,546 · Granted Aug 5, 2025

Audio automatic mixer with frequency weighting

Inventor: Asbjorn Therkelsen (Nesbru, NO)
Assignee: CISCO TECHNOLOGY, INC.
G10L21/0308G10L25/18H04R3/005H04R5/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,912
App. No.
17/689,546
Granted
Aug 5, 2025
Kind
B2
Abstract

A method is provided that is performed at a system including multiple speech collectors to collect speech from a talker to produce corresponding ones of multiple audio signals that each convey speech energy: for each audio signal: separating high-frequency speech energy from low-frequency speech energy; and determining a first energy level of the high-frequency speech energy; and determining a preferred audio signal among the multiple audio signals for subsequent processing at least based on the first energy level of each audio signal.

Claims (60)

1. A method comprising:

at a controller of a system that includes speech collectors to convert speech from a talker to corresponding ones of audio signals that each convey speech energy:

performing signal processing on each audio signal by:

filtering the speech energy into high-frequency speech energy and low-frequency speech energy;

determining a high-frequency energy level of the high-frequency speech energy;

determining a low-frequency energy level of the low-frequency speech energy;

normalizing the high-frequency energy level with respect to a first sum of the high-frequency speech energy across the audio signals, to produce a normalized high-frequency energy level;

normalizing the low-frequency energy level with respect to a second sum of the low-frequency speech energy across the audio signals, to produce a normalized low-frequency energy level;

weighting each of the normalized high-frequency energy level and the normalized low-frequency energy level to produce a weighted normalized high-frequency energy level and a weighted normalized low-frequency energy level; and

summing the weighted normalized high-frequency energy level and the weighted normalized low-frequency energy level to produce an individual energy level; and

selecting as a preferred audio signal whichever of the audio signals has a highest individual energy level for encoding into a data packet.

2. The method of claim 1 , further comprising:

selecting a preferred speech collector among the speech collectors that produced the preferred audio signal.

3. The method of claim 1 , wherein each speech collector is one of (i) a microphone beam formed using a microphone array, and (ii) a microphone that is not part of the microphone array.

4. The method of claim 3 , wherein the speech collectors include one or more microphone beams formed using the microphone array and one or more microphones that are not part of the microphone array.

5. The method of claim 1 , wherein:

the speech energy spans a frequency band of 0 kHz to 20 kHz.

6. The method of claim 1 , wherein filtering includes filtering such that the high-frequency speech energy contains frequencies for voiced hard consonants and the low-frequency speech energy does not contain the frequencies for the voiced hard consonants.

7. The method of claim 1 , wherein filtering includes filtering the high-frequency speech energy into an audio frequency band from 2 kHz to 8 kHz.

8. The method of claim 1 , wherein:

the weighted normalized high-frequency energy level is weighted more heavily than the weighted normalized low-frequency energy level.

9. The method of claim 1 , wherein:

filtering includes filtering such that the high-frequency speech energy contains voiced hard sounds and the low-frequency speech energy contains voiced soft sounds.

10. The method of claim 1 , wherein:

filtering includes using a high-frequency bandpass filter to produce the high-frequency speech energy and a low-frequency bandpass filter to produce the low-frequency speech energy.

11. The method of claim 1 , further comprising:

receiving an indicator that indicates a presence of the speech; and

performing the signal processing of each audio signal while the indicator indicates the presence of the speech.

12. An apparatus comprising:

speech collectors to convert speech from a talker to corresponding ones of audio signals that each convey speech energy; and

a controller coupled to the speech collectors and configured to perform:

signal processing on each audio signal by:

filtering the speech energy into high-frequency speech energy and low-frequency speech energy;

determining a high-frequency energy level of the high-frequency speech energy;

determining a low-frequency energy level of the low-frequency speech energy;

normalizing the high-frequency energy level with respect to a first sum of the high-frequency speech energy across the audio signals, to produce a normalized high-frequency energy level;

normalizing the low-frequency energy level with respect to a second sum of the low-frequency speech energy across the audio signals, to produce a normalized low-frequency energy level;

weighting each of the normalized high-frequency energy level and the normalized low-frequency energy level to produce a weighted normalized high-frequency energy level and a weighted normalized low-frequency energy level; and

summing the weighted normalized high-frequency energy level and the weighted normalized low-frequency energy level to produce an individual energy level; and

selecting as a preferred audio signal whichever of the audio signals has a highest individual energy level for encoding into a data packet.

13. The apparatus of claim 12 , wherein each speech collector is one of (i) a microphone beam formed using a microphone array, and (ii) a microphone that is not part of the microphone array.

14. The apparatus of claim 13 , wherein the speech collectors include one or more microphone beams formed using the microphone array and one or more microphones that are not part of the microphone array.

15. The apparatus of claim 12 , wherein:

the speech energy spans a frequency band of 0 kHz to 20 kHz.

16. The apparatus of claim 12 , wherein the weighted normalized high-frequency energy level is weighted more heavily than the weighted normalized low-frequency energy level.

17. A non-transitory computer readable medium encoded with instructions that, when executed by a controller of a system that includes speech collectors configured to convert speech from a talker to corresponding ones of audio signals that each convey speech energy, cause the controller to perform:

performing signal processing on each audio signal by:

filtering the speech energy into high-frequency speech energy and low-frequency speech energy;

determining a high-frequency energy level of the high-frequency speech energy;

determining a low-frequency energy level of the low-frequency speech energy;

normalizing the high-frequency energy level with respect to a first sum of the high-frequency speech energy across the audio signals, to produce a normalized high-frequency energy level;

normalizing the low-frequency energy level with respect to a second sum of the low-frequency speech energy across the audio signals, to produce a normalized low-frequency energy level;

weighting each of the normalized high-frequency energy level and the normalized low-frequency energy level to produce a weighted normalized high-frequency energy level and a weighted normalized low-frequency energy level; and

summing the weighted normalized high-frequency energy level and the weighted normalized low-frequency energy level to produce an individual energy level; and

selecting as a preferred audio signal whichever of the audio signals has a highest individual energy level for encoding into a data packet.

18. The non-transitory computer readable medium of claim 17 , wherein each speech collector is one of (i) a microphone beam formed using a microphone array, and (ii) a microphone that is not part of the microphone array.

19. The non-transitory computer readable medium of claim 17 , wherein:

the speech energy spans a frequency band of 0 kHz to 20 kHz.

20. The non-transitory computer readable medium of claim 17 , wherein:

the weighted normalized high-frequency energy level is weighted more heavily than the weighted normalized low-frequency energy level.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: THERKELSEN, ASBJORN
To: CISCO TECHNOLOGY, INC.
Reel/Frame 059199/0316 →
Continuity (1)
Related Publication 20230290370A1 · Sep 14, 2023
References Cited (22)
US 3992584A · Dugan · 1976 [cited by applicant]
US 4817155A · Briar · 1989 [cited by examiner]
US 4864627A · Dugan · 1989 [cited by applicant]
US 6125343A · Schuster · 2000 [cited by examiner]
US 8160877B1 · Nucci · 2012 [cited by examiner]
US 9363598B1 · Yang · 2016 [cited by applicant]
US 11483646B1 · Pan · 2022 [cited by examiner]
US 20050175190A1 · Tashev · 2005 [cited by examiner]
US 20080147415A1 · Schnell · 2008 [cited by examiner]
US 20080300702A1 · Gomez · 2008 [cited by examiner]
US 20090274318A1 · Ishibashi et al. · 2009 [cited by applicant]
US 20170026740A1 · Kirsch · 2017 [cited by examiner]
US 20190090052A1 · Radmanesh et al. · 2019 [cited by applicant]
US 20200380954A1 · Fan · 2020 [cited by examiner]
CN 1375096A · 2002 [cited by examiner]
WO WO2016100422A1 · 2016 [cited by examiner]
Vamsynagh Pedamallu, “Microphone Array Wiener Beamforming with emphasis on Reverberation”, Blekinge Institute of Technology, Jan. 2012, 73 pages. [cited by applicant]
Biamp, “Automixer basics”, last updated: Aug. 13, 2020, 10 pages, retrieved from Internet Mar. 8, 2022; https://support.biamp.com/General/Audio/Automixer_basics. [cited by applicant]
Paul Gunia, “Why Automatic Mixing is Crucial for Conferencing”, Shure, Oct. 1, 2018, 7 pages; https://www.shure.com/en-US/conferencing-meetings/ignite/why-automatic-mixing-is-crucial-for-conferencing. [cited by applicant]
Dan Dugan, “Dugan Automatic Microphone Mixers”, Dan Dugan Sound Design, 7 pages, retrieved from Internet Mar. 8, 2022; https://dandugan.com/products/. [cited by applicant]
Biamp, “Gain Sharing Auto Mixer”, 3 pages, retrieved from Internet Mar. 8, 2022; https://tesira-help.biamp.com/Component_Objects/Audio/Mixers/Gain_Sharing_Automixer.htm. [cited by applicant]
Biamp, “Gain Sharing vs. Gating Automixer”, last updated Jul. 25, 2016, 5 pages, retrieved from Internet Mar. 8, 2022; https://support.biamp.com/Tesira/Programming/Gain_Sharing_vs.Gating_Automixer. [cited by applicant]
Cited By (1)
US 12,621,389