IP Library › Granted Patent US 9,210,269
Granted Patent B2
US 9,210,269 · App. 13/664,640 · Granted Dec 8, 2015

Active speaker indicator for conference participants

Inventors: Yanghua Liu (San Jose, CA); Weidong Chen (Palo Alto, CA); Biren Gandhi (San Jose, CA); Raghurama Bhat (Cupertino, CA); Joseph Fouad Khouri (San Jose, CA); John Joseph Houston (San Jose, CA); Brian Thomas Toombs (San Jose, CA)
Assignee: Cisco Technology, Inc.
H04M3/563G10L17/00H04N7/15
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,210,269
App. No.
13/664,640
Granted
Dec 8, 2015
Kind
B2
Abstract

In one embodiment, a method includes receiving requests to join a conference from a plurality of user devices proximate a first endpoint. The requests include a username. The method also includes receiving an audio signal for the conference from the first endpoint. The first endpoint is operable to capture audio proximate the first endpoint. The method also includes transmitting the audio signal to a second endpoint, remote from the first endpoint. The method also includes identifying, by a processor, an active speaker proximate the first endpoint based on information received from the plurality of user devices.

Claims (81)

1. A system, comprising:

a processor; and

a non-transitory computer-readable storage medium embodying software that is operable when executed by the processor to:

receive requests to join a conference from a plurality of user devices proximate a first endpoint, the requests comprising a username;

receive an audio signal for the conference from the first endpoint, the first endpoint operable to capture audio proximate the first endpoint;

transmit the audio signal to a second endpoint, remote from the first endpoint;

receive a plurality of audio energy values from the plurality of user devices proximate the first endpoint, the audio energy values associated with the audio signal;

identify an active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and

transmit an identity of the identified active speaker to the second endpoint while continuing to transmit audio signals received from the first endpoint to the second endpoint wherein: the active speaker is a first active speaker; identify a second active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and transmit an identity for both the first active speaker and the second active speaker to the second endpoint.

2. The system of claim 1 , wherein:

the software is further operable when executed to identify the active speaker proximate the first endpoint by comparing the plurality of audio energy values.

3. The system of claim 2 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by calibrating the plurality of user devices.

4. The system of claim 2 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by:

identifying a greatest audio energy value user device of the plurality of user devices; and

identifying the username transmitted by the greatest audio energy value user device as the active speaker.

5. The system of claim 1 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by:

receiving a first audio energy value from a first user device of the plurality of user devices;

receiving a second audio energy value from a second user device of the plurality of user devices; and

comparing the first audio energy value with the second audio energy value.

6. A method, comprising:

receiving requests to join a conference from a plurality of user devices proximate a first endpoint, the requests comprising a username;

receiving an audio signal for the conference from the first endpoint, the first endpoint operable to capture audio proximate the first endpoint;

transmitting the audio signal to a second endpoint, remote from the first endpoint;

receiving a plurality of audio energy values from the plurality of user devices, the audio energy values associated with the audio signal;

identifying, by a processor, an active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and

transmit an identity of the identified active speaker to the second endpoint while continuing to transmit audio signals received from the first endpoint to the second endpoint wherein: the active speaker is a first active speaker; identifying a second active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and transmitting an identity for both the first active speaker and the second active speaker to the second endpoint.

7. The method of claim 6 , wherein:

identifying the active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices comprises comparing the plurality of audio energy values.

8. The method of claim 7 , wherein identifying the active speaker proximate the first endpoint based on information received from the plurality of user devices further comprises calibrating the plurality of user devices.

9. The method of claim 7 , wherein identifying the active speaker proximate the first endpoint based on information received from the plurality of user devices further comprises:

identifying a greatest audio energy value user device of the plurality of user devices; and

identifying the username transmitted by the greatest audio energy value user device as the active speaker.

10. The method of claim 6 , wherein identifying the active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices comprises:

receiving a first audio energy value from a first user device of the plurality of user devices;

receiving a second audio energy value from a second user device of the plurality of user devices; and

comparing the first audio energy value with the second audio energy value.

11. One or more non-transitory computer-readable storage media embodying software that is operable when executed by a processor to:

receive requests to join a conference from a plurality of user devices proximate a first endpoint, the requests comprising a username;

receive an audio signal for the conference from the first endpoint, the first endpoint operable to capture audio proximate the first endpoint;

transmit the audio signal to a second endpoint, remote from the first endpoint;

receive a plurality of audio energy values from the plurality of user devices, the audio energy values associated with the audio signal;

identify an active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and

transmit an identity of the identified active speaker to the second endpoint while continuing to transmit audio signals received from the first endpoint to the second endpoint wherein: the active speaker is a first active speaker; identify a second active speaker proximate the first endpoint based on the audio energy values received from the plurality of user devices; and transmit an identity for both the first active speaker and the second active speaker to the second endpoint.

12. The media of claim 11 , wherein:

the software is further operable when executed to identify the active speaker proximate the first endpoint by comparing the plurality of audio energy values.

13. The media of claim 12 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by calibrating the plurality of user devices.

14. The media of claim 12 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by:

identifying a greatest audio energy value user device of the plurality of user devices; and

identifying the username transmitted by the greatest audio energy value user device as the active speaker.

15. The media of claim 11 , wherein the software is further operable when executed to identify the active speaker proximate the first endpoint by:

receiving a first audio energy value from a first user device of the plurality of user devices;

receiving a second audio energy value from a second user device of the plurality of user devices; and

comparing the first audio energy value with the second audio energy value.

16. A system, comprising:

a processor; and

a non-transitory computer-readable storage medium embodying software that is operable when executed by the processor to:

receive requests to join a conference from a plurality of user devices proximate a first endpoint, the requests comprising a username;

receive registration audio signals associated with a plurality of users;

generate voice identification information for the plurality of users based on the received registration audio signals;

store the voice identification information in a database;

receive an audio signal for the conference from the first endpoint, the first endpoint operable to capture audio proximate the first endpoint;

transmit the audio signal to a second endpoint, remote from the first endpoint;

identify an active speaker proximate the first endpoint based on the audio signal and the voice identification information; and

transmit an identity of the identified active speaker to the second endpoint while continuing to transmit audio signals received from the first endpoint to the second endpoint wherein: the active speaker is a first active speaker; identify a second active speaker proximate the first endpoint based on the audio signals received from the first endpoint; and transmit an identity for both the first active speaker and the second active speaker to the second endpoint.

17. The system of claim 16 , wherein the software is further operable when executed to:

select a subset of the plurality of users; and

identify the active speaker proximate the first endpoint based on the audio signal and the voice identification information for the subset of the plurality of users.

18. The system of claim 17 , wherein the software is further operable when executed to:

select the subset of the plurality of users based on the received requests to join the conference.

19. The system of claim 17 , wherein:

the database further stores location information associated with the plurality of users; and

the software is further operable when executed to:

determine a location of the first endpoint; and

select the subset of the plurality of users based on the location information associated with the plurality of users.

20. The system of claim 16 , wherein the software is further operable when executed to:

receive active speaker detection feedback, the feedback indicating the accuracy of the active speaker identification; and

update the voice identification information based on the active speaker detection feedback.

21. The system of claim 16 , wherein the software is further operable when executed to:

receive a video signal for the conference from the first endpoint;

process the video signal to produce a processed video signal that includes active speaker identification information; and

transmit the processed video signal to the second endpoint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2012
From: LIU, YANGHUA; CHEN, WEIDONG; GANDHI, BIREN; BHAT, RAGHURAMA; KHOURI, JOSEPH FOUAD; HOUSTON, JOHN JOSEPH; TOOMBS, BRIAN THOMAS
To: CISCO TECHNOLOGY, INC.
Reel/Frame 029216/0060 →
Continuity (1)
Related Publication 20140118472A1 · May 1, 2014