IP Library Granted Patent US 8,126,705
Granted Patent B2
US 8,126,705 · App. 12/615,239 · Granted Feb 28, 2012

System and method for automatically adjusting floor controls for a conversation

Assignee: Palo Alto Research Center Incorporated
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,126,705
App. No.
12/615,239
Granted
Feb 28, 2012
Kind
B2
Abstract

A system and method for automatically adjusting floor controls for a conversation is provided. Audio streams are received, which each originate from an audio source. Floor controls for a current configuration including at least a portion of the audio streams are maintained. Conversational characteristics shared by two or more of the audio sources are determined. Possible configurations for the audio streams are identified based on the conversational characteristics. An analysis of the current configuration and the possible configurations is performed. A change threshold is applied to the analysis. When the analysis satisfies the change threshold, the floor controls are automatically adjusted. The audio streams are mixed into one or more outputs based on the adjusted floor controls.

Claims (44)

1. A system for identifying participants in a conversation, comprising:

input modules to receive a plurality of audio streams each originating from an audio source;

a floor analysis module to maintain floor controls for a current conversational configuration comprising at least a portion of the audio sources, to analyze the audio sources for conversational characteristics shared by two or more of the audio sources, to determine possible conversational configurations of the audio sources based on the conversational characteristics, to assign a probability to at least one of the possible conversational configurations, to select one of the possible conversational configurations based on the probability as a most probable conversation configuration, and to adjust the floor controls for the current conversational configuration according to the most probable conversational configuration;

an audio mixer to mix the audio streams based on the adjusted floor controls and providing the mixed audio streams as output; and

a processor to execute the input modules, the floor analysis module, and the audio mixer.

2. A system according to claim 1 , wherein the output provides at least one of signals capable of driving an audio sound reproduction mechanism and signals capable of electronic recording.

3. A system according to claim 1 , wherein the floor controls are automatically adjusted when the current conversational configuration is maintained for a minimum number of timeslices and the most probable conversational configuration is different from the current conversational configuration.

4. A system according to claim 1 , wherein the floor controls are automatically adjusted when the current conversational configuration is different from the most probable conversational configuration for a minimum number of consecutive timeslices.

5. A system according to claim 1 , wherein the floor controls are automatically adjusted when the current conversational configuration is different from the most probable conversational configuration for more than a minimum net number of timeslices.

6. A system according to claim 1 , wherein the conversational characteristics are determined by at least one of a turn taking analysis, a referential action analysis, and a responsive action analysis.

7. A system according to claim 6 , further comprising:

a turn taking analysis module to perform the turn taking analysis by aligning the audio streams from two or more of the sources, monitoring the aligned audio stream, and identifying the conversational characteristics from the aligned audio stream.

8. A system according to claim 6 , further comprising:

a referential action analysis module to perform the referential action analysis by determining that at least one of the audio streams occurs during a time window and identifying at least one conversational characteristic comprising one or more of a name and a name variant for at least one of the other audio sources during the time window.

9. A system according to claim 6 , further comprising:

a responsive action analysis module to perform the responsive action analysis by determining that at least one of the audio streams occurs during a time window, by identifying one or more conversational characteristics comprising a backchannel response during the time window, and by matching the back channel response to vocalizations of another audio stream.

10. A system according to claim 1 , wherein the conversational characteristics comprise at least one of sustained periods of overlapping speech, a lack of correlation for a beginning of speech, speech beginning at a transition relevance place, communication identifiers, change in speech volume, speech energy amplitude, common content, prosody, and backchannel responses.

11. A computer-implemented method for identifying participants in a conversation, comprising:

receiving a plurality of audio streams each originating from an audio source;

maintaining floor controls for a current conversational configuration comprising at least a portion of the audio sources;

analyzing the audio streams for conversational characteristics shared by two or more of the audio sources and determining possible conversational configurations of the audio sources based on the conversational characteristics;

assigning a probability to at least one of the possible conversational configurations and selecting one of the possible conversational configurations based on the probability as a most probable conversation configuration;

adjusting the floor controls for the current conversational configuration according to the most probable conversational configuration; and

mixing the audio streams based on the adjusted floor controls and providing the mixed audio streams as output.

12. A computer-implemented method according to claim 11 , wherein the output provides at least one of signals capable of driving an audio sound reproduction mechanism and signals capable of electronic recording.

13. A computer-implemented method according to claim 11 , wherein the floor controls are automatically adjusted when the current conversational configuration is maintained for a minimum number of timeslices and the most probable conversational configuration is different from the current conversational configuration.

14. A computer-implemented method according to claim 11 , wherein the floor controls are automatically adjusted when the current conversational configuration is different from the most probable conversational configuration for a minimum number of consecutive timeslices.

15. A computer-implemented method according to claim 11 , wherein the floor controls are automatically adjusted when the current conversational configuration is different from the most probable conversational configuration for more than a minimum net number of timeslices.

16. A computer-implemented method according to claim 11 , wherein the conversational characteristics comprise at least one of sustained periods of overlapping speech, a lack of correlation for a beginning of speech, speech beginning at a transition relevance place, communication identifiers, change in speech volume, speech energy amplitude, common content, prosody, and backchannel responses.

17. A computer-implemented method according to claim 11 , wherein the conversational characteristics are determined by at least one of a turn taking analysis, a referential action analysis, and a responsive action analysis.

18. A computer-implemented method according to claim 17 , further comprising:

performing the turn taking analysis, comprising:

aligning the audio streams from two or more of the sources;

monitoring the aligned audio streams; and

identifying the conversational characteristics from the aligned audio stream.

19. A computer-implemented method according to claim 17 , further comprising:

performing the referential action analysis, comprising:

determining that at least one of the audio streams occurs during a time window; and

identifying at least one conversational characteristic comprising one or more of a name and a name variant for at least one of the other audio sources during the time window.

20. A computer-implemented method according to claim 17 , further comprising:

performing the responsive action analysis, comprising:

determining that at least one of the audio streams occurs during a time window;

identifying one or more conversational characteristics comprising a backchannel response during the time window; and

matching the back channel response to vocalizations of another audio stream.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2015
From: PALO ALTO RESEARCH CENTER INCORPORATED
To: III HOLDINGS 6, LLC
Reel/Frame 036198/0543 →
Continuity (3)
Continuation 10414912 · Apr 16, 2003
Provisional Application 60450724 · Feb 28, 2003
Related Publication 20100057445A1 · Mar 4, 2010