AUTOMATED ASSISTANTS WITH CONFERENCE CAPABILITIES
Techniques are described related to enabling automated assistants to enter into a “conference mode” in which they can “participate” in meetings between multiple human participants and perform various functions described herein. In various implementations, an automated assistant implemented at least in part on conference computing device(s) may be set to a conference mode in which the automated assistant performs speech-to-text processing on multiple distinct spoken utterances, provided by multiple meeting participants, without requiring explicit invocation prior to each utterance. The automated assistant may perform semantic processing on first text generated from the speech-to-text processing of one or more of the spoken utterances, and generate, based on the semantic processing, data that is pertinent to the first text. The data may be output to the participants at conference computing device(s). The automated assistant may later determine that the meeting has concluded, and may be set to a non-conference mode.
1 . A method implemented using one or more processors, comprising:
setting an automated assistant implemented at least in part on one or more conference computing devices to a conference mode in which the automated assistant performs speech-to-text processing on multiple distinct spoken utterances provided by multiple participants during a meeting facilitated by the one or more conference computing devices, without requiring explicit invocation of the automated assistant prior to each of the multiple distinct spoken utterances;
capturing, via one or more microphones of the one or more conference computing devices, a first spoken utterance provided by a first participant of the multiple participants;
capturing, via the one or more microphones, a second spoken utterance provided by a second participant of the multiple participants, wherein the second participant is different from the first participant;
performing speech-to-text processing on the first spoken utterance and the second spoken utterance to generate, respectively, a first text and a second text;
automatically performing, by the automated assistant, semantic processing on both the first text and the second text to combine content from the first text and the second text into a single consolidated query;
generating, by the automated assistant, based on results of the single consolidated query, data that is pertinent to the meeting; and
outputting, by the automated assistant at one or more of the conference computing devices, the data that is pertinent to the results while the automated assistant is in the conference mode.
2 . The method of claim 1 , wherein performing the semantic processing to combine the content includes identifying a first topic from the first text and identifying a second topic from the second text, and wherein the single consolidated query is generated based on a relationship between the first topic and the second topic.
3 . The method of claim 1 , wherein the first text includes a location and the second text includes a time, and the automated assistant generates the single consolidated query to include both the location from the first participant and the time from the second participant.
4 . The method of claim 1 , further comprising maintaining a meeting dialog context that includes at least the first text provided by the first participant, and wherein performing the semantic processing on the second text includes using the meeting dialog context to disambiguate one or more tokens of the second text provided by the second participant prior to generating the single consolidated query.
5 . The method of claim 1 , further comprising identifying an output modality used by a first conference computing device of the one or more conference computing devices, wherein the data that is pertinent to the results is output at a frequency that is selected based on the identified output modality.
6 . The method of claim 1 , wherein the single consolidated query comprises a search query, and the results comprise search results.
7 . The method of claim 1 , further comprising: monitoring, by the automated assistant, the meeting for a pause of at least a predetermined time interval; and wherein the data that is pertinent to the results is output in response to detecting the pause.
8 . A system comprising one or more processors and memory storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:
setting an automated assistant implemented at least in part on one or more conference computing devices to a conference mode in which the automated assistant performs speech-to-text processing on multiple distinct spoken utterances provided by multiple participants during a meeting facilitated by the one or more conference computing devices, without requiring explicit invocation of the automated assistant prior to each of the multiple distinct spoken utterances;
capturing, via one or more microphones of the one or more conference computing devices, a first spoken utterance provided by a first participant of the multiple participants;
capturing, via the one or more microphones, a second spoken utterance provided by a second participant of the multiple participants, wherein the second participant is different from the first participant;
performing speech-to-text processing on the first spoken utterance and the second spoken utterance to generate, respectively, a first text and a second text;
automatically performing, by the automated assistant, semantic processing on both the first text and the second text to combine content from the first text and the second text into a single consolidated query;
generating, by the automated assistant, based on results of the single consolidated query, data that is pertinent to the meeting; and
outputting, by the automated assistant at one or more of the conference computing devices, the data that is pertinent to the results while the automated assistant is in the conference mode.
9 . The system of claim 8 , wherein performing the semantic processing to combine the content includes identifying a first topic from the first text and identifying a second topic from the second text, and wherein the single consolidated query is generated based on a relationship between the first topic and the second topic.
10 . The system of claim 8 , wherein the first text includes a location and the second text includes a time, and the automated assistant generates the single consolidated query to include both the location from the first participant and the time from the second participant.
11 . The system of claim 8 , the operations further comprising maintaining a meeting dialog context that includes at least the first text provided by the first participant, and wherein performing the semantic processing on the second text includes using the meeting dialog context to disambiguate one or more tokens of the second text provided by the second participant prior to generating the single consolidated query.
12 . The system of claim 8 , the operations further comprising identifying an output modality used by a first conference computing device of the one or more conference computing devices, wherein the data that is pertinent to the results is output at a frequency that is selected based on the identified output modality.
13 . The system of claim 8 , wherein the single consolidated query comprises a search query, and the results comprise search results.
14 . The system of claim 8 , the operations further comprising: monitoring, by the automated assistant, the meeting for a pause of at least a predetermined time interval; and wherein the data that is pertinent to the results is output in response to detecting the pause.
15 . At least one non-transitory computer-readable medium comprising instructions that, when executed by one or more processors, cause the one or more processors to perform operations comprising:
setting an automated assistant implemented at least in part on one or more conference computing devices to a conference mode in which the automated assistant performs speech-to-text processing on multiple distinct spoken utterances provided by multiple participants during a meeting facilitated by the one or more conference computing devices, without requiring explicit invocation of the automated assistant prior to each of the multiple distinct spoken utterances;
capturing, via one or more microphones of the one or more conference computing devices, a first spoken utterance provided by a first participant of the multiple participants;
capturing, via the one or more microphones, a second spoken utterance provided by a second participant of the multiple participants, wherein the second participant is different from the first participant;
performing speech-to-text processing on the first spoken utterance and the second spoken utterance to generate, respectively, a first text and a second text;
automatically performing, by the automated assistant, semantic processing on both the first text and the second text to combine content from the first text and the second text into a single consolidated query;
generating, by the automated assistant, based on results of the single consolidated query, data that is pertinent to the meeting; and
outputting, by the automated assistant at one or more of the conference computing devices, the data that is pertinent to the results while the automated assistant is in the conference mode.
16 . The at least one non-transitory computer-readable medium of claim 15 , wherein performing the semantic processing to combine the content includes identifying a first topic from the first text and identifying a second topic from the second text, and wherein the single consolidated query is generated based on a relationship between the first topic and the second topic.
17 . The at least one non-transitory computer-readable medium of claim 15 , wherein the first text includes a location and the second text includes a time, and the automated assistant generates the single consolidated query to include both the location from the first participant and the time from the second participant.
18 . The at least one non-transitory computer-readable medium of claim 15 , the operations further comprising maintaining a meeting dialog context that includes at least the first text provided by the first participant, and wherein performing the semantic processing on the second text includes using the meeting dialog context to disambiguate one or more tokens of the second text provided by the second participant prior to generating the single consolidated query.
19 . The at least one non-transitory computer-readable medium of claim 15 , the operations further comprising identifying an output modality used by a first conference computing device of the one or more conference computing devices, wherein the data that is pertinent to the results is output at a frequency that is selected based on the identified output modality.
20 . The at least one non-transitory computer-readable medium of claim 15 , the operations further comprising: monitoring, by the automated assistant, the meeting for a pause of at least a predetermined time interval; and wherein the data that is pertinent to the results is output in response to detecting the pause.