IP Library › Granted Patent US 10,536,288
Granted Patent B1
US 10,536,288 · App. 15/841,016 · Granted Jan 14, 2020

Network conference management and arbitration via voice-capturing devices

Inventors: Jonathan Alan Leblang (Menlo Park, CA); Milo Oostergo (Issaquah, WA); James L. Ford (Bellevue, WA); Kevin Crews (Seattle, WA)
Assignee: Amazon Technologies, Inc.
H04L12/1822G10L15/22H04L12/1818H04L67/306H04M3/563H04W40/244G10L13/00G10L15/1815G10L2015/223H04M2242/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,536,288
App. No.
15/841,016
Granted
Jan 14, 2020
Kind
B1
Abstract

Systems and methods are provided for managing a conference call with multiple voice-enabled and voice-capturing devices, such as smart speakers. Reproduced, duplicate voice commands can cause unexpected results in a conference call. The voice commands can be determined to be received from the same conference call. A voice command for a particular voice-enabled device can be selected based on an energy level of an audio signal, event data, time data, and/or user identification.

Claims (75)

1. A system for managing a conference meeting, the system comprising:

an electronic data store; and

one or more computer hardware processors in communication with the electronic data store, the one or more computer hardware processors configured to execute computer-executable instructions to at least:

receive a first audio signal from a first voice-enabled device;

identify a first user profile based on the first audio signal, wherein identifying the first user profile comprises performing speaker recognition on the first audio signal and using a first user voice profile;

receive a second audio signal from a second voice-enabled device;

identify the first user voice profile from the second audio signal;

generate a group of voice-enabled devices based on identifying the first user voice profile from both the first audio signal and the second audio signal, wherein the group comprises the first voice-enabled device and the second voice-enabled device, wherein the first voice-enabled device and the second voice-enabled device are in different rooms, wherein voice input from a conference call participant is received by the first voice-enabled device and the second voice-enabled device, and wherein the first voice-enabled device is associated with a first account different than a second account associated with the second voice-enabled device;

receive a third audio signal from the first voice-enabled device;

identify a voice command from the third audio signal;

determine, using the group, that the voice command was also received by the second voice-enabled device;

determine that the voice command corresponds to a command to leave a conference call associated with a meeting;

identify a first user profile based on the third audio signal, wherein identifying the first user profile comprises identifying the first user voice profile from the third audio signal;

identify an association between the first user profile and the first voice-enabled device;

select the first voice-enabled device, from the first voice-enabled device and the second voice-enabled device, based on the association between the first user profile and the first voice-enabled device; and

execute the voice command, wherein execution of the first voice command causes the first voice-enabled device to disconnect from the conference call.

2. The system of claim 1 , wherein the one or more computer hardware processors are further configured to:

receive, from the first voice-enabled device, an indication that a first beacon associated with the first user profile is received by the first voice-enabled device, wherein the first beacon is transmitted to the first voice-enabled device via at least one of radio-frequency identification (RFID) or Bluetooth, wherein the presence of the first beacon indicates that the user of the first user profile is near the first voice-enabled device.

3. The system of claim 1 , wherein the one or more computer hardware processors are further configured to:

retrieve event data associated with the meeting;

retrieve, from the event data, a first user profile identifier; and

retrieve the first user profile using the first user profile identifier, wherein a first user associated with the first user profile is invited to the meeting.

4. The system of claim 1 , wherein performing speaker recognition on the first audio signal comprises:

comparing the first audio signal to a baseline voice model, wherein comparing the first audio signal to the baseline model identifies a first difference between the first audio signal and the baseline model; and

selecting, from a plurality of user voice profiles, the first user voice profile, wherein selecting the first user voice profile further comprises:

identifying that the first difference is within a first threshold of a first feature of the first user voice profile.

5. A computer-implemented method comprising:

receiving a first audio signal from a first voice-enabled device;

identifying a first user profile based on the first audio signal, wherein identifying the first user profile comprises performing speaker recognition on the first audio signal and using a first user voice profile;

receiving a second audio signal from a second voice-enabled device;

identifying the first user voice profile from the second audio signal;

generating a group of voice-enabled devices based on identifying the first user voice profile from both the first audio signal and the second audio signal, wherein the group comprises the first voice-enabled device and the second voice-enabled device, wherein voice input from a conference call participant is received by the first voice-enabled device and the second voice-enabled device, and wherein the first voice-enabled device is associated with a first account different than a second account associated with the second voice-enabled device;

receiving a third audio signal from the first voice-enabled device;

identifying a voice command from a third audio signal;

determining, using the group, that the voice command was also received by the second voice-enabled device;

identifying a first user profile based on the third audio signal, wherein identifying the first user profile comprises identifying the first user voice profile from the third audio signal;

determining that the first user profile is associated with the first voice-enabled device;

selecting the first voice-enabled device, from the first voice-enabled device and the second voice-enabled device, based on the determining that the first user profile is associated with the first voice-enabled device; and

executing the voice command for the first voice-enabled device instead of the second voice-enabled device.

6. The computer-implemented method of claim 5 , further comprising:

receiving, from the first voice-enabled device, an indication that a first beacon associated with the first user profile is received by the first voice-enabled device.

7. The computer-implemented method of claim 6 , wherein the first beacon is transmitted to the first voice-enabled device via radio-frequency identification (RFID) or Bluetooth.

8. The computer-implemented method of claim 5 , further comprising:

determining that both the first audio signal and the second audio signal were received within a threshold period of time.

9. The computer-implemented method of claim 5 , wherein identifying the first user profile based on the first audio signal comprises:

identifying the first user voice profile from the first audio signal, wherein the first user voice profile comprises a feature of a first user's speech utterance; and

accessing a second entry that indicates an association between the first user voice profile and the first user profile.

10. The computer-implemented method of claim 5 , wherein the first user profile comprises a first entry, and wherein the first entry further indicates that the first voice-enabled device is registered to the first user profile.

11. A system comprising:

an electronic data store; and

one or more computer hardware processors in communication with the electronic data store, the one or more computer hardware processors configured to execute computer-executable instructions to at least:

receive a first audio signal from a first voice-enabled device;

identify a first user profile based on the first audio signal, wherein identifying the first user profile comprises performing speaker recognition on the first audio signal and using a first user voice profile;

receive a second audio signal from a second voice-enabled device;

identify the first user voice profile from the second audio signal;

generate a group of voice-enabled devices based on identifying the first user voice profile from both the first audio signal and the second audio signal, wherein the group comprises the first voice-enabled device and the second voice-enabled device, wherein voice input from a conference call participant is received by the first voice-enabled device and the second voice-enabled device, and wherein the first voice-enabled device is associated with a first account different than a second account associated with the second voice-enabled device;

receive a third audio signal from the first voice-enabled device;

identify a voice command from the third audio signal;

determine, using the group, that the voice command was also received by the second voice-enabled device;

identify a first user profile based on the third audio signal;

determine that the first user profile is associated with the first voice-enabled device;

select the first voice-enabled device, from the first voice-enabled device and the second voice-enabled device, based at least on determining that the first user profile is associated with the first voice-enabled device; and

execute the voice command for the first voice-enabled device instead of the second voice-enabled device.

12. The system of claim 11 , wherein the one or more computer hardware processors are further configured to:

receive, from the first voice-enabled device, an indication that a first beacon associated with the first user profile is received by the first voice-enabled device.

13. The system of claim 12 , wherein the first voice-enabled device is located in a room with a speaker device separate from the first voice-enabled device, and wherein audio for a conference call session is provided in the room by the speaker device.

14. The system of claim 11 , wherein the computer hardware processor is further configured to at least:

retrieve event data associated with a meeting;

retrieve the first user profile based on information in the event data.

15. The system of claim 11 , wherein identifying the first user profile based on the first audio signal comprises:

identifying the first user voice profile from the first audio signal, wherein the first user voice profile comprises a feature of a first user's speech utterance, the first user voice profile associated with the first user profile.

16. The system of claim 15 , wherein the one or more computer hardware processors are further configured to:

prompt a user to conduct a voice registration process;

receive voice input from the voice registration process; and

generate the first user voice profile based on the voice input.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 17, 2018
From: LEBLANG, JONATHAN ALAN; OOSTERGO, MILO; FORD, JAMES L.; CREWS, KEVIN
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 045833/0061 →
Cited By (9)
US 12,211,515 US 12,284,048 US 12,334,078 US 12,475,883 US 12,587,576 US 12,602,446 US 12,652,212 US 12,687,624 US 12,707,245