IP Library › Granted Patent US 11,115,630
Granted Patent B1
US 11,115,630 · App. 16/359,520 · Granted Sep 7, 2021

Custom and automated audio prompts for devices

Inventors: Elliott Lemberger (Santa Monica, CA); John Modestine (Los Angeles, CA); Kevin Park (Los Angeles, CA); Richard Carter Mosher (Atherton, CA); Trevor Grolle (Mesa, AZ); Kirk David Bacon (Westminster, CA)
Assignee: Amazon Technologies, Inc.
H04N7/186G06F3/167G06K9/00671G06K9/00711G06K9/00771H04L67/125G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,115,630
App. No.
16/359,520
Granted
Sep 7, 2021
Kind
B1
Abstract

A network-connected security device is communicatively coupled to an audio/video (A/V) recording and communication device having a camera and a speaker. A method receives video data captured by the camera, and performs an object recognition algorithm upon the received video data to identify an object therein. The method performs a table lookup using the identified object, into a data structure that associates objects with at least one description of a predefined voice message. The method selects a description of a predefined voice message associated with the identified object, and transmits the selected description's predefined voice message to the A/V recording and communication device for output through the speaker.

Claims (91)

1. A method comprising:

receiving audio prompt data from a user device;

receiving, from the user device, a request to associate the audio prompt data with an object;

storing first identifier data associated with the audio prompt data;

storing second identifier data associated with the object;

receiving first image data generated by an electronic device;

determining that the first image data represents the object;

based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and

sending the audio prompt data to the electronic device.

2. The method as recited in claim 1 , further comprising:

receiving second image data representing the object,

and wherein

the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.

3. The method as recited in claim 1 , wherein the object is a person, and wherein the method further comprises:

receiving second image data representing the person,

and wherein:

the second identifier data represents an identity of the person; and

the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data.

4. The method as recited in claim 1 , wherein the first identifier data represents a description of the audio prompt data, and wherein the method further comprises:

based at least in part on the determining that the first image data represents the object, determining that the description is associated with the second identifier data,

and wherein the selecting the audio prompt data comprises selecting, based at least in part on the description being associated with the second identifier data, the audio prompt data.

5. The method as recited in claim 1 , further comprising:

determining that the first identifier data is associated with the audio prompt data;

determining that the first identifier data is associated with additional audio prompt data;

determining a first value associated with the audio prompt data;

determining a second value associated with the additional audio prompt data; and

determining that the first value is greater than the second value,

and wherein the selecting the audio prompt data is further based at least in part on the determining that the first value is greater than the second value.

6. The method as recited in claim 1 , further comprising:

sending a message to the user device, the message including at least an image represented by the first image data and a description of the audio prompt data; and

receiving, from the user device, a request to output the audio prompt data,

and wherein the sending the audio prompt data to the electronic device is based at least in part on the receiving the request to output the audio prompt data.

7. The method as recited in claim 1 , further comprising:

receiving audio data generated by the electronic device; and

identifying user speech represented by the audio data,

and wherein the selecting the audio prompt data is further based at least in part on the identifying the user speech.

8. The method as recited in claim 1 , wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the user device, at least:

the first identifier data associated with the audio prompt data; and

the second identifier associated with the object.

9. The method as recited in claim 1 , wherein the storing the first identifier data associated with the audio prompt data comprises storing the first identifier data that represents a description associated with the audio prompt data, the description including one or more words that identify the object.

10. An electronic device comprising:

a camera;

one or more speakers;

one or more processors; and

one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the electronic device to perform operations comprising:

receiving, from a system, audio prompt data;

receiving, from the system, a request to associate the audio prompt data with an object;

storing first identifier data associated with the audio prompt data;

storing second identifier data associated with the object;

generating first image data using the camera;

determining that the first image data represents the object;

based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and

outputting, using the one or more speakers, sound represented by the audio prompt data.

11. The electronic device as recited in claim 10 , the one or more computer-readable media storing further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising:

receiving second image data representing the object,

and wherein

the determining that the first image data represents the object comprises determining, based at least in part on the second image data, that the first image data represents the object.

12. The electronic device as recited in claim 10 , wherein the second identifier data represents in identity of the object, the object being a person, and wherein the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising:

receiving second image data representing the person

and wherein:

the determining that the first image data represents the person comprises determining, based at least in part on the second image data, the identity of the person represented by the first image data; and

the selecting the audio prompt data comprises selecting, based at least in part on the identity, the audio prompt data.

13. The electronic device as recited in claim 10 , wherein the receiving the request to associate the audio prompt data with the object comprises receiving, from the system, at least:

the first identifier data associated with the audio prompt data; and

the second identifier data associated with the object.

14. The electronic device as recited in claim 10 , wherein:

the first identifier data represents a description associated with the audio prompt data, the description including one or more words that identify the object;

the one or more computer-readable media store further instructions that, when executed by the one or more processors, cause the electronic device to perform further operations comprising determining, based at least in part on the first image data representing the object, that the description includes the one or more words that identify the object; and

the selecting the audio prompt data is based at least in part on the determining that the description includes the one or more words that identify the object.

15. A method comprising:

receiving, from a system, audio prompt data;

receiving, from the system, a request to associate the audio prompt data with an object;

storing identifier data associated with the audio prompt data, the identifier data representing a description that includes one or more words that identify the object;

generating first image data using a camera;

determining that the first image data represents the object;

based at least in part on the determining that the first image data represents the object, selecting the audio prompt data; and

outputting sound represented by the audio prompt data.

16. The method as recited in claim 15 , wherein the receiving the request to associate the audio prompt data with the object comprises at least receiving, from the system, the identifier data associated with the audio prompt data.

17. The method as recited in claim 15 , further comprising:

receiving second image data representing the object,

wherein the determining that the first image data represents the object comprises at least:

analyzing the first image data using at least the second image data; and

determining that the first image data represents the object.

18. The method as recited in claim 15 , wherein the determining that the first image data represents the object comprises at least

determining that the first image data represents one or more characteristics; and

determining that the one or more characteristics are associated with the object.

19. The method as recited in claim 15 , wherein the selecting the audio prompt data comprises at least:

determining that the identifier data represents the description;

determining that the description includes the one or more words that identify the object; and

selecting the audio prompt data based at least in part on the description including the one or more words that identify the object.

20. The method as recited in claim 15 , wherein the outputting of sound represented by the audio prompt data comprises outputting the sound that includes user speech, the audio prompt data representing the user speech.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 9, 2019
From: LEMBERGER, ELLIOTT; MODESTINE, JOHN; PARK, KEVIN; MOSHER, RICHARD CARTER; GROLLE, TREVOR; BACON, KIRK DAVID
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 048835/0046 →
Continuity (1)
Provisional Application 62649504 · Mar 28, 2018
Cited By (4)
US 12,423,394 US 12,625,218 US 12,647,624 US 12,705,962