IP Library Granted Patent US 9,031,847
Granted Patent B2
US 9,031,847 · App. 13/297,116 · Granted May 12, 2015

Voice-controlled camera operations

Inventors: Raman Kumar Sarin (Redmond, WA); Joseph H. Matthews, III (Woodinville, WA); James Kai Yu Lau (Bellevue, WA); Monica Estela Gonzalez Veron (Seattle, WA); Jae Pum Park (Bellevue, WA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/22H04N5/232G10L2015/223G06F3/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,031,847
App. No.
13/297,116
Granted
May 12, 2015
Kind
B2
Abstract

A computing device (e.g., a smart phone, a tablet computer, digital camera, or other device with image capture functionality) causes an image capture device to capture one or more digital images based on audio input (e.g., a voice command) received by the computing device. For example, a user's voice (e.g., a word or phrase) is converted to audio input data by the computing device, which then compares (e.g., using an audio matching algorithm) the audio input data to an expected voice command associated with an image capture application. In another aspect, a computing device activates an image capture application and captures one or more digital images based on a received voice command. In another aspect, a computing device transitions from a low-power state to an active state, activates an image capture application, and causes a camera device to capture digital images based on a received voice command.

Claims (40)

1. A computer-implemented method comprising:

receiving audio input at a computing device;

while the computing device is in a low-power state, performing a single voice recognition event by:

comparing the received audio input with an expected audio command associated with an image capture operation, the expected audio command comprising at least one word; and

based on the comparing, determining that the received audio input matches the at least one word; and

in response to performing the single voice recognition event, without receiving further audio input, both transitioning the computing device from a low-power state to an active state and causing an image capture device to capture one or more digital images.

2. The method of claim 1 wherein the audio input comprises voice data from one or more human voices.

3. The method of claim 1 wherein the comparing is based at least in part on training.

4. The method of claim 1 wherein the comparing is based at least in part on previously recognized audio commands.

5. The method of claim 1 wherein the comparing is based at least in part on one or more contextual cues.

6. The method of claim 1 wherein the expected audio command is based at least in part on a user modification.

7. The method of claim 6 wherein the expected audio command comprises a default audio command, and wherein the user modification comprises a modification of the default audio command.

8. The method of claim 1 wherein the expected audio command is a member of a set of audio commands associated with functionality of the computing device, the set of audio commands further comprising a keep-image command, a delete-image command, a record-video command, a stop-recording command, and a show-photo command.

9. The method of claim 1 wherein the expected audio command is associated with a macro comprising a series of plural functions of the computing device.

10. The method of claim 1 wherein the expected audio command is a member of a set of audio commands associated with functionality of a plurality of computing devices comprising a smart phone, a digital camera, a tablet computer, and a gaming console.

11. The method of claim 1 wherein the computing device comprises a smart phone, and wherein the comparing is performed by the smart phone.

12. The method of claim 1 wherein the comparing is performed at least in part at a remote server.

13. The method of claim 1 wherein the comparing is performed on a dedicated processor.

14. The method of claim 1 wherein the computing device comprises a smart phone, and wherein the smart phone comprises the image capture device.

15. The method of claim 1 further comprising, based on the determining, activating an image capture application on the computing device.

16. One or more computer-readable memory or storage devices having stored thereon computer-executable instructions operable to cause a mobile computing device to perform a method comprising:

receiving a voice command associated with an image capture application, wherein the mobile computing device comprises an image capture device controlled by the image capture application; and

responsive to the received voice command, and without receiving additional voice commands:

performing a single voice recognition event by determining that the received voice command includes an expected word or group of words; and

upon performing the single voice recognition event:

transitioning the mobile computing device from a low-power state to an active state;

activating the image capture application; and

capturing one or more digital images with the image capture device.

17. The computer-readable memory or storage devices of claim 16 wherein the one or more digital images comprise video images.

18. The computer-readable memory or storage devices of claim 16 wherein activating the image capture application comprises transitioning the image capture application from a background state to a foreground state.

19. A mobile device that includes one or more processors, a camera device, plural output devices, memory and storage media, the storage media storing computer-executable instructions for causing the mobile device to perform a method comprising:

receiving an image capture voice command when the mobile device is in a low-power state, the mobile device having stored thereon computer-executable instructions corresponding to a plurality of applications operable to be executed on the mobile device, the plurality of applications comprising an image capture application and one or more other applications;

capturing one or more digital images with the camera device through speech recognition while the mobile device remains in a device locked state by:

determining that the received image capture voice command comprises a match for an expected voice command associated with the image capture application using speech recognition; and

responsive to the determining:

transitioning from the low-power state to an active state;

activating the image capture application on the mobile device;

outputting feedback by one or more of the output devices, the feedback indicating that the received image capture voice command has been recognized; and

causing the camera device to capture one or more digital images.

20. The computer-readable memory or storage devices of claim 16 , wherein the activating the image capture application and capturing one or more digital images with the image capture device are performed while the mobile computing device remains in a device locked state.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 9, 2014
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034544/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2011
From: SARIN, RAMAN KUMAR; MATTHEWS, JOSEPH H., III; LAU, JAMES KAI YU; PARK, JAE PUM; GONZALEZ VERON, MONICA ESTELA
To: MICROSOFT CORPORATION
Reel/Frame 027231/0635 →
Continuity (1)
Related Publication 20130124207A1 · May 16, 2013