IP Library › Granted Patent US 10,318,016
Granted Patent B2
US 10,318,016 · App. 14/294,328 · Granted Jun 11, 2019

Hands free device with directional interface

Inventors: Davide Di Censo (San Mateo, CA); Stefan Marti (Oakland, CA)
Assignee: HARMAN INTERNATIONAL INDUSTRIES, INCORPORATED
G06F3/0346G06F3/012G06F3/013G06F3/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,318,016
App. No.
14/294,328
Granted
Jun 11, 2019
Kind
B2
Abstract

Embodiments provide a non-transitory computer-readable medium containing computer program code that, when executed, performs an operation. The operation includes detecting a user action requesting an interaction with a first device and originating from a source. Additionally, embodiments determine a direction in which the source is located, relative to a current position of the first device. A response to the user action is also determined, based on a current state of the first device. Embodiments further include outputting the determined response substantially in the determined direction in which the source is located.

Claims (59)

1. A non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform the steps of:

detecting a hands-free user action requesting an interaction with a first device and originating from a source;

determining a direction in which the source is located relative to a current position of the first device;

determining a response to the hands-free user action based on a current state of the first device; and

outputting the determined response as a steerable sound beam substantially in the determined direction in which the source is located.

2. The non-transitory computer-readable medium of claim 1 , wherein the response comprises an audible response, and wherein the response is output from the first device as a steerable beam of sound oriented in the determined direction in which the source location is located.

3. The non-transitory computer-readable medium of claim 2 , wherein detecting the hands-free user action requesting the interaction with the first device comprises:

detecting, by operation of one or more sensor devices of the first device, that a user gaze is substantially oriented in the direction of the first device, comprising:

capturing one or more images that include the source;

analyzing the captured one or more images to identify a face within one of the one or more images; and

determining whether the user gaze is substantially oriented in the direction of the first device, based on the identified face within the one or more images.

4. The non-transitory computer-readable medium of claim 1 , wherein the response comprises one or more frames, and wherein outputting the determined response substantially in the determined direction in which the source is located further comprises:

determining a physical surface within a viewing range of the source; and

projecting the one or more frames onto the physical surface, using a projector device of the first device.

5. The non-transitory computer-readable medium of claim 1 , wherein the hands-free user action comprises a voice command, and the operation further comprising:

analyzing the voice command to determine a user request corresponding with the voice command; and

processing the user request to produce a result,

wherein the determined response provides at least an indication of the produced result.

6. The non-transitory computer-readable medium of claim 5 , wherein processing the user request to produce the result further comprises generating an executable query based on the user request, and wherein processing the user request to produce the result further comprises executing the executable query to produce query results, and wherein determining the response to the hands-free user action is performed using a text-to-speech synthesizer based on text associated with at least a portion of the query results.

7. The non-transitory computer-readable medium of claim 1 , wherein the direction in which the source is located is determined via at least one of a microphone and an image sensor, and the response is outputted via one or more directional loudspeakers.

8. A non-transitory computer-readable medium containing instructions that, when executed by a processor, cause the processor to perform the steps of:

detecting a triggering event comprising at least one of:

detecting a voice trigger; and

detecting a user gaze in a direction of a first device;

determining a direction to a source of the triggering event relative to a current position of the first device; and

initiating an interactive voice dialogue by outputting an audible response as a steerable sound beam substantially in the determined direction in which the source of the triggering event is located.

9. The non-transitory computer-readable medium of claim 8 , the operation further comprising:

detecting, by operation of one or more sensors of the first device, that the user gaze is oriented in the direction of the first device.

10. The non-transitory computer-readable medium of claim 8 , the operation further comprising:

analyzing the voice trigger to determine a user request corresponding with the voice trigger; and

processing the user request to produce a result,

wherein the determined audible response provides at least an indication of the produced result.

11. The non-transitory computer-readable medium of claim 10 , wherein processing the user request to produce the result further comprises generating an executable query based on the user request, and wherein processing the user request to produce the result further comprises executing the executable query to produce query results, and wherein determining an audible response to the hands-free user action is performed using a text-to-speech synthesizer based on text associated with at least a portion of the query results.

12. The non-transitory computer-readable medium of claim 8 , the operation further comprising:

capturing one or more images that include a depiction of the source of the triggering event using one or more sensor devices of the first device; and

authenticating the source of the triggering event based on a comparison of at least a portion of the captured one or more images and a predefined image.

13. The non-transitory computer-readable medium of claim 8 , the operation further comprising:

authenticating the source of the triggering event based on a comparison of the voice trigger and a predefined voice recording.

14. The non-transitory computer-readable medium of claim 8 , wherein the triggering event is detected via at least one of a microphone and an image sensor, and the audible response is outputted via one or more directional loudspeakers.

15. An apparatus, comprising:

a computer processor;

a memory containing a program that, when executed by the computer processor, causes the computer processor to perform the steps of:

detecting a hands-free user action originating from a source;

determining a direction in which the source is located relative to a current position of the apparatus;

determining a response to the hands-free user action; and

outputting the determined response as a steerable sound beam substantially in the determined direction in which the source is located.

16. The apparatus of claim 15 , wherein the response is outputted via a beam-forming speaker array.

17. The apparatus of claim 15 , wherein the response is outputted via one or more actuated directional speakers.

18. The apparatus of claim 15 , the operation further comprising:

upon detecting that a connection with a body-mounted audio output device associated with the source of the triggering event is available, outputting the determined response over the connection for playback using the body-mounted audio output device.

19. The apparatus of claim 15 , wherein determining the direction in which the source is located, relative to a current position of the apparatus, further comprises:

capturing, via a first sensor of the apparatus, one or more images;

processing the one or more images to identify the source; and

determining the direction in which the source is located based on a location of the source within the one or more images and a position of the first sensor.

20. The apparatus of claim 15 , the operation further comprising:

capturing, via one or more sensors, one or more images including the source; and

authenticating the source by comparing at least a portion of the one or more images to a predefined image of a user.

21. The apparatus of claim 20 , wherein comparing the at least a portion of the one or more images to the predefined image comprises performing facial recognition on the at least a portion of the one or more images and the predefined image of the user.

22. The apparatus of claim 20 , wherein comparing the at least a portion of the one or more images to the predefined image comprises performing a retinal scan on the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2014
From: DI CENSO, DAVIDE; MARTI, STEFAN
To: HARMAN INTERNATIONAL INDUSTRIES, INCORPORATED
Reel/Frame 033016/0099 →
Continuity (1)
Related Publication 20150346845A1 · Dec 3, 2015