IP Library Granted Patent US 10,048,748
Granted Patent B2
US 10,048,748 · App. 14/077,821 · Granted Aug 14, 2018

Audio-visual interaction with user devices

Inventors: Venkatraman Sridharan (Sunnyvale, CA); Robert Jacob Kirk (Sunnyvale, CA)
Assignee: EXCALIBUR IP, LLC
G06F3/012G06F3/013G06F3/017G06F3/0304G06F3/167G10L15/22G06F2203/0381G10L21/10G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,048,748
App. No.
14/077,821
Filed
Nov 12, 2013
Granted
Aug 14, 2018
Kind
B2
Art Unit
2621
USPC
345/158
Abstract

A user device is enabled by an audio-visual assistant for audio-visual interaction with a user. The audio-visual assistant enables the user device to track the user's eyes and face to determine objects on the screen that the user is currently observing. Various tasks can be executed on the objects based on further input provided by the user. The user can provide further inputs via facial gestures, voice or combinations thereof for executing the various tasks.

Claims (69)

1. A method comprising:

accessing, by a processor of a hand-held user device comprising a display screen, an audio-visual assistant being executed on the hand-held user device, the audio-visual assistant comprising an optical detection function executable on the hand-held user device for receiving visual input of a user on the hand-held user device and an audio detection function executable on the hand-held user device for receiving voice input that enables audio interaction by the user with the hand-held user device;

identifying, by the processor of the hand-held device via the audio-visual assistant, a user's face;

receiving, by the processor of the hand-held device, data associated with the identified user's face, the data including eye tracking data for tracking a user's gaze;

mapping, by the processor of the hand-held device, based on the received data, the user's gaze with respect to different portions of the display screen for determining a portion of the display screen the user is observing;

determining, by the processor of the hand-held device via the audio-visual assistant, the portion of the display screen currently being observed by the user based on the received visual input;

identifying, by the processor of the hand-held device, a user interface element for executing tasks, the user interface element displayed within the determined portion of the display screen and displayed proximate a visually observable cursor that is separate from the identified user interface element and that tracks a movement of the user's gaze;

receiving, by the processor of the hand-held device, a voice input as a command from the user to control the user interface element identified by the received visual input;

selecting, by the processor of the hand-held device from a plurality of tasks, a task to be associated with the identified user interface element, the plurality of tasks corresponding to a respective plurality of commands associated with the audio-visual assistant; and

executing, by the processor of the hand-held device based at least on the received command from the user, the selected task associated with the identified user interface element.

2. The method of claim 1 , wherein the audio-visual assistant is configured to receive a first subset of the plurality of commands as input via the optical detection function.

3. The method of claim 1 , wherein the audio-visual assistant is configured to receive a second subset of the plurality of commands as input via the audio detection function.

4. The method of claim 1 , wherein the audio-visual assistant is configured to receive a third subset of the plurality of commands as a combination of inputs via the optical detection function and via the audio detection function.

5. The method of claim 4 , wherein the command is comprised in the third subset such that the identified user interface element is selected in response to the input received via the optical detection function and the task is selected in response to the input received via the audio detection function.

6. The method of claim 4 , further comprising:

converting, by the processor, the input to the audio detection function to text.

7. The method of claim 6 , further comprising:

identifying, by the processor, a command in program code that maps to the text.

8. The method of claim 7 , executing the selected task associated with the user interface element further comprising:

executing, by the processor, a code block associated with the command.

9. The method of claim 1 , determining the portion of the display screen further comprising:

storing, by the processor, information regarding the positions in a data storage of the hand-held user device; and

displaying, by the processor, a cursor upon conclusion of the calibration on the portion of the display screen currently being observed by the user.

10. The method of claim 1 , wherein the user interface element is associated with a software application being executed on the hand-held user device.

11. A hand-held user device comprising:

at least one processor;

a display screen;

a non-transitory computer-readable storage medium for tangibly storing thereon program logic for execution by the processor, the program logic comprising:

accessing logic, executed by the processor, for accessing an audio-visual assistant on the hand-held user device, the audio-visual assistant comprising an optical detection function for receiving visual input of a user the hand-held user device and an audio detection function for receiving voice input that enables audio interaction by the user with the hand-held user device;

identifying logic, executed by the processor, for identifying via the audio-visual assistant, a user's face;

receiving logic, executed by the processor, for receiving data associated with the identified user's face, the data including eye tracking data for tracking a user's gaze;

mapping logic, executed by the processor, for mapping, based on the received data, the user's gaze with respect to different portions of the display screen for determining a portion of the display screen the user is observing;

determining logic, executed by the processor for determining via the audio-visual assistant, the portion of the display screen currently being observed by the user based on the received visual input and displayed proximate a visually observable separate cursor that tracks a movement of the user's gaze;

identifying logic, executed by the processor, for identifying, a user interface element for executing tasks, the user interface element displayed within the determined portion of the display screen and displayed proximate a visually observable cursor that is separate from the identified user interface element and that tracks a movement of the user's gaze;

receiving logic, executed by the processor, for receiving a voice input as a command from the user to control the user interface element identified by the received visual input;

selecting logic, executed by the processor, for selecting from a plurality of tasks, a task to be associated with the identified user interface element, the plurality of tasks corresponding to a respective plurality of commands associated with the audio-visual assistant; and

logic for executing, by the processor based at least on the received command from the user, the selected task associated with the identified user interface element.

12. The hand-held user device of claim 1 , the receiving logic further comprising:

visual input receiving logic, executed by the processor for receiving a first subset of the plurality of commands as input via the optical detection function.

13. The hand-held user device of claim 1 , the receiving logic further comprising:

audio input receiving logic, executed by the processor for receiving a second subset of the plurality of commands as input via the audio detection function.

14. The hand-held user device of claim 1 , the receiving logic further comprising:

combination input receiving logic, executed by the processor, for receiving a third subset of the plurality of commands as a combination of inputs via the optical detection function and the audio detection function.

15. The hand-held user device of claim 14 , further comprising:

voice converting logic, executed by the processor, for converting the input received by the audio detection function to text; and

command identifying logic, executed by the processor, for identifying a command in program code that maps to the text.

16. The hand-held user device of claim 1 , wherein the calibrating logic further comprises:

storing logic, executed by the processor, for storing information regarding the positions in a data storage of the hand-held user device; and

cursor displaying logic, executed by the processor, for displaying a cursor upon conclusion of the calibration on the portion of the display screen currently being observed by the user.

17. A non-transitory computer readable storage medium tangibly encoded with computer-executable instructions, that when executed by a processor of a hand-held user device, cause the hand-held user device to perform a method comprising:

accessing an audio-visual assistant being executed on the hand-held user device comprising a display screen, the audio-visual assistant comprising an optical detection function for receiving visual input of a user on the hand-held user device and an audio detection function for receiving voice input that enables audio interaction by the user with the hand-held user device;

identifying via the audio-visual assistant, a user's face;

receiving data associated with the identified user's face, the data including eye tracking data for tracking a user's gaze;

mapping, based on the received data, the user's gaze with respect to different portions of the display screen for determining a portion of the display screen the user is observing;

determining via the audio-visual assistant being executed on the hand-held user device, the portion of a display screen currently being observed by the user based on the received visual input;

identifying a user interface element for executing tasks, the user interface element displayed within the determined portion of the display screen and displayed proximate a visually observable cursor that is separate from the identified user interface element and that tracks a movement of the user's gaze;

receiving a voice input as a command from the user to control the user interface element identified by the received visual input;

selecting from a plurality of tasks, a task to be associated with the identified user interface element, the plurality of tasks corresponding to a respective plurality of commands associated with the audio-visual assistant; and

executing based at least on the received command from the user, the selected task associated with the identified user interface element.

18. The non-transitory computer readable storage medium of claim 17 , wherein the audio-visual assistant is configured to receive a subset of the plurality of commands as input via the optical detection function.

19. The non-transitory computer readable storage medium of claim 17 , wherein the audio-visual assistant is configured to receive a subset of the plurality of commands as input via the audio detection function.

20. The non-transitory computer readable storage medium of claim 17 , wherein the audio-visual assistant is configured to receive a subset of the plurality of commands as a combination of input via the optical detection function and input via the audio detection function.

21. The non-transitory computer readable storage medium of claim 20 , further comprising processor-executable instructions for:

converting input received by the audio detection function to text; and

executing a code block associated with the command.

22. The non-transitory computer readable storage medium of claim 17 , further comprising processor-executable instructions for:

storing information regarding the positions in a data storage of the hand-held user device.

23. The non-transitory computer readable storage medium of claim 22 , further comprising processor-executable instructions for:

displaying a cursor on the portion of the display screen currently being observed by the user upon conclusion of the calibration.

Assignments (9)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNOR NAME PREVIOUSLY RECORDED AT REEL: 052853 FRAME: 0153. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 29, 2021
From: R2 SOLUTIONS LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 056832/0001 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED ON REEL 053654 FRAME 0254. ASSIGNOR(S) HEREBY CONFIRMS THE RELEASE OF SECURITY INTEREST GRANTED PURSUANT TO THE PATENT SECURITY AGREEMENT PREVIOUSLY RECORDED. Recorded Dec 30, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: R2 SOLUTIONS LLC
Reel/Frame 054981/0377 →
RELEASE OF SECURITY INTEREST IN PATENTS Recorded Jul 8, 2020
From: STARBOARD VALUE INTERMEDIATE FUND LP
To: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
Reel/Frame 053654/0254 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 25, 2020
From: EXCALIBUR IP, LLC
To: R2 SOLUTIONS LLC
Reel/Frame 053459/0059 →
PATENT SECURITY AGREEMENT Recorded Jun 5, 2020
From: ACACIA RESEARCH GROUP LLC; AMERICAN VEHICULAR SCIENCES LLC; BONUTTI SKELETAL INNOVATIONS LLC; CELLULAR COMMUNICATIONS EQUIPMENT LLC; INNOVATIVE DISPLAY TECHNOLOGIES LLC; LIFEPORT SCIENCES LLC; LIMESTONE MEMORY SYSTEMS LLC; MERTON ACQUISITION HOLDCO LLC; MOBILE ENHANCEMENT SOLUTIONS LLC; MONARCH NETWORKING SOLUTIONS LLC; NEXUS DISPLAY TECHNOLOGIES LLC; PARTHENON UNIFIED MEMORY ARCHITECTURE LLC; R2 SOLUTIONS LLC; SAINT LAWRENCE COMMUNICATIONS LLC; STINGRAY IP SOLUTIONS LLC; SUPER INTERCONNECT TECHNOLOGIES LLC; TELECONFERENCE SYSTEMS LLC; UNIFICATION TECHNOLOGIES LLC
To: STARBOARD VALUE INTERMEDIATE FUND LP, AS COLLATERAL AGENT
Reel/Frame 052853/0153 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038950/0592 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2016
From: EXCALIBUR IP, LLC
To: YAHOO! INC.
Reel/Frame 038951/0295 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2016
From: YAHOO! INC.
To: EXCALIBUR IP, LLC
Reel/Frame 038383/0466 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2013
From: SRIDHARAN, VENKATRAMAN; KIRK, ROBERT JACOB
To: YAHOO! INC.
Reel/Frame 031585/0433 →
Continuity (1)
Related Publication 20150130716A1 · May 14, 2015
Cited By (24)
US 12,197,817 US 12,200,297 US 12,211,502 US 12,219,314 US 12,236,952 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,360,747 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,422,933 US 12,423,917 US 12,436,619 US 12,477,470 US 12,504,863 US 12,567,415 US 12,608,171 US 12,619,452 US 12,620,179