METHOD FOR REFINING CONTROL BY COMBINING EYE TRACKING AND VOICE RECOGNITION
The invention is a method for combining eye tracking and voice-recognition control technologies to increase the speed and/or accuracy of locating and selecting objects displayed on a display screen for subsequent control and operations.
1 . A method comprising:
determining an area on a display screen at which a user is gazing;
recognizing a spoken word or plurality of spoken words;
associating said spoken word or plurality of spoken words with objects displayed on said display screen;
limiting said objects displayed on said display screen to said area on said screen at which a user is gazing;
associating said objects displayed on said display screen in said area on a screen at which said user is gazing with said spoken word or plurality of spoken words.
2 . A method as in claim 1 further comprising:
determining a level of confidence in said associating said objects displayed on said display screen in said area on a screen at which said user is gazing with said spoken word or plurality of spoken words;
comparing said level of confidence with a predetermined level of confidence value and if greater than said predetermined level of confidence value, accepting the association of said spoken word or plurality of spoken words with said objects displayed on said display screen in said area on a screen which said user is gazing.
3 . A method as in claim 1 further comprising:
determining said level of confidence value based on the accuracy of the gaze coordinates, the noise of the gaze coordinates, the confidence level in the gaze coordinates, the location of the objects on the screen, or any combination thereof.
4 . A method as in claim 1 further comprising:
determining a level of probability in said associating said objects displayed on said display screen in said area on a screen at which said user is gazing with recognizing said spoken word or plurality of spoken words;
comparing said level of probability with a predetermined level of probability value and if greater than said predetermined level of probability value, accepting the association of said spoken word or plurality of spoken words with said objects displayed on said display screen in said area on a screen at which said user is gazing.
5 . A method as in claim 4 further comprising:
determining said level of probability based on the confidence level of the voice recognition, the distance from the gaze fixation to each object, the duration of the gaze fixation, the time elapsed between the gaze fixation and the emission of the voice command, or any combination thereof.
6 . A method comprising:
determining the objects present in an area on a display screen at which said user is gazing,
building a vocabulary of a voice recognition engine based on said objects,
recognizing a spoken word or plurality of spoken words using said vocabulary;
associating said objects present in the gazed area with said spoken word or plurality of spoken words.
7 . A method as in claim 6 further comprising
updating said vocabulary of said voice recognition engine on every fixation of said user.