IP Library Granted Patent US 11,908,475
Granted Patent B1
US 11,908,475 · App. 18/167,158 · Granted Feb 20, 2024

Systems, methods and non-transitory computer readable media for human interface device accessibility

Inventor: Alexander Dunn (Wilmington, MA)
Assignee: CEPHABLE INC.
G10L15/22G10L15/187G10L15/25G10L2015/088G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,908,475
App. No.
18/167,158
Granted
Feb 20, 2024
Kind
B1
Abstract

A method, system, and non-transitory computer readable media for converting input from a user into a human interface device (HID) output to cause a corresponding action at an mapped device includes receiving one or more user input from a user at an input device, analyzing the user input selecting a command from a command profile that maps at least one of the received user inputs to one or more mapped tasks, executing the one or more mapped tasks associated with the selected command, and causing one or more corresponding actions at one or more mapped devices associated with the one or more mapped tasks.

Claims (168)

1. A method comprising:

receiving, by a computing device, audio data, wherein the audio data comprises speech of a user and other audio;

isolating, by the computing device, speech of the user from the audio data;

identifying, by the computing device, at least one of a one or more vowel sound and a diphthong in the isolated speech of the user;

analyzing, by the computing device, the isolated speech of the user and generating a plurality of predictions, wherein each prediction of the plurality of predictions comprises at least one of a word and a phrase in the isolated speech of the user;

performing, by the computing device, a post processing on the plurality of predictions and generating a plurality of processed predictions based on the plurality of predictions, wherein performing the post processing on the plurality of predictions and generating a plurality of processed predictions based on the plurality of predictions comprises:

comparing, by the computing device, at least one word and phrase of the plurality of predictions to one or more relative words and phrases in a word-phrase graph;

calculating, by the computing device, a correction score for at least one prediction of the plurality of predictions and at least one relative word and phrase of the one or more relative words and phrases;

determining, by the computing device, whether the correction score exceeds a threshold correction score; and

generating, by the computing device, a processed prediction,

wherein the processed prediction comprises the at least one relative word and phrase used to calculate the correction score if the correction score exceeds the threshold correction score, and

wherein the processed prediction comprises the at least one prediction used to calculate the correction score if the correction score does not exceed the threshold correction score;

selecting, by the computing device, a first command in a command profile, the first command defining a user input comprising at least one of the vowel sound and the diphthong identified in the isolated speech of the user and the one or more processed predictions of the plurality of processed predictions, wherein a command comprises at least one user input mapped to at least one mapped task associated with a mapped device and to be executed by the computing device;

executing, by the computing device, the at least one mapped task of the first command; and

causing, by the computing device, at least one corresponding action at the mapped device associated with the one or more mapped tasks of the first command

wherein a correction score is calculated according to:

R

W

-

D

L

W

-

D

+

T

F

H

wherein RW represents a weight of the relative word or phrase,

LW represents a weight of the word or phrase of the at least one prediction,

D represents a phonetic distance between the relative word and the word or phrase of the at least one prediction,

F represents a total number of times the word or phrase of the at least one prediction has been detected in previous iterations of post processing,

T represents a number of transitions made from the word or phrase of the at least one prediction to the relative word or phrase, and

H represents the threshold correction value.

2. The method of claim 1 , wherein isolating the speech of the user from the audio data comprises:

identifying, by the computing device, a plurality of characteristics of the audio data, wherein characteristics of the audio data comprise at least one of volume, clarity, and consistency; and

comparing, by the computing device, a characteristic of the plurality of characteristics to one or more other characteristics of the plurality of characteristics of the audio data.

3. The method of claim 1 , wherein the word-phrase graph comprises:

a plurality of word-phrase nodes corresponding to each relative word or phrase of the one or more relative words or phrases, wherein each word-phrase node comprises:

a phonetic representation of the corresponding relative word or phrase;

one or more phonetic components of the corresponding relative word or phrase;

a syllable count of the corresponding relative word or phrase;

a weighting value;

a session frequency reflecting a number of times the corresponding relative word or phrase has been detected in a particular session;

a total frequency reflecting a number of times the corresponding relative word or phrase has been detected in previous iterations of the post processing; and

an average frequency reflecting an average number of times the corresponding relative word or phrase is detected in the particular session; and

a plurality of edges connecting pairs of word-phrase nodes of the plurality of word-phrase nodes, wherein each edge of the plurality of edges comprises:

a number of phonetic parts shared by each word-phrase node of the pair of word-phrase nodes connected by each edge, and

a number of transitions.

4. The method of claim 1 , wherein performing the post processing further comprises:

determining, by the computing device, that at least one prediction of the one or more predictions begins with a dictation phrase; and

detecting, by the computing device, a second command corresponding to the dictation phrase.

5. The method of claim 1 , wherein selecting the command comprises:

applying, by the computing device and to each processed prediction, one of:

a start-of-chain with removal mode,

a mid-chain with removal mode, and

a mid-chain with no removal mode.

6. The method of claim 1 , wherein executing the one or more mapped tasks of the first command comprises:

storing, by the computing device, the one or more mapped tasks of the first command in a Human Interface Device (HID) state and a HID queue, wherein the HID state comprises one or more mapped tasks being executed by the computing device, and wherein the HID queue comprises one or more mapped tasks to be executed by the computing device after the one or more mapped tasks in the HID state are executed;

storing, by the computing device, one or more HID packets in a HID report after executing a mapped task of the one or more mapped tasks of the first command; and

communicating, by the computing device, the HID report to the mapped device.

7. The method of claim 1 , wherein executing one or more mapped tasks of the first command comprises executing, by the computing device, a remote task of a third-party system, wherein executing a remote task at a third-party system comprises:

identifying, by the computing device, one or more remote task access tokens associated with the first command;

communicating, by the computing device, a remote task request and the remote task access token to the third-party system; and

receiving, by the computing device, access to the third-party system.

8. The method of claim 7 , wherein the mapped device is associated with the third-party system.

9. One or more non-transitory computer readable media comprising instructions that, when executed by a computing system, cause the computing system to:

receive audio data, wherein the audio data comprises speech of a user and other audio;

isolate the speech of the user from the audio data;

identify at least one of one or more vowel sounds and a diphthong in the isolated speech of the user;

analyze the isolated speech of the user and generate a plurality of predictions, wherein each prediction of the plurality of predictions comprises at least one of a word and a phrase in the isolated speech of the user;

perform a post processing on the plurality of predictions and generate a plurality of processed predictions based on the plurality of predictions, wherein the instructions that, when executed by the computing system, cause the computing system to perform the post processing on the plurality of predictions and generate the processed prediction comprise instructions that cause the computing system to:

compare the at least one word and phrase of the plurality of predictions to one or more relative words and phrases in a word-phrase graph;

calculate a correction score for at least one prediction of the plurality of predictions and at least one relative word and phrase of the one or more relative words and phrases;

determine whether the correction score exceeds a threshold correction score; and

generating a processed prediction,

wherein the processed prediction comprises the at least one relative word and phrase used to calculate the correction score if the correction score exceeds the threshold correction score, and

wherein the processed prediction comprises the at least one at least one prediction used to calculate the correction score if the correction score does not exceed the threshold correction score;

select a first command in a command profile, the first command defining a user input comprising at least one of the one or more vowel sounds and the diphthong identified in the isolated speech of the user and the one or more processed predictions of the plurality of processed predictions, wherein a command comprises a user input mapped to at least one mapped task associated with a mapped device;

execute the at least one mapped task of the first command; and

cause at least one corresponding action at the mapped device associated with the at least one mapped tasks of the first command,

wherein the word-phrase graph comprises:

a plurality of word-phrase nodes corresponding to each relative word or phrase of the one or more relative words or phrases, wherein each word-phrase node comprises:

a phonetic representation of the corresponding relative word or phrase;

one or more phonetic components of the corresponding relative word or phrase;

a syllable count of the corresponding relative word or phrase;

a weighting value;

a session frequency reflecting a number of times the corresponding relative word or phrase has been detected in a particular session;

a total frequency reflecting a number of times the corresponding relative word or phrase has been detected in previous interactions of the post processing; and

an average frequency reflecting an average number of times the corresponding relative word or phrase is detected in the particular session; and

a plurality of edges connecting pairs of word-phrase nodes of the plurality of word-phrase nodes, wherein each edge of the plurality of edges comprises:

a number of phonetic parts shared by a pair of each word-phrase node of the pair of word-phrase nodes connected by each edge, and a number of transitions.

10. The one or more non-transitory media of claim 9 , wherein the instructions that, when executed by the computing system, cause the computing system to isolate the speech of the user comprises instructions that cause the computing system to:

identify a plurality of characteristics of the audio data, wherein characteristics of the audio data comprise at least one of volume, clarity, and consistency; and

compare a characteristic of the plurality of characteristics to one or more other characteristics of the plurality of characteristics of the audio data.

11. The one or more non-transitory media of claim 9 , wherein the correction score is calculated according to:

R

W

-

D

L

W

-

D

+

T

F

H

wherein RW represents a weight of the relative word or phrase,

LW represents a weight of the word or phrase of the at least one prediction,

D represents a phonetic distance between the relative word and the word or phrase of the at least one prediction,

F represents a total number of times the word or phrase of the at least one prediction has been detected in previous iterations of the post processing,

T represents a number of transitions made from the word or phrase of the at least one prediction to the relative word or phrase, and

H represents the threshold correction score.

12. The one or more non-transitory media of claim 9 , wherein the instructions that, when executed by the computing system, cause the computing system to perform a post processing further comprise instructions that cause the computing system to:

determine that the at least one prediction of the one or more predictions begins with a dictation phrase; and

detect a second command corresponding to the dictation phrase.

13. The one or more non-transitory media of claim 9 , wherein instructions that, when executed by the computing system, cause the computing system to select a command comprise instructions that cause the computing system to:

apply, to each processed prediction, one of:

a start-of-chain with removal mode, a mid-chain with removal mode, or a mid-chain with no removal mode.

14. The one or more non-transitory media of claim 9 , wherein instructions that, when executed by the computing system, cause the computing system to execute one or more mapped tasks of the first command comprises instructions that cause the computing system to execute a remote task of a third-party system, wherein instructions that cause the computing system to execute a remote task at a third-party system comprise instructions to:

identify one or more remote task access tokens associated with the first command;

communicate a remote task request and the remote task access token to the third-party system; and

receive access to the third-party system.

15. A system comprising:

one or more processors; and

at least one memory operatively coupled to the processor and storing instructions that, when executed by at least one processor of the one or more processors, cause the system to:

receive audio data, wherein the audio data comprises speech of a user and other audio;

isolate the speech of the user from the audio data;

identify at least one of one or more vowel sounds and a diphthong in the isolated speech of the user;

analyze the isolated speech of the user and generate a plurality of predictions, wherein each prediction of the plurality of predictions comprises at least one of a word and a phrase in the isolated speech of the user;

perform a post processing on the plurality of predictions and generate a plurality of processed predictions based on the plurality of predictions;

calculate a correction score for at least one prediction of the plurality of predictions and at least one relative word and phrase of the one or more relative words and phrases;

select a first command in a command profile, the first command defining a user input comprising at least one of the one or more vowel sounds and the diphthong identified in the isolated speech of the user and the one or more processed predictions of the plurality of processed predictions, wherein a command comprises a user input mapped to at least one mapped task associated with a mapped device;

execute the at least one mapped task of the first command; and

cause at least one corresponding action at the mapped device associated with the at least one mapped tasks of the first command,

wherein the correction score is calculated according to:

R

W

-

D

L

W

-

D

+

T

F

H

wherein RW represents a weight of the relative word or phrase,

LW represents a weight of the word or phrase of the at least one prediction,

D represents a phonetic distance between the relative word and the word or phrase of the at least one prediction,

F represents a total number of times the word or phrase of the at least one prediction has been detected in previous iterations of the post processing,

T represents a number of transitions made from the word or phrase of the at least one prediction to the relative word or phrase, and

H represents the threshold correction score.

16. The system of claim 15 , wherein instructions that, when executed by at least one processor of the one or more processors, cause the system to execute one or more mapped tasks of the first command comprise instructions that cause the at least one processor to execute a remote task of a third-party system, wherein instructions that cause the at least one processor to execute a remote task at a third-party system comprise instructions to:

identify one or more remote task access tokens associated with the first command;

communicate a remote task request and the remote task access token to the third-party system; and

receive access to the third-party system.

Assignments (3)
CHANGE OF NAME Recorded Dec 13, 2023
From: ENABLED PLAY, INC.
To: CEPHABLE INC.
Reel/Frame 065990/0535 →
MERGER Recorded Jun 29, 2023
From: ENABLED PLAY LLC
To: ENABLED PLAY, INC.
Reel/Frame 064108/0449 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2023
From: DUNN, ALEXANDER
To: ENABLED PLAY LLC
Reel/Frame 062649/0576 →