Systems, methods, and storage media for performing actions based on utterance of a command
Systems and methods for recognizing and executing spoken commands using speech recognition. Exemplary implementations may: store actionable phrases; obtain audio information representing sound captured by a mobile client computing platform associated with a user; detect any spoken instances of a predetermined keyword present in the sound represented by the audio information; perform speech recognition on the sound represented by the audio information; identify an utterance of an individual actionable phrase in speech temporally adjacent to the spoken instance of the predetermined keyword that is present in the sound represented by the audio information; perform natural language processing to identify an individual command uttered temporally adjacent to the spoken instance of the predetermined keyword that is present in the sound represented by the audio information; and effectuate performance of instructions corresponding to the command.
1 . A system configured to recognize and execute spoken commands using speech recognition, the system comprising:
electronic storage media configured to store predetermined actionable phrases, individual ones of the predetermined actionable phrases correlating to individual known commands, wherein the known commands are known documentation commands; and
one or more processors configured by machine-readable instructions to:
receive, via a mobile client computing platform, a word or a phrase to define as a predetermined keyword that indicates utterance of one or more commands temporally adjacent to utterance of the word or the phrase;
obtain audio information representing sound captured by the mobile client computing platform associated with a user;
convert the sound to digital signals, wherein the digital signals include noise;
filter the digital signals of the noise;
cause an audio encoder to encode the digital signals to an audio file according to an audio file format;
detect one or more spoken occurrences of the predetermined keyword present in the sound by performing word recognition;
responsive to determination of the one or more spoken occurrences of the predetermined keyword:
determine whether the speech includes one or more utterances of one or more of the predetermined actionable phrases temporally adjacent to a spoken occurrence;
responsive to determining that the speech includes an utterance of a predetermined actionable phrase, obtain a known command correlated to the predetermined actionable phrase; and
responsive to determining that the speech lacks one or more utterances of one or more of the predetermined actionable phrases, perform natural language processing to identify a new command uttered temporally adjacent to the spoken occurrence; and
effectuate performance of instructions corresponding to the new command or the known command.
2 . The system of claim 1 , wherein the one or more processors are further configured by the machine-readable instructions to:
transmit the instructions to the mobile client computing platform.
3 . The system of claim 1 , wherein the mobile client computing platform includes one or more of a microphone, the audio encoder, a speaker, or a processor.
4 . The system of claim 1 , wherein the known commands include one or more of taking a note, opening a file, reciting information, setting a calendar date, sending information, or sending requests.
5 . The system of claim 1 , wherein the predetermined actionable phrases are frequently used phrases.
6 . The system of claim 1 , wherein the one or more processors are further configured by the machine-readable instructions to obtain requests to add one or more additional predetermined actionable phrases to the electronic storage, and effectuate storage of the one or more additional predetermined actionable phrases to the electronic storage.
7 . The system of claim 1 , wherein the one or more processors are further configured by the machine-readable instructions to obtain requests to alter one or more of the predetermined actionable phrases.
8 . The system of claim 1 , wherein the predetermined keyword is the word.
9 . The system of claim 1 , wherein the predetermined keyword is the phrase.
10 . A method configured to recognize and execute spoken commands using speech recognition, the method comprising:
storing predetermined actionable phrases, individual ones of the predetermined actionable phrases correlating to individual known commands, wherein the known commands are known documentation commands;
receiving, via a mobile client computing platform, a word or a phrase to define as a predetermined keyword that indicates utterance of one or more commands temporally adjacent to utterance of the word or the phrase;
obtaining audio information representing sound captured by the mobile client computing platform associated with a user;
converting the sound to digital signals, wherein the digital signals include noise;
filtering the digital signals of the noise;
causing an audio encoder to encode the digital signals to an audio file according to an audio file format;
detecting one or more spoken occurrences of the predetermined keyword present in the sound by performing word recognition;
responsive to determination of the one or more spoken occurrences of the predetermined keyword:
determining whether the speech includes one or more utterances of one or more of the predetermined actionable phrases temporally adjacent to the spoken occurrence;
responsive to determining that the speech includes an utterance of a predetermined actionable phrase, obtaining a known command correlated to the predetermined actionable phrase; and
responsive to determining that the speech lacks one or more utterances of one or more of the predetermined actionable phrases, performing natural language processing to identify a new command uttered temporally adjacent to the spoken occurrence; and
effectuating performance of instructions corresponding to the new command or the known command.
11 . The method of claim 10 , further comprising:
transmitting the instructions to the mobile client computing platform.
12 . The method of claim 10 , wherein the mobile client computing platform includes one or more of a microphone, the audio encoder, a speaker, or a processor.
13 . The method of claim 10 , wherein the known commands include one or more of taking a note, opening a file, reciting information, setting a calendar date, sending information, or sending requests.
14 . The method of claim 10 , wherein the converting, the filtering, and the causing are performed by the mobile client computing platform.
15 . The method of claim 10 , wherein the predetermined actionable phrases are frequently used phrases.
16 . The method of claim 10 , further comprising obtaining requests to add one or more additional predetermined actionable phrases to the electronic storage, and effectuating storage of the one or more additional predetermined actionable phrases to the electronic storage.
17 . The method of claim 10 , further comprising obtaining requests to alter one or more of the predetermined actionable phrases.
18 . The method of claim 10 , wherein the predetermined keyword is the word.
19 . The method of claim 10 , wherein the predetermined keyword is the phrase.
20 . A system configured to recognize and execute spoken commands using speech recognition, the system comprising:
electronic storage media configured to store a predetermined actionable phrase corresponding to a known command, and a predetermined keyword;
one or more processors configured by machine-readable instructions to:
convert sound to a digital signal;
cause an audio encoder to encode the digital signal to an audio file;
detect a spoken occurrence of the predetermined keyword in the sound;
responsive to detection of the spoken occurrence:
determine whether the speech includes an utterance of the predetermined actionable phrase temporally adjacent to the spoken occurrence;
responsive to determining that the speech includes the utterance of the predetermined actionable phrase, identify the known command corresponding to the predetermined actionable phrase; and
responsive to determining that the speech lacks the utterance of the predetermined actionable phrase, perform natural language processing to identify a new command uttered temporally adjacent to the spoken occurrence; and
effectuate performance of instructions corresponding to the new command or the known command.