Systems and methods to facilitate intent determination of a command by grouping terms based on context
Systems and methods to group terms based on context to facilitate determining intent of a command are disclosed. Exemplary implementations to train a model: obtain a set of writings within a particular knowledge domain; obtain a vector generation model that generates vectors for individual instances of the terms in the set of writings; generate a first set of vectors that represent the instances of a first term and other vectors that represent instances of the other terms of the set of writings; train the vector generation model to group the vectors of a similar context in a space of a vector space; obtain a transcript include a new term generated from user audio dictation; generate a new vector that represent the instance of the new term; obtain the space; compare the new vector with the space; utilize the new term as the first term.
1 . A system configured to group terms based on context, the system comprising:
one or more processors configured by machine-readable instructions to:
obtain a set of writings, wherein individual writings within the set of writings include terms, wherein the terms are part of a lexicon;
obtain a vector generation model that generates vectors for instances of individual ones of the terms that are included in the set of writings and that are part of the lexicon, the vectors numerically representing text of the individual ones of the terms and contexts of the instances of the individual ones of the terms;
use the vector generation model to generate the vectors, such that:
sets of the vectors are generated, wherein the vectors of individual ones of the sets represent instances of one of the terms, wherein each of the vectors numerically represent (i) text of a corresponding one of the terms, and (ii) a context in which an instance of the corresponding one of the terms is used; and
train the vector generation model to group the vectors spatially within a vector space based on other contexts similar to the ones on which generation of the vectors was based.
2 . The system of claim 1 , wherein the terms include words and/or phrases.
3 . The system of claim 1 , wherein the set of writings include transcripts of notes of a user, books, articles, theses, and/or transcribed lectures.
4 . The system of claim 1 , wherein the contexts of the instances of the individual ones of the terms involve other ones of the terms and/or syntactic relationships with the instances of the individual ones of the terms.
5 . The system of claim 1 , wherein training the vector generation model includes determining co-occurrence probability between instances of different ones of the terms and some of the terms that are part of the lexicon.
6 . The system of claim 1 , wherein training the vector generation model includes determining mutual information between instances of different ones of the terms and some of the terms that are part of the lexicon.
7 . The system of claim 1 , wherein the contexts of the instances of the individual ones of the terms includes location of a user, title of the user, and/or user behavior of the user.
8 . A system configured to utilize a term grouping model to determine intent of a command with a new spoken term, the system comprising:
one or more processors configured by machine-readable instructions to:
obtain a transcript, the transcript including instances of transcribed terms, the transcribed terms including an instance of a new transcribed term;
facilitate communication with the term grouping model, the term grouping model configured to group terms based on a particular context;
generate, via the term grouping model, a set of vectors that numerically represent text of the transcribed terms included in the transcript and contexts of the instances of the transcribed terms included in the transcript, the set of the vectors including a first vector that numerically represents text of the new transcribed term and context of the instance of the new transcribed term;
obtain, from the term grouping model, a first space, wherein the first space is included in a vector space, wherein the first space represents text of transcribed terms that are in the contexts of the instances of the transcribed terms included in the transcript, the first space including the set of the vectors except the first vector and including a second vector that represents text of an alternative term to the new transcribed term in the contexts of the instances of the transcribed terms included in the transcript;
compare the set of the vectors to the first space to determine whether the first vector is related to the first space and whether the first vector correlates to one of the vectors in the first space, such that the first vector is determined to be related to the first space and a correlation between the first vector and the second vector is determined, wherein the correlation indicates equivalency;
and
utilize, based on the correlation, the text of the alternative term by replacing the instance of the new transcribed term with the text of the alternative term.
9 . The system of claim 8 , wherein the first space is a result of training the term grouping model to group vectors that were generated based on a similar context.
10 . The system of claim 9 , wherein training the term grouping model to group the vectors that were generated based on the similar context includes obtaining a set of writings, wherein individual writings within the set of writings include terms, the terms including instances of the alternative term.
11 . A method configured to group terms based on context, the method comprising:
obtaining a set of writings, wherein individual writings within the set of writings include terms, wherein the terms are part of a lexicon;
obtaining a vector generation model that generates vectors for instances of individual ones of the terms that are included in the set of writings and that are part of the lexicon, the vectors numerically representing text of the individual ones of the terms and contexts of the instances of the individual ones of the terms;
using the vector generation model to generate the vectors, such that:
sets of the vectors are generated, wherein the vectors of individual ones of the sets represent instances of one of the terms, wherein each of the vectors numerically represent (i) text of a corresponding one of the terms, and (ii) a context in which an instance of the corresponding one of the terms is used; and
training the vector generation model to group the vectors spatially within a vector space based on other contexts similar to the ones on which generation of the vectors was based.
12 . The method of claim 11 , wherein the terms include words and/or phrases.
13 . The method of claim 11 , wherein the set of writings include transcripts of notes of a user, books, articles, theses, and/or transcribed lectures.
14 . The method of claim 11 , wherein the contexts of the instances of the individual ones of the terms involve other ones of the terms and/or syntactic relationships with the instances of the individual ones of the terms.
15 . The method of claim 11 , wherein training the vector generation model includes determining co-occurrence probability between instances of different ones of the terms and some of the terms that are part of the lexicon.
16 . The method of claim 11 , wherein training the vector generation model includes determining mutual information between instances of different ones of the terms and some of the terms that are part of the lexicon.
17 . The method of claim 11 , wherein the contexts of the instances of the individual ones of the terms includes location of a user, title of the user, and/or user behavior of the user.
18 . A method configured to utilize a term grouping model to determine intent of a command with a new spoken term, the method comprising:
obtaining a transcript, the transcript including instances of transcribed terms, the transcribed terms including an instance of a new transcribed term;
facilitating communication with the term grouping model, the term grouping model configured to group terms based on a particular context;
generating, via the term grouping model, a set of vectors that numerically represent text of the transcribed terms included in the transcript and contexts of the instances of the transcribed terms in the transcript, the set of the vectors including a first vector that numerically represents text of the new transcribed term and context of the instance of the new transcribed term;
obtaining, from the term grouping model, a first space, wherein the first space is included in a vector space, wherein the first space represents text of transcribed terms that are in the contexts of the instances of the transcribed terms included in the transcript, the first space including the set of the vectors except the first vector and including a second vector that represents text of an alternative term to the new transcribed term in the contexts of the instances of the transcribed terms included in the transcript;
comparing the set of the vectors to the first space to determine whether the first vector is related to the first space and whether the first vector correlates to one of the vectors in the first space, such that the first vector is determined to be related to the first space and a correlation between the first vector and the second vector is determined, wherein the correlation indicates equivalency,
and
utilizing, based on the correlation, the text of the alternative term by replacing the instance of the new transcribed term with the text of the alternative term.
19 . The method of claim 18 , wherein the first space is a result of training the term grouping model to group vectors that were generated based on a similar context.
20 . The method of claim 19 , wherein training the term grouping model to group the vectors that were generated based on the similar context includes obtaining a set of writings, wherein individual writings within the set of writings include terms, the terms including instances of the alternative term.