System and methods for prospective legal research
The present invention is directed towards systems and methods for conducting prospective legal research, which comprises receiving an initiated user question at a graphical user interface comprising one or more search terms and performing query expansion on the received search query. One or more documents that are responsive to the expanded search query are then identified, and from the set of responsive documents, a subset of documents that reference future development are then identified. The one or more responsive documents that reference future development are grouped into one or more document clusters and a topic is identified for each of the one or more document clusters. The one or more document clusters and the associated topics are then presented at the graphical user interface.
1 . A computer-implemented method for conducting prospective legal research comprising:
preprocessing a set of documents with a functional weighting scheme placing an increased weight on one or more documents within a recent time interval;
receiving an initiated user question at a graphical user interface comprising one or more search terms;
performing query expansion on the received search query;
identifying one or more documents from the set of pre-processed documents that are responsive to the expanded search query;
identifying a plurality of responsive documents that reference future development by determining whether the one or more responsive documents contains at least one of a future term and a relevant feature, wherein a relevant feature comprises at least one of a prospective legal phrase, a rare temporal phrase, an entity tag and a part of speech tag, wherein a first of the plurality of responsive documents that reference future development contains a future date comprising an explicit future date, a second of the plurality of responsive documents that reference future development contains a future date comprising a future date phrase and a third of the plurality of responsive documents that reference future development contains a future date comprising a future date range;
calculating a cosine similarity using term-frequency-inverse document frequency vectorization for the one or more responsive documents;
grouping, using the calculated cosine similarity, the one or more responsive documents that reference future development into one or more document clusters;
identifying a topic for each of the one or more document clusters from the calculated cosine similarity; and
determining an optimal number of topics for latent Dirichlet allocation (LDA) modeling;
setting the optimal number of topics to between four and six for LDA modeling;
presenting the one or more document clusters and the associated topics at the graphical user interface, wherein the graphic user interface includes therein a timeline depicting future dates relating to the at least one of the one or more response documents that reference future developments, including at least one solid interval and at least one fuzzy interval on the timeline each of the intervals corresponding to at least one of the explicit future date of the first of the plurality of responsive documents that reference future developments the future date phrase of the second of the plurality of responsive documents that reference future developments, and the future date range of the third of the plurality of responsive documents that reference future developments.
2 . The computer-implemented method of claim 1 wherein presenting at the graphical user interface includes presenting a representation of a similarity measure to a selected document.
3 . The computer-implemented method of claim 1 , comprising
setting a maximum limit on a number of documents within the set of documents for preprocessing;
preprocessing the one or more responsive documents with natural language processing, wherein natural language processing includes stemming and punctuation removal;
performing LDA topic modelling on the one or more documents preprocessed with natural language processing; and
selecting one or more dominant terms within each topic;
determining a strong weighting within the one or more documents after LDA topic modeling and wherein a future term comprises at least one of a modal verb, a common prospective term and an uncommon prospective phrase.
4 . The computer-implemented method of claim 1 wherein grouping the one or more responsive documents that reference future development into one or more document clusters is completed based on at least one of matching keywords, matching subjects, matching entities, matching unstructured text, matching authorship, matching quotes, matching dates, related dates, volume of documents, tagging relationships and direct connections between documents.
5 . The computer-implemented method of claim 1 wherein the timeline has a first range of dates and a scale having a second range of dates smaller than the first range of dates, the graphic user interface further including a second timeline having a start and end corresponding to the second range of dates.
6 . The computer-implemented method of claim 2 wherein the representation of the similarity measure to the selected document is based on the calculated cosine similarity using term-frequency-inverse document frequency vectorization.
7 . The computer-implemented method of claim 5 wherein the second timeline depicts the plurality of responsive documents that reference future developments sorted based on whether documents are official, unofficial, or legislative.
8 . Non-computer readable media comprising program code stored thereon for execution by a programmable processor to perform a method for conducting prospective legal research comprising:
program code for preprocessing a set of documents with a functional weighting scheme placing an increased weight on one or more documents within a recent time interval;
program code for receiving an initiated user question at a graphical user interface comprising one or more search terms;
program code for performing query expansion on the received search query;
program code for identifying one or more documents that are responsive to the expanded search query;
program code for determining an optimal number of topics for latent Dirichlet allocation (LDA) modeling;
program code for setting the optimal number of topics to between four and six for LDA modeling;
program code for identifying a plurality of responsive documents that reference future development by determining whether the one or more responsive documents contains at least one of a future term and a relevant feature, wherein a relevant feature comprises at least one of a prospective legal phrase, a rare temporal phrase, an entity tag and a part of speech tag, wherein a first of the plurality of responsive documents that reference future development contains a future date comprising an explicit future date, a second of the plurality of responsive documents that reference future development contains a future date comprising a future date phrase, and a third of the plurality of responsive documents that reference future development contains a future date comprising a future date range;
program code for calculating a cosine similarity using term-frequency-inverse document frequency vectorization for the one or more responsive documents;
program code for grouping, using the calculated cosine similarity, the one or more responsive documents that reference future development into one or more document clusters;
program code for identifying a topic for each of the one or more document clusters from the calculated cosine similarity; and
program code for presenting the one or more document clusters and the associated topics at the graphical user interface, wherein the graphic user interface includes therein a timeline depicting future dates relating to the at least one of the one or more response documents that reference future developments, including at least one solid interval and at least one fuzzy interval on the timeline, each of the intervals corresponding to at least one of the explicit future date of the first of the plurality of responsive documents that reference future developments the future date phrase of the second of the plurality of responsive documents that reference future
developments, and the future date range of the third of the plurality of responsive documents that reference future developments.
9 . The computer readable media of claim 8 wherein the program code for presenting at the graphical user interface further includes program code for presenting a representation of a similarity measure to a selected document.
10 . The computer readable media of claim 8 further comprising
program code for setting a maximum limit on a number of documents within the set of documents for preprocessing;
program code for preprocessing the one or more responsive documents with natural language processing, wherein natural language processing includes stemming and punctuation removal;
program code for performing LDA topic modelling on the one or more documents preprocessed with natural language processing; and
program code for selecting one or more dominant terms within each topic;
program code for determining a strong weighting within the one or more documents after LDA topic modeling and wherein a future term comprises at least one of a modal verb, a common prospective term and an uncommon prospective phrase.
11 . The computer readable media of claim 9 wherein the representation of the similarity measure to the selected document is based on the calculated cosine similarity using term-frequency-inverse document frequency vectorization.
12 . The computer readable media of claim 9 wherein the program code for grouping the one or more responsive documents that reference future development into one or more document clusters is completed based on at least one of matching keywords, matching subjects, matching entities, matching unstructured text, matching authorship, matching quotes, matching dates, related dates, volume of documents, tagging relationships and direct connections between documents.
13 . A system for conducting prospective legal research comprising:
a server including a processor configured to:
preprocess a set of documents with a functional weighting scheme placing an increased weight on one or more documents within a recent time interval;
receive an initiated user question at a graphical user interface comprising one or more search terms;
perform query expansion on the received search query;
identify one or more documents that are responsive to the expanded search query;
determine an optimal number of topics for latent Dirichlet allocation (LDA) modeling;
set the optimal number of topics to between four and six for LDA modeling;
identify a plurality of responsive documents that reference future development by determining whether the one or more responsive documents contains at least one of a future term and a relevant feature, wherein a relevant feature comprises at least one of a prospective legal phrase, a rare temporal phrase, an entity tag and a part of speech tag, wherein a first of the plurality of responsive documents that reference future development contains a future date comprising an explicit future date, a second of the plurality of responsive documents that reference future development contains a future date comprising a future date phrase, and a third of the plurality of responsive documents that reference future development contains a future date comprising a future date range;
calculate a cosine similarity using term-frequency-inverse document frequency vectorization for the one or more responsive documents;
group, using the calculated cosine similarity, the one or more responsive documents that reference future development into one or more document clusters;
identify a topic for each of the one or more document clusters from the calculated cosine similarity; and
present the one or more document clusters and the associated topics at the graphical user interface, wherein the graphic user interface includes therein a timeline depicting future dates relating to the at least one of the one or more response documents that reference future developments, including at least one solid interval and at least one fuzzy interval on the timeline, each of the intervals corresponding to at least one of the explicit future date of the first of the plurality of responsive documents that reference future developments the future date phrase of the second of the plurality of responsive documents that reference future developments, and the future date range of the third of the plurality of responsive documents that reference future developments.
14 . The system of claim 13 wherein the server including the processor is further configured to present at the graphical user interface a representation of a similarity measure to a selected document.
15 . The system of claim 13 wherein the server including the processor is further configured to
set a maximum limit on a number of documents within the set of documents for preprocessing;
preprocess the one or more responsive documents with natural language processing, wherein natural language processing includes stemming and punctuation removal;
perform LDA topic modelling on the one or more documents preprocessed with natural language processing; and
select one or more dominant terms within each topic;
determine a strong weighting within the one or more documents after LDA topic modeling and wherein a future term comprises at least one of a modal verb, a common prospective term and an uncommon prospective phrase.
16 . The system of claim 13 wherein grouping the one or more responsive documents that reference future development into one or more document clusters is completed based on at least one of matching keywords, matching subjects, matching entities, matching unstructured text, matching authorship, matching quotes, matching dates, related dates, volume of documents, tagging relationships and direct connections between documents.
17 . The system of claim 13 wherein the timeline has a first range of dates and a scale having a second range of dates smaller than the first range of dates, the graphic user interface further including a second timeline having a start and end corresponding to the second range of dates.
18 . The system of claim 14 wherein the representation of the similarity measure to the selected document is based on the calculated cosine similarity using term-frequency-inverse document frequency vectorization.
19 . The system of claim 17 wherein the second timeline depicts the plurality of responsive documents that reference future developments sorted based on whether documents are official, unofficial, or legislative.