System and methods for screening obstructive sleep apnea during wakefulness using anthropometric information and tracheal breathing sounds
Novel systems and methods use extracted features from audio signals of patient breathing sounds taken during periods of full wakefulness in order to classify patients as either having, or not having, obstructive sleep apnea using a Random Forest machine learning algorithm. In some embodiments, the features are preprocessed then modeled. The models with a high correlation to the response and low overlap percentages are selected for use in the Random Forest classification process.
1 . A method for improving computational analysis of physiological signals to derive a machine-learning obstructive sleep apnea (OSA) screening tool from a heterogeneous initial dataset containing confounding anthropometric variables, said method comprising:
(a) storage, in non-transitory computer readable memory, of an initial dataset that comprises, for each of a plurality of human test subjects from whom respective audio recordings were taken during periods of wakefulness, a respective subject dataset in which there is stored at least:
a known OSA severity parameter of the test subject;
anthropometric data identifying different anthropometric parameters of the test subject; and
audio data containing stored audio signals from said respective audio recording of said test subject;
(b) extraction, by said computer, of at least spectral and bi-spectral features from the audio data of the subject datasets;
(c) selection of a training dataset from said initial dataset, and from said training dataset, grouping together the subject datasets from a first high-severity group of said test subjects whose known OSA severity parameters are above a first threshold, and grouping together the subject datasets of a second low-severity group of said test subjects whose known OSA severity parameters are below a second threshold that is lesser than said first threshold;
(d) based on the anthropometric data, division of the subject datasets from each of the high-severity and low-severity groups into a plurality of anthropometrically distinct subsets, among which the subject data sets in each subset share a common confounding anthropometric variable, thereby partitioning the heterogeneous initial dataset for targeted modeling;
(e) derivation, by said computer, of input for a classifier training procedure at least partly by:
(i) for each anthropometrically distinct subset, filtering said extracted spectral and bi-spectral features down to a selected subset of features;
(ii) for each anthropometrically distinct subset, creating a respective pool of predictive models for predicting OSA severity values, of which each predictive model in said pool is based on a different combination of features from the selected subset of features; and
(iii) for each anthropometrically distinct subset, evaluating the predictive models of the respective pool against one another to reduce the respective pool down to a reduced subset of the predictive models, by, for each predictive model in the pool: (1) dividing a response variable of the model into a plurality of severity segments; (2) calculating a unique correlation value for each of the plurality of severity segments to account for regional non-linearity; (3) determining an average of the calculated correlation values; and (4) filtering the pool of predictive models to generate the reduced subset based on said average of the calculated correlation values;
(f) using said derived input, training by said computer of a classifier for the purpose of classifying a human patient as either OSA or non-OSA based on a respective patient dataset that contains at least:
anthropometric data identifying different anthropometric parameters of the patient; and
at least one of either:
audio data containing stored audio signals from a respective audio recording of said patient during a period of full wakefulness; and/or
feature data concerning features already extracted from said audio signals from the respective audio recording of said patient; and
(g) storage in the same or another non-transitory computer readable memory:
said trained classifier;
for each anthropometrically distinct subset, identification of a respective feature set, from among said extracted spectral and bi-spectral features, that is to be used for classification of patients whose anthropometric parameters overlap those of the subject datasets of said anthropometrically distinct subset; and
statements and instructions executable by one or more computer processors to:
read the anthropometric data of the patient dataset;
match the anthropometric data of the patient dataset to one or more of the anthropometrically distinctive subsets to identify one or more respective feature sets associated therewith;
based on identification of said one or more respective feature sets, select which particular features are required from the audio data or feature data of the patient dataset; and
input said particular features to said trained classifier to classify said patients as either OSA or non-OSA.
2 . The method of claim 1 wherein division of the response variable in step (e)(iii) is based on real OSA severity values from the subject datasets of the test subjects.
3 . The method of claim 2 wherein step (e)(iii) further comprises, for each of said predictive models in each pool:
calculating average OSA severity values for the first set of different OSA severity groups;
determining whether said average OSA severity values increase or decrease in matching sequence to said first set of different OSA severity groups; and
removing from said pool any predictive model whose average OSA severity values do not increase or decrease in matching sequence to said first set of different OSA severity groups.
4 . The method of claim 3 wherein said average OSA severity values are calculated averages of predicted OSA severity values from the predictive model.
5 . The method of claim 4 wherein step (e)(iii) further comprises, for each of said predictive models in each pool:
dividing datapoints of each of said predictive models in said pool into a second set of different OSA severity groups based on the predicted OSA severity values from the predictive model;
calculating average real OSA severity values for the second set of different OSA severity groups;
determining whether said average real OSA severity values increase or decrease in matching sequence to said second set of different OSA severity groups; and
removing from said pool any predictive model whose average real OSA severity values do not increase or decrease in matching sequence to said second set of different OSA severity groups.
6 . The method of claim 5 wherein step (e)(iii) further comprises:
for each predictive model in each pool, calculating an overlap percentage between adjacent OSA severity groups in both the first and second sets of different OSA severity groups;
for each predictive model in each pool, calculating an average of the overlap percentages of said predictive model for both the first and second sets of different OSA severity groups.
7 . The method of claim 6 wherein step (e)(iii) further comprises:
for each predictive model in each pool, calculating an overall average of the average overlap percentages from both the first and second sets of different OSA severity groups; and
filtering each pool down to said reduced subset of the predictive models therein based on which of said predictive models have a lesser overall average overlap percentage than other models from said pool.
8 . The method of claim 2 wherein step (e)(iii) further comprises, for each predictive model in each pool, calculating an overlap percentage between adjacent OSA severity groups in each set of different OSA severity groups.
9 . The method of claim 8 wherein step (e)(iii) further comprises, for each set of different OSA severity groups for each predictive model in each pool, calculating an average of the overlap percentages of said predictive model.
10 . The method of claim 9 wherein step (e)(iii) further comprises, using said average of the overlap percentages of each predictive model, filtering each pool down to said subset of the predictive models therein based on which of said predictive models have a lesser average overlap between said adjacent OSA severity groups than other models from said pool.
11 . The method of claim 1 wherein step (c) comprises also selecting a blind dataset from said initial dataset, and step (e)(iii) further comprises, from each pool, running each predictive model from said pool using the subject datasets from the blind dataset to calculate estimated OSA severity values for said blind dataset, and assessing an accuracy of said estimated OSA severity values against real OSA severity values from said blind dataset.
12 . The method of claim 11 wherein step (f) comprises, for each predictive model in each pool, at least one training procedure having a classifier training step and a classifier testing step, of which the classifier training step comprises running the classifier with model-predicted OSA severity values from a subset of the training dataset, and the classifier testing step comprises rerunning of the classifier with model-predicted OSA severity values from the same or a different subset of the training dataset, and evaluating classification results from said rerunning off the classifier against real OSA severity values from said same or different subset of the training dataset.
13 . The method of claim 1 wherein step (c) comprises also selecting a blind dataset from said initial dataset, and step (f) comprises, for each predictive model in the reduced subset of the predictive models, at least one training procedure having a classifier training step and a classifier validation step, of which the classifier training step comprises running the classifier with model-predicted OSA severity values from a subset of the training dataset, and the classifier validation step comprises rerunning the classifier with model-predicted OSA severity values from the same or a different subset of the training dataset, and comparing classification results from said rerunning of the classifier against real OSA severity values from said same or different subset of the training dataset to evaluate performance of the predictive model.
14 . A method of performing an obstructive sleep apnea (OSA) screening test on a patient, said method comprising:
(a) obtaining one or more computer readable media on which there is stored:
a trained classifier derived in the manner recited in at least steps (a) through (f) of claim 1 ;
for said patient, a patient dataset of the type recited in step (f) of claim 1 ; and
statements and instructions executable by one or more computer processors;
(b) through execution of said statements and instructions by said one or more computer processors:
(i) reading the anthropometric data of said patient dataset;
(ii) running, multiple times, a trained classifier derived in the manner recited in at least steps (a) through (f) of claim 1 , each time starting with input composed of or derived from a different feature combination, comprised of at least spectral and bi-spectral features, particularly selected or derived from the patient dataset for a different anthropometric parameter read from the anthropometric data of said patient dataset, and for each run of said trained classifier, deriving therefrom a respective classification result classifying the patient as either OSA or non-OSA;
(iii) based on the classification results from step (b) (ii), deriving a final classification result for the patient; and
(iv) storing or displaying said final classification.
15 . Non-transitory computer readable memory having stored thereon statements and instructions executable by one or more processors to, when executed, perform at least step (b) of claim 14 .
16 . The method of claim 14 comprising:
prior to step (a), recording tracheal breathing sounds of said patient during said period of wakefulness, thereby obtaining said respective audio recording of said patient, and digitally storing the audio data containing the stored audio signals from said respective audio recording; and
after step (g), diagnosing said patient as having OSA based on the classification from the trained classifier.
17 . Non-transitory computer readable memory having stored thereon executable statements and instructions configured to, when executed by the one or more processors, perform at least steps (e) and (f) of claim 1 .
18 . The method of claim 1 wherein a total quantity of predictive models created in step (e) comprises at least 2,925 predictive models.
19 . The method of claim 1 wherein a total quantity of predictive models created in step (e) comprises at least 17,550 predictive models.
20 . The method of claim 1 wherein a total quantity of predictive models created in step (e) comprises at least 80,730 predictive models.