Sleep assistance through generation and evaluation of generative sleep content and/or sleep improvement interventions including training, adjusting, mediating, and/or integrating outputs of one or more AI models
Disclosed are a system, a device, and/or a method of sleep assistance through generation and evaluation of generative sleep content and/or sleep improvement interventions including training, adjusting, mediating, and/or integrating outputs of one or more AI models. In one embodiment, a system includes an earbud generating an audio generation request through a voice interface, which is parsed by a sleep assistance server to extract a narrative prompt. A generative content server inputs the narrative prompt into an artificial neural network to generate a text data, and then text-to-speech model to generate a narrative audio. A content integration routine may overlay the narrative audio data with additional music and/or physiological guidance data. A generative audio data is returned to the earbud to assist in achieving sleep. The system may evaluate physiological data to determine effectiveness of the sleep assistance audio, and may augment, fine-tune, and/or retrain one or more AI models.
1 . A system for generating customized audio for increasing effectiveness of relaxation and/or sleep, the system comprising:
a network;
an earbud wearable by a user comprising:
a speaker,
a microphone,
a processor of the earbud,
a memory of the earbud that is a non-transient computer readable memory of the earbud,
a wireless network interface controller communicatively coupled to the network, and
a request generation routine comprising computer readable instructions that when executed on the processor of the earbud generate an audio generation request through a voice interface implemented on the speaker and the microphone;
a sleep assistance server communicatively coupled to the network, comprising:
a processor of the sleep assistance server,
a memory of the sleep assistance server that is a non-transient computer readable memory of the sleep assistance server,
a request agent comprising computer readable instructions that when executed receive the audio generation request,
a content prompt extraction routine comprising computer readable instructions that when executed:
parse the audio generation request to extract a narrative prompt, and
store the narrative prompt in the memory of the sleep assistance server; and
a generative content server communicatively coupled to the network, comprising:
a processor of the generative content server,
a memory of the generative content server,
a text generation routine comprising computer readable instructions that when executed:
input the narrative prompt into a text generation model comprising an artificial neural network of the text generation model,
wherein the artificial neural network comprising a plurality of nodes comprising a set of input nodes, a set of hidden nodes, and a set of output nodes, and
generate a text data as the output of the artificial neural network of the text generation model,
a voice generation routine comprising computer readable instructions that when executed:
input the text data into a text-to-speech model, and
store a narrative audio data as an output of the text-to-speech model,
a content integration routine comprising computer readable instructions that when executed generate a generative audio data comprising an overlay of the narrative audio data and at least one of a music audio data, a physiological guidance data, and an ambient audio data, and
a generative content engine comprising computer readable instructions that when executed transmit the generative audio data to the earbud of the user to assist the user in achieving at least one of relaxation and sleep through the customized audio.
2 . The system of claim 1 ,
wherein the earbud further comprising:
an inertial measurement unit, and
a physiological signal agent comprising computer readable instructions that when executed gather a physiological data of the user from the inertial measurement unit of the earbud worn by the user while the generative audio data plays sound on the speaker of the earbud,
wherein the physiological data comprising at least one of motion of the user comprising a heartbeat, a respiration, and a macro movement; and
wherein the sleep assistance server further comprising:
a state monitoring routine comprising computer readable instructions that when executed utilize the physiological data gathered from the inertial measurement unit to determine the user is in at least one of a sleep state and an awake state over a time period,
a sleep metric routine comprising computer readable instructions that when executed determine one or more sleep metrics for a sleep session comprising a sleep onset latency value, a length of a sleep period during the time period, a ratio of the sleep period to an awake period during the time period, a number of rapid eye movement (REM) periods, a length of REM periods, a number of Non-REM periods, a length of Non-REM periods, a number of interstitial awake periods, and a length of the awake period during the time period, and
a sleep content evaluation routine comprising computer readable instructions that when executed:
extract one or more generative elements of at least one of the narrative audio data and the text data associated with the narrative audio data, and
store one or more generative elements in association with an effectiveness value and an effectiveness rating based on the one or more sleep metrics.
3 . The system of claim 1 , wherein the generative content server further comprising:
a general augmentation routine comprising computer readable instructions that when executed:
querying a general augment data comprising at least one of general augment narrative data, general augment music data, general augment ambient data, and general augment voice data,
extract a subset of the general augment data based on textual association with the audio generation request, and
load the subset of the general augment data into at least one of an input prompt of an artificial neural network and a context window of the artificial neural network,
wherein the input prompt comprises at least one of the narrative prompt, a music prompt, and an ambient prompt, and
wherein the artificial neural network is at least one of an artificial neural network of the text generation model.
4 . The system of claim 1 , further comprising:
a specific augmentation routine comprising computer readable instructions that when executed:
query a user specific augment data comprising at least one of user augment narrative data, user augment music data, user augment ambient data, and user augment voice data,
extract a subset of the user specific augment data relevant to the audio generation request,
overwrite at least some of the subset of the general augment data within at least one of the input prompt of the artificial neural network and a context window of the artificial neural network, and
load the subset of the user specific augment data into at least one of the input prompt of the artificial neural network and the context window of the artificial neural network.
5 . The system of claim 1 ,
wherein the generative content server further comprising:
a model training routine comprising computer readable instructions that when executed retrain the artificial neural network of the text generation model based on at least one of the effectiveness value and the effectiveness rating,
wherein retraining comprises adjusting a parameter of at least one of the artificial neural network of the text generation model,
wherein adjusting the parameter comprising modifying a node weight of an artificial neural network (ANN) node,
wherein tuning the parameter adjusts a weight value of at least one node of the set of input nodes, the set of hidden nodes, the set of output nodes,
wherein the text generation model is a large language model, and
wherein the text data input into a voice synthesizer comprising a custom voice model; and
a guidance prioritization subroutine comprising computer readable instructions that when executed constrain the output of the artificial neural network of a narrative generation model to produce the text data in which a text clause of the text data is temporally associable with a physiological guidance element of a physiological guidance template,
wherein the text-to-speech model generates the narrative audio such that a voiceover of the text data is temporally associated with a physiological guidance element.
6 . A method for generating customized audio for increasing effectiveness of relaxation and/or sleep, the method comprising:
receiving an audio generation request through a voice interface collected on a microphone of an earbud;
parsing the audio generation request to extract a narrative prompt and a physiological guidance prompt;
storing the narrative prompt in a computer readable memory;
determining a narrative modifier from the audio generation request comprising at least one of a narrative style description and a narrative genre description;
inputting the narrative prompt into a text generation model comprising an artificial neural network of the text generation model;
generating a text data as an output of the artificial neural network of the text generation model;
inputting the text data into a text-to-speech model;
outputting a narrative audio data;
storing the physiological guidance prompt in the computer readable memory;
determining a physiological guidance modifier from the audio generation request comprising a physiological guidance type,
wherein the physiological guidance type comprising at least one of a respiration rate, a respiration pattern, and a heart rate;
inputting the physiological guidance prompt into a physiological guidance model comprising an artificial neural network of the physiological guidance generation model,
generating a physiological guidance template as an output of the artificial neural network of the physiological guidance model;
wherein the physiological guidance comprising one or more physiological
generating a generative audio data comprising an overlay of the narrative audio data and audio generated from the physiological guidance template; and
transmitting the generative audio data to the earbud of a user to assist the user in achieving at least one of relaxation and sleep through the customized audio.
7 . The method of claim 6 , further comprising:
parsing the audio generation request to further determine if a music prompt is included within the audio generation request;
generating, if the music prompt was not present when the audio generation request was parsed, the music prompt by inputting the narrative prompt into a text-music relation model relating a text to at least one of musical elements, a music style description, and a music genre description;
optionally extracting a music modifier comprising at least one of a music filter, the music style description, and the music genre description;
storing in the computer readable memory the music prompt and optionally at least one of the music style description, the music genre description, and the music filter;
inputting the music prompt into a music generation model comprising an artificial neural network of the music generation model,
wherein the music generation model trained with training data comprising associations between text tokens and musical elements; and
generating a music audio data as an output of the artificial neural network of the music generation model,
wherein the generative audio data further comprising an overlay of the music audio data.
8 . The method of claim 7 , further comprising:
parsing the audio generation request to further determine an ambient sound prompt;
storing the ambient sound prompt in the computer readable memory;
receiving an ambient modifier comprising at least one of an ambient filter and an ambient style description;
inputting the ambient sound prompt and the ambient modifier into an ambient sound generation model comprising an artificial neural network of the ambient sound generation model,
wherein the ambient sound generation model trained with training data comprising associations between text tokens and sound elements; and
generating an ambient audio data as an output of the artificial neural network of the ambient generation model,
wherein the generative audio data further comprising an overlay of the ambient audio data.
9 . The method of claim 8 , further comprising:
querying a general augment data comprising at least one of general augment narrative data, general augment music data, general augment ambient data, and general augment voice data;
extracting a subset of the general augment data based on textual association with the audio generation request; and
loading the subset of the general augment data into at least one of an input prompt of an artificial neural network and a context window of the artificial neural network,
wherein the input prompt comprises at least one of the narrative prompt, the music prompt, and the ambient prompt, and
wherein the artificial neural network is at least one of an artificial neural network of the text generation model, an artificial neural network of the music generation model, and an artificial neural network of the ambient sound generation model.
10 . The method of claim 9 , further comprising:
querying a user specific augment data comprising at least one of user augment narrative data, user augment music data, user augment ambient data, and user augment voice data,
extracting a subset of the user specific augment data relevant to the audio generation request;
overwriting at least some of the subset of the general augment data within at least one of the input prompt of the artificial neural network and the context window of the artificial neural network; and
loading the subset of the user specific augment data into at least one of the input prompt of the artificial neural network and the context window of the artificial neural network.
11 . The method of claim 10 , further comprising:
constraining the output of the artificial neural network of the ambient sound generation model to produce an ambient audio data in which an ambient element is temporally associated with a physiological guidance element;
constraining the output of the artificial neural network of the music generation model to produce the music audio data in which a musical element is temporally associated with a physiological guidance element; and
constraining the output of the artificial neural network of the text generation model to produce the text data in which a text clause of the text data is temporally associable with a physiological guidance element,
wherein the text-to-speech model generates the narrative audio such that a voiceover of the text data is temporally associated with a physiological guidance element.
12 . The method of claim 11 , further comprising:
gathering a physiological data of the user from a sensor of an earbud worn by the user while the generative audio data plays sound on a speaker of the earbud;
determining the user is in a sleep state;
determining one or more sleep metrics comprising a sleep onset latency value, a length of a sleep period during a time period, a ratio of the sleep period to an awake period during the time period, a number of rapid eye movement (REM) periods, a length of REM periods, a number of Non-REM periods, a length of Non-REM periods, a number of interstitial awake periods, and a length of the awake period during the time period;
extracting one or more generative elements of at least one of the generative audio data, the narrative audio data, the text data associated with the narrative audio data, the music audio data, the ambient audio data, and the one or more physiological guidance elements;
storing the one or more generative elements in association with an effectiveness value and an effectiveness rating; and
retraining at least one of the artificial neural network of the text generation model, the artificial neural network of the music generation model, the artificial neural network of the physiological guidance model and the artificial neural network of the ambient generation model,
wherein retraining comprises adjusting a parameter of at least one of the artificial neural network of the text generation model, the artificial neural network of the music generation model, and the artificial neural network of the physiological guidance model and the artificial neural network of the ambient generation model, and
wherein adjusting the parameter comprising modifying a node weight of an artificial neural network (ANN) node.
13 . The method of claim 12 , further comprising:
querying a prerecorded audio comprising at least one of a recorded narrative audio, a recorded music audio, a recorded ambient audio, and a recorded physiological guidance audio, and
integrating the prerecorded audio with the generative audio,
wherein the artificial neural network each comprising a plurality nodes comprising a set of input nodes, a set of hidden nodes, and a set of output nodes,
wherein tuning the parameter adjusts a weight value of at least one node of the set of input nodes, the set of hidden nodes, and the set of output nodes,
wherein the text generation model is a large language model, and
wherein the text data input into a voice synthesizer comprising a custom voice model.