From prompt to voice: How an AI-generated meditation is made
Text input, language model, speech synthesis, audio storage: Using MindIsland as an example, here is how an AI meditation is created and where the technology hits its limits.

An AI-generated meditation is created in four stages: A user describes what they want to meditate on, a language model writes a meditation script from it, speech synthesis reads it aloud, and a server generates and stores the finished audio file. The MindIsland app discloses in its privacy policy which services it uses for this. That makes it a good example for tracing the chain from input to voice. MindIsland is owned by ThreeFlavors LLC.
Stage 1: The input
It starts with a short text. MindIsland asks: “What would you like to meditate on?” Typical inputs are sentences like “Big presentation tomorrow, thoughts won’t slow down” or “Hard conversation with my partner.” According to a screenshot, there is also a voice input icon next to typing.
The user also sets the parameters:
- the length: 3, 5, or 10 minutes
- the voice, such as “Luna – Female”
- optional background music
- the posture: chair, floor, lying down, or standing
These details matter because without instructions, a language model would have no reason to tailor a script to exactly five minutes or to a person standing on a train.
Anyone who does not want to enter anything gets a daily meditation that is generated anew every morning. It also goes through the same chain, just without a personal starting text.
Stage 2: The language model writes the script
According to the privacy policy, the meditation script is generated by the Google Gemini API. A large language model like Gemini generates text word by word based on probabilities it has learned from large amounts of text. It does not “understand” stress in the human sense, but it can respond to it in linguistically appropriate ways.
For that to become a usable meditation, the model needs a fixed structure. At MindIsland, every session consists of three parts:
- The problem is acknowledged and named.
- The trigger is gently reframed.
- Breath and silence lead into relaxation.
There is also a fixed opening: inhale for four seconds, a short pause, exhale for six seconds. Such instructions constrain the model. That is intentional. The clearer the framework, the more predictable and calmer the result sounds.
Exactly how the maker phrases its instructions to the model is not public. Only the structure that becomes audible in the end is known.
Stage 3: Text becomes voice
A synthetic voice reads the finished script aloud. According to the privacy policy, MindIsland uses the Inworld service for this. Modern text-to-speech systems generate speech with neural networks. They calculate pitch, tempo, and emphasis so that the voice sounds much more natural than earlier computer voices.
For meditations, tempo is crucial. A calm script only feels calm if pauses fall in the right places. The maker speaks of “warm, natural voices” and several voices to choose from. It does not say exactly how many there are.
After that, the optional background music is added. According to the website, it is mixed so that it “ducks” under the spoken words. In audio engineering, this process is called ducking: The music gets quieter as soon as speech begins.
Stage 4: Generating and storing the audio
According to the privacy policy, the audio file is generated and stored on Amazon Web Services servers in the EU, specifically in Ireland. From there, the app downloads it to the phone.
The last ten meditations remain cached on the device. They can therefore be played without a network connection, such as on an airplane. Creating a new meditation, by contrast, only works online because the text model and speech synthesis run in the cloud.
Other background services are Firebase for push notifications and Resend for emails. They have nothing to do with the meditation itself.
What the technology cannot do
Generative AI has known weaknesses, and they apply here too.
- No understanding in the human sense: The model reacts to words, not to the person behind them. A brief sentence provides only limited context.
- Possible inaccuracies: Language models can choose phrasings that do not quite fit. Fixed structures reduce this, but they do not rule it out.
- Not a professional: A generated meditation does not replace counseling. MindIsland itself points out that the app is not a medical device and does not offer therapy, diagnosis, or treatment. The app says: “AI-generated, not a substitute for professional advice.”
- Language: Currently, only English meditations are documented.
Privacy: What happens when you type
Anyone who writes their worries into a text field reveals personal information. That text must be sent to the language model, or no suitable meditation can be created. It is therefore worth reading the privacy policy of such a service and checking which providers are involved.
MindIsland names these providers openly and advertises GDPR compliance. Users can export their data or delete their account at any time. Regardless, a simple rule applies: Names, addresses, or other identifying details do not belong in a meditation prompt. “Argument with my brother” is enough as a starting point.
More information about the app is available on the MindIsland website.
The bottom line
An AI meditation is the result of a chain of input, text model, speech synthesis, and cloud storage. At MindIsland, those are Google Gemini, Inworld, and AWS in Ireland. Quality depends above all on how well the fixed framework guides the model: length, posture, breathing opening, and the sequence of acknowledging, reframing, and relaxing. The technology can deliver a calm, suitable script. It does not replace human support for serious problems.