FIELD GUIDE / VOICE DIRECTION
Gemini TTS Prompting Guide
A useful direction tells the speaker what to do with the words. This Gemini TTS prompting guide shows how to separate spoken text from delivery instructions, choose a voice, place vocal cues, and build a two-person exchange. Copy an example or open it in the studio, then adapt one setting at a time.
Explore the examplesStart with the words you want to hear
Put spoken content in the script field. Use the voice description for the way it should sound, and scene direction for the situation around a conversation. If you put a sentence such as “read this gently” in the spoken script, you risk asking the model to say the instruction aloud. Keeping these jobs separate makes a draft easier to understand and easier to revise when a take needs work.
Start with two or three sentences that represent the final material. Include any name, technical word, or transition that might cause trouble. Listen once for overall tone, then again for individual words. A short test gives you a clear question to answer: does this speaker sound appropriate for this passage? You do not need to configure every control before finding out.
The working method throughout this Gemini TTS prompting guide is simple: write a clear passage, choose a voice, add one useful direction, and listen. Keep the text unchanged while comparing voices. Keep the voice unchanged while comparing styles. This makes it easier to identify the setting that caused an improvement instead of guessing between several simultaneous changes.
Gemini TTS prompting guide: voice instructions
A good voice description names a delivery the listener could recognise. “Friendly teacher, clear and patient” is easier to interpret than a long collection of adjectives. For a story, try “quiet narrator, measured pace, curious at the end.” For a tutorial, try “calm explanation, crisp instructions, brief pauses between steps.” These examples describe a performance without changing the words being spoken.
Avoid contradictory directions such as very fast, extremely relaxed, energetic, whispered, and flat all at once. Decide which quality matters most for the scene. If you need a change partway through a long reading, consider making separate segments with distinct direction. That is easier to review than asking one large block of instructions to govern several different scenes.
A gentle delivery
Use a short voice instruction to change tone without adding spoken stage directions.
Take a moment to settle into your chair. <breath> There is no need to hurry. Notice the light coming through the window, and let the next breath arrive in its own time.
Voice direction: Gentle, reassuring delivery with a relaxed pace.
Nonfiction narration
Explain one idea at a comfortable pace, keeping the key distinction easy to follow.
A useful habit starts with a clear cue. Think of the moment you put a cup beside the kettle. That small action can remind you to fill your water bottle for the day. <short pause> Begin with one change you can repeat. Once it feels familiar, connect the next step to the same routine.
Voice direction: Clear nonfiction narrator. Friendly and steady, with gentle emphasis on practical actions.
Let meaning guide emotion and pace
A voice can sound concerned without sounding theatrical. Give it a reason for the emotion: a reassuring explanation to a worried listener, or a curious question asked during an interview. Brief scene context often communicates more than repeatedly demanding a stronger emotional effect. Keep the feeling appropriate to the text; an upbeat promotional tone can distract from a careful safety explanation or a reflective story.
Pace is also about sentence structure. If a line feels rushed, shorten the sentence or split an idea into two parts before choosing a slower setting. If it feels lifeless, remove unnecessary qualifications and put the main action earlier. A natural recording usually begins with language that a person could comfortably say. This Gemini TTS prompting guide treats script editing as part of voice direction, not an afterthought.
Gemini TTS prompting guide: two speakers
Switch to Two speakers, choose a voice for each role, and assign each line to speaker one or speaker two. Write only the spoken words in each turn. You do not need to prefix the text with “Speaker 1:” because the interface already stores that assignment. A shared scene description can explain where the conversation takes place and how the speakers relate to each other.
Give each participant a consistent role. An interviewer might ask short, curious questions while a guest gives measured answers. A learner might express uncertainty while a teacher explains one idea at a time. A useful contrast comes from their purpose as well as their voices. Avoid making both speakers repeat the same information simply to extend the exchange.
Read the dialogue aloud in order. Every answer should respond to the preceding line, and each pause should help the exchange feel intentional. The application supports two speakers per generation, so a script with a larger cast needs to be adapted or produced in separate requests. Those requests are not automatically merged into a finished episode.
A short interview
A host asks a focused question and a guest answers with a practical example.
Speaker 1: What helped you finish your first project? Speaker 2: Making the first version smaller. I chose one problem and tested it with a friend. Speaker 1: What changed after that conversation? Speaker 2: I rewrote the opening. Once the purpose was clear, the rest was much easier to explain.
Scene direction: A relaxed interview. The host is curious; the guest answers thoughtfully without rushing.
A question and answer lesson
Let a learner ask for clarification, then have a teacher explain a single concept.
Speaker 1: Why does the moon look different each night? Speaker 2: We see different amounts of its sunlit side as it travels around Earth. Speaker 1: So its shape does not actually change? Speaker 2: Exactly. <short pause> The moon stays round. Our view of the light changes.
Scene direction: A friendly science lesson. Keep both speakers clear and conversational.
A scene in two voices
Two characters disagree about what they have found. Keep their voices consistent across the exchange.
Speaker 1: You left this envelope on the bench, didn’t you? Speaker 2: No. <short pause> I was about to ask you the same thing. Speaker 1: Then how did they know I would be here? Speaker 2: Look at the time on the letter. Whoever wrote it is still waiting.
Scene direction: A quiet station at night. Two people speak softly; curiosity grows into concern.
Revise one cause at a time
If a name sounds wrong, start by checking the spelling and nearby punctuation. Try a clearer written form in a short test, then listen again. If a sentence sounds too intense, simplify the voice direction or remove a vocal cue. If the recording sounds monotonous, first check whether the script has natural sentence variation and a clear point of emphasis.
If a turn is missing or assigned incorrectly, check the speaker selector on every line and make sure no required turn is empty. Review the displayed character count before submitting. The studio accepts up to 5,000 characters per generation; long passages should be split at meaningful boundaries. Dialogue length includes the submitted turns rather than treating every speaker as a separate allowance.
Keep a simple record of the chosen voice, model, direction, and the exact text used for a successful take. This Gemini TTS prompting guide provides starting examples, but your own reviewed settings are the best reference for the next segment of the same project. Consistency is easier when the inputs are deliberate and the finished audio is checked.
Questions before you start
Do the examples in this Gemini TTS prompting guide cost credits?
Reading, copying, and applying a template do not submit a generation request. Credits are involved when you choose Generate speech after signing in. The estimate beside the button updates with the model and script, so review it before each new take.
Can I put all direction in the script?
Use the separate voice description and scene fields for general instructions. Keep the script focused on spoken words and supported inline cues. This separation also helps you copy or revise the actual narration without accidentally carrying spoken stage directions into the next version.
Will identical settings produce identical recordings?
Do not rely on a new generation being an exact copy of a previous take. Keep the WAV file you want to use, and compare new results before replacing it. Voice previews demonstrate a neutral example; they do not guarantee the delivery of every custom passage.
Can I use this Gemini TTS prompting guide for both models?
Yes, the examples use the controls available in this application for Flash and Flash-Lite. Ordinary examples keep your selected model when replacing a draft. Explicit model presets can switch it. Use the model comparison page when deciding which option to test first.
Where should I go after this Gemini TTS prompting guide?
Choose a voice in the library, then apply the example closest to your task. The YouTube and audiobook pages explain how to prepare segments for those workflows. For an exchange between two people, use the dialogue page and replace its sample lines with your own.