Write an audio drama prompt in the order the audience will hear it. Set the room tone, introduce each speaker with a useful voice cue, write the dialogue as it should be performed, and place sound changes where they happen.
This guide focuses on scene text and performance direction for FanficLab. If you already have an excerpt from AO3, start with How to Turn an AO3 Fic Into Audio Drama.
Build the prompt in scene order
Use this basic sequence:
Prompt = Sound Effects + Speaker A Voice and Emotion Description + Speaker A Line + Speaker B Voice Description + Speaker B Line
Start with the opening sound, add the first speaker's voice and emotion, then write the line. Move through the scene one beat at a time:
- Opening sound: establish the location, ambience, music, weather, crowd, or silence.
- Speaker cue: name the character and add only the voice, emotion, accent, or delivery details that matter.
- Spoken line: write the exact words, using punctuation to guide the rhythm.
- Next audible beat: introduce the next speaker, action, or sound change when it occurs.
Turn visual cues into sound
Translate visual cues into something the audience can hear. If a character is nervous, write the physical signs: "Her breath catches; she answers too quickly, trying to sound calm." Useful direction covers voice texture, emotion, delivery, audible action, and sound changes. Place each cue beside the line it affects.
When you cue music, include its entrance, level, and exit: "Upbeat intro music fades in before the opening line, stays under the voices, then fades out after the final sign-off."
For a hurt/comfort reunion between original characters, a compact prompt might read:
Rain taps against the fire escape. Mira (hoarse, holding back tears) catches her breath as the apartment door opens. Wet footsteps stop in the hall. Mira says quietly, "You came back." Jon (tired, gentle, almost whispering) answers, "You asked me to." His coat drops to the floor. Leave a long pause before the room tone fades.
Direct multilingual scenes clearly
Write each spoken line in the language you want the audience to hear. Your voice, emotion, action, and scene directions can use another language if that’s easier to work with. For example:
A man says warmly: “Todo estará bien.”
In a multilingual scene, keep character names consistent and mark every language switch. Choose a preset or reference voice that fits the language and performance because voice capabilities vary.
Keep the scene within 2,800 characters
Each FanficLab generation accepts up to 2,800 characters, including spaces and punctuation, and can produce about two minutes of audio. Use one clear emotional turn, leave room for pauses and reactions, and split longer chapters into separate scenes.
Example 1: Build an action scene in audible layers
This Monster Hunt prompt starts with weather, gives its two speakers distinct voices, and places physical sounds between their lines.
Above a raging sea, continuous howling gale-force wind and the crashing roar of towering waves smashing down; a thunderous eruption of water and a deep booming roar. Varmak (adult male, rough booming voice, commanding, battle-hardened, epic sci-fi style) shouts with challenging authority, “Cast the nets! Don’t hesitate! Hook the edge of its ring segment and cinch the harness tight—hold the line, we take the Kraken alive!” Paul (young adult male, tense strained breath, low determined voice, epic sci-fi style) murmurs inwardly, “I must not fear. This sea has taught me.” Pounding, surging swim strokes, then a sharp metallic hook clamping into the scales and the groan of straining ropes drawing taut; a huge water-churning bellow of pain with heavy roiling currents that slowly weakens into exhausted, defeated snorts. Paul shouts with sudden commanding force, “We’ve got it—it’s caught!” Varmak laughs loudly with pride, “Look what we’ve landed! The Kraken is ours! From this day on, he is one of us!”
Wind and waves set the scale. The hook, ropes, and weakening roar keep the action easy to follow by ear.
Example 2: Direct one voice and one crowd
This single-character prompt directs the commentator's delivery and times the crowd reaction to the goal:
Inside a huge football stadium, with the deafening roar of tens of thousands of fans throughout the background. The commentator (middle-aged male, British accent, rich and penetrating voice, classic sports commentary, extremely exhilarated) shouts in a rapid, soaring, full-throated tone: “OH, HE SCORES!!! WHAT A GOAL! He beats two men and buries it in the top corner—UNBELIEVABLE! The stadium is on its feet!!!” He draws out the word “GOAL” with a voice slightly hoarse from excitement, and the crowd’s cheering erupts at the moment of the goal and continues to the end.
The speed, projection, vocal strain, elongated "GOAL," and crowd timing all sit beside the line they affect.
Example 3: Write reactions as part of the dialogue rhythm
For comedy, write the reaction beats too. This scene cues its music, room tone, pauses, voice cracks, laughter, and applause:
Opens with a classic upbeat sitcom intro tune—bright electric-guitar strumming, cheerful bass, crisp drums, playful and bouncy; plays for a few seconds, then fades out. Inside a pet store, relaxed atmosphere, very faint room tone. Dave (man around 30, American accent, low voice, forced composure, pretend expert) says smugly, “Oh, this Golden Retriever is so well taken care of, scientifically raised since a pup. How old’s the little guy?” Lily (young woman around 25, American accent, bright lively voice, blunt) sincerely blurts out, “Sir… that’s an alpaca.” Immediately, classic sitcom audience laughter bursts out, fading from loud to soft. Dave freezes, struggles to save face, and his voice cracks: “I… I know that. I mean it… it really looks like a Golden Retriever.” Audience laughter erupts louder. Lily nods earnestly: “Yeah, it thinks so too—that’s why it’s been staring at you this whole time.” End with audience laughter mixed with scattered applause.
The pauses leave room for the joke, and each burst of laughter sets up the next line.
Reference audio in FanficLab
In Voice mode, choose or upload one clean reference clip for each recurring character. You can guide up to three voices. Use the same role name beside the character's lines so FanficLab can map the clip correctly. The clip guides the voice; the prompt supplies this scene's emotion and delivery. See the References and Characters guide for recording tips.
Use timeline control when a beat must land at a specific moment
Add a time range before a line when it needs to land inside a specific window:
Total audio duration: 20 seconds. A quiet concert hall with faint audience murmurs and soft strings. Ethan (quiet, curious, almost a whisper):
[2.8s:8.4s]"I didn't expect the room to feel so alive." Lena (warm, speaking softly):[10.2s:18.2s]"The musicians are breathing together." Let the strings linger after her line.
Timeline control helps with reveals, interruptions, entrances, music cues, and fixed-duration scenes. Leave gaps for breaths and transitions.
Checklist before you generate
Check these details before you generate:
- What does the audience hear first?
- Can you tell who's speaking from the character name and voice cue?
- Does each change in emotion sit beside the line it affects?
- Do the dialogue, sound effects, and music appear in the order they happen?
- Does the scene have enough room to breathe?
- Do voice-reference names match the role names in the prompt?

