Making Music on Gendia: Two Tracks a Go, and Everything You Can Do After
Gendia's music generator runs Suno V3.5 to V5 at 40–60 credits for two tracks. Plus the parts nobody mentions, covers, extending a track, adding vocals, and splitting a song into stems.
The thing to know before you start: every music run gives you two tracks, not one. Same prompt, two different songs, one price. So the real cost per usable result is half what the number says, and you should listen to both before deciding your prompt was wrong.
The versions, and what they cost
All prices are for one run, which returns two tracks:
| Version | Credits | Notes |
|---|---|---|
| V3.5 | 40 | The cheap one. Fine for drafting an idea |
| V4 | 50 | |
| V4.5 | 50 | Up to 8 minutes |
| V4.5+ | 60 | Richer sound, up to 8 minutes |
| V5 | 60 | The current best, more expressive, and faster |
The spread is 40 to 60. That's narrow enough that there's rarely a good reason to draft on V3.5 unless you're generating a lot of options, unlike video, where the range is a hundredfold and drafting cheap is essential.
Eight minutes is worth pausing on. Most AI music tools give you a fragment. This gives you something you could put under a whole video.
Writing a music prompt
Four things, in roughly this order:
Genre and era. "Lo-fi hip hop", "80s synthwave", "acoustic folk", "orchestral". This does the heaviest lifting.
Instruments. "Fingerpicked acoustic guitar, upright bass, brushed drums." Naming instruments is the most reliable way to steer the sound.
Tempo and energy. "Slow and unhurried", "driving", "90 BPM". Vague tempo is the most common reason a track doesn't fit the video you made it for.
Mood and use. "Warm and nostalgic", "tense", "background music for a product video, should not pull focus."
That last clause matters more than people expect. Music written to be listened to fights a voiceover. Say what the track is for.
Vocals or not. Say so explicitly. "Instrumental, no vocals" is the single most useful phrase in the whole guide if you're scoring video, because unasked-for vocals will wreck a narration track.
What you can do afterwards
This is the part most people never find, and it's where the tool stops being a novelty.
Extend, 40 to 60 credits. Continue a track you generated. This is how you get from a good 2-minute piece to the exact length your video needs.
Upload & extend, 40 to 60. Same, but on audio you brought in yourself rather than generated here.
Upload & cover, 40 to 60. Take existing audio and re-record it in a different style. Your own demo, rebuilt as something else.
Add vocals, 60 (V4.5+ and V5). Put AI vocals over an instrumental.
Add instrumental, 60 (V4.5+ and V5). The reverse: backing music under a vocal.
Lyrics only, 5 credits. Words, no music. Cheap enough to iterate on properly before you spend anything generating audio.
Timestamped lyrics, free. Word timings for a track you already have. This is what you want for captions and subtitles, and it costs nothing.
Stem separation:
- Vocals and instrumental, 40 credits. Two files. This is how you get an instrumental version of a track you like, or isolate a vocal to sit over different music.
- Full split, 200 credits. Twelve or more files, one per instrument. Expensive, and worth it only if you're actually going to mix.
A workflow that uses the cheap steps
- Lyrics first, 5 credits. Iterate until the words are right.
- Draft on V3.5, 40 credits, two tracks. Is the direction right?
- Generate properly on V5, 60, two more tracks.
- Extend, 60, to reach the length you need.
- Timestamped lyrics, free, for captions.
- Stem separate, 40, if you need an instrumental bed under narration.
Under 300 credits for a finished, captioned, correctly-lengthed piece of music with an instrumental version of it. That's about two Nano Banana Pro images.
Two things worth knowing
Listen to both tracks. Every time. They can differ a lot, and the second one is often the better one. People discard a run based on the first track and regenerate for no reason.
Generate the music to fit the video, not the reverse. Decide the length and the mood after the edit exists, then extend to match. Cutting a video to fit music you generated first is the harder job.
Try it
Open the music generator and describe something specific, genre, two instruments, a tempo, and whether you want vocals. Forty credits on V3.5 gets you two tracks and a much better sense of what your prompt should say.
Then: the timeline editor · text to speech
- #AIMusicGenerator
- #SunoGuide
- #AISongGenerator
- #TextToMusic
- #AIBackgroundMusic
- #StemSeparationAI
- #ExtendASongAI
- #RoyaltyFreeAIMusic
- #GendiaTutorial
- #Gendia



