How to Make a Song with AI: From Prompt to Finished Track
Most guides to AI music stop at write a prompt and listen. This one is for people who need a track built to a specific timeline, not entertainment: what to decide before you open the generator, how to write the style prompt, how many attempts a usable take really needs, and where the model’s job ends and manual editing begins.

In this article
I make songs with AI because a client video needs music cut to a specific timeline, and hiring a composer for every job is not realistic when you are also the editor, the colourist and the producer. What follows is the order that actually works: what to decide before generating, how to write the style prompt, how to hand over lyrics, how many attempts a usable take really takes, and where the model stops and manual editing starts. No excitement about the first lucky generation, only what actually saves time.
What to decide before you generate
The common mistake is opening the generator and typing something like “a great song for a video” into the style field. The model takes that literally, and the result is not about anything specific. Before writing a prompt, lock in four things.
- Genre. Not “music”, something specific: synth-pop, lo-fi hip hop, acoustic, orchestral epic. The more precise the genre, the less the model has to guess on your behalf.
- Tempo. Even a rough slow, medium or fast narrows the result more than genre does. For a video with a cut-driven edit, tempo often matters more than genre.
- Mood. One or two words: tense, warm, triumphant, nostalgic. Mood shapes the arrangement more than any list of instruments will.
- Language. Suno sings in dozens of languages, not only English. State the one you want. Leave it out and the generator picks for you, and it will not always pick the one you meant.
Make these four calls before opening the generator, not while you are in it. Otherwise you end up rewriting the prompt in circles, spending credits on guessing rather than on results.
How to write the style prompt
The style prompt is not creative writing, it is a comma-separated list of specific traits: genre, tempo, mood, key instruments, a reference sound if you have one. For example: synth-pop, mid tempo, warm nostalgic mood, analogue synths, live drums, female vocal. The more concrete the wording, the more predictable the result. Vague adjectives with no noun attached get interpreted however the model feels like that day.
It is worth stating what should not be there too: no aggressive electronic elements, no choir, no abrupt transitions. Negative instructions work less reliably than positive ones, but sometimes that is the only way to cut something the model defaults to for a given genre.
How to supply lyrics and tag verse and chorus
If you need a track with real lyrics rather than an instrumental, write the lyrics in full beforehand and hand them over as they are, rather than hoping the generator writes something usable on its own. Tagging the structure, verse, chorus, bridge, helps the model hold a shape instead of drifting into one long section with no real hook.
The practice is simple: mark each block before it starts, verse 1, chorus, verse 2, chorus, bridge, final chorus. Break lines clearly within a block so it is obvious where one phrase ends and the next begins. If you want a chorus that repeats with identical wording several times, paste that exact chorus into every spot it should appear. Otherwise the model may write a new variation instead of repeating the old one.
How many attempts a usable take really needs
In practice you rarely get a finished result on the first generation. One generation returns two versions at once and costs about 5 credits, so trying different style prompts is cheaper than it sounds: the 50 free daily credits are roughly 10 attempts a day, and on Pro’s 2,500 credits a month you are counting in the hundreds.
The order that works: lock the lyrics and leave them alone, vary only the style prompt between attempts. That way it is clear what is actually changing the sound, not the text or the structure, just the phrasing of the style. A usable version usually turns up within three to five generations once genre and mood were decided ahead of time, and takes noticeably longer if the style wording stayed vague.
Stems, and what to do with them
When a track works structurally but not in balance, vocal too loud, drums too quiet under a specific scene, stems solve it. Paid tiers export up to 12 separate WAV tracks: vocal, drums, bass and so on, individually. That opens up fixes that a single mixed-down file simply does not allow.
For video work, stems matter most where music has to duck under a voiceover or synced dialogue: mute or pull the vocal track and keep only the instrumental bed, with no need to regenerate the whole track. From there it is ordinary work in the editing program, moving faders and automating volume around specific scenes.
Remixes, covers and uploading your own audio
A track does not always need to start from nothing. Suno accepts your own audio, up to 8 minutes on Free and up to 30 minutes on Premier, and can turn it into a remix or a cover. For client work that matters when someone already has a rough guitar part or a voice memo of a melody: instead of describing the idea in words, you upload the source and ask the model to rework it in the style you need.
Verse and chorus tagging still applies here: if the uploaded material already has song structure, mark where the verse and chorus sit in the source so the model does not reorder them while processing. Instrumental material with no lyrics does not need this step.
Voices, Custom Models and My Taste, and when they earn their keep
On 9 September 2026 Suno replaced its model lineup with v6, retrained on licensed music under deals with Warner Music, BMG and Believe after copyright lawsuits from major labels. Paid plans get the flagship v6 and the more unpredictable v6-wild, the free plan gets the faster v6-mini. Three features carried over from the previous release still change the workflow for regular work: Voices lets you use your own voice once it has been verified, useful when a series of videos needs one recognisable vocal rather than a new voice every generation. Custom Models let you tune the model to your own sound. My Taste adjusts generation to your history of liked tracks if you use the service regularly.
None of this matters for a one-off job, a plain style prompt and lyrics are enough there. They start paying off once music is not a one-time need but part of a recognisable sound across a series of videos, at which point Voices and Custom Models save you from rebuilding the sound from scratch on every new clip.
Where the AI stops and manual work begins
The model covers writing the music and the lyrics, the arrangement, the vocal performance. It does not cover fitting the track exactly to a video’s timing, cutting and stretching are still done by hand in the editing program, the final loudness balance against the rest of the video’s audio, and the judgement call on whether the mood actually fits the scene, which is made by a person looking at the finished cut, not just listening to the track on its own.
One thing learned from client work: generate the track once at least a rough cut exists, not before. That way you know the length you need, where the accent should land, and whether the chorus should line up with a specific shot. Music built for an existing cut almost always lands better than a cut built around music made in advance.
If you need commercial rights
If the track is for a client rather than for yourself, generate it on a paid tier from the start. The free plan carries no commercial rights at all, and upgrading afterwards does not apply retroactively to tracks you already made on it.
By this point you are probably also paying for a chat model, an image model and a video model. Adding a fifth subscription for music, and a fifth point where a card can bounce, is its own kind of tax on the work. We run our music generation through SYNTX instead, one bill instead of five, one card that actually clears. Check the catalogue for the model you need before you pay, the line-up shifts from month to month. The full breakdown of plans, downloads and rights is in the guide to downloading Suno tracks, and if the job needs vocals in a language other than English, this piece on non-English vocals covers what to expect.
Frequently asked questions
Where do I start if I have never made music with AI before?
With four decisions made before you generate: genre, tempo, mood, language. After that, write a concrete style prompt as a comma-separated list, not a general phrase like “great music”. If you need lyrics, prepare them in full beforehand and tag verse and chorus.
How many attempts does it take to get a usable track?
In practice, three to five generations once genre and mood are decided upfront, and noticeably more if the style prompt stays vague. One generation returns two versions and costs about 5 credits, so trying several variations is not expensive.
How do I tag lyrics with verse and chorus?
Mark each block before it starts: verse 1, chorus, verse 2, and so on. If the chorus should repeat with identical wording, paste it in full at every spot it appears, otherwise the model may write a new variation rather than repeat the old one.
What if the vocal covers the voiceover in my video?
Export stems, up to 12 separate WAV tracks on paid tiers, and mute or lower the vocal track, keeping the instrumental. That is faster than regenerating the whole track from scratch.
Can I get a fully finished track with zero manual work?
The model covers the shape and the sound completely. Fitting it exactly to a video’s timing, balancing it against the rest of the video’s audio, and deciding whether the mood suits a specific scene are still handled by a person in the editing program.
Do I need a paid plan for commercial use?
Yes. The free plan has no commercial rights at all, and upgrading later does not grant them retroactively to tracks already made. If the track is for a client, generate it on a paid tier from the beginning.
Need a video or an AI system built for your task?
Text: Artem Shutkin, AIVFX studio