All guides
Making Music with AI
26 August 2026 · 6 min read
AI music generation rewards musical vocabulary more than any other kind of prompting. You do not need to be a producer, but knowing how to name a tempo, a texture or a song section changes results dramatically, because those terms map onto real distinctions in the training data.
Genre is the highest-leverage word
Genre carries an enormous amount of implicit information: instrumentation, tempo range, production style, song structure and mood all come bundled with it. Naming a genre precisely does more work than a paragraph of description.
Be as specific as you can. Ambient is a broad target; dark ambient with granular textures is a narrow one. Rock covers decades; seventies-style hard rock with overdriven guitars and live drums does not. The narrower the genre reference, the more predictable the result.
Specify tempo and energy separately
Tempo and energy are related but distinct, and conflating them is a common mistake. A slow track can be intense, and a fast one can be relaxed. If you want something specific, say both: a slow tempo with a heavy, driving feel, or an up-tempo track with a light, airy character.
Giving an approximate tempo in beats per minute works well when you have a target in mind, particularly if the music has to sit against video of a known length.
- Name the genre as narrowly as you can.
- Give tempo separately from energy or mood.
- List the instruments that should carry the track.
- Say what the track is for, since function implies arrangement.
Instrumental or with lyrics
Decide this before you write the prompt, because it changes everything else. Instrumental tracks are more reliable and better suited to backing music, since there is no vocal to sit awkwardly against dialogue or narration.
If you want lyrics, supplying your own generally beats asking the model to invent them. Model-written lyrics tend towards generic imagery, while your own words give the track a specific point of view. Providing lyrics also lets you control song structure directly, since the sections you write become the sections the model builds.
Think about structure
Songs have parts, and naming them helps. An intro, a verse, a chorus and an outro is a structure the model can follow. Without any structural instruction you often get something that develops pleasantly but never resolves, which is fine as background and unsatisfying as a song.
For music intended to sit under video or narration, ask explicitly for a track that stays in one mood without dramatic changes. Generated music often introduces a build or a key change that fights whatever it is scoring.
Generate several and choose
Variance between generations from the same prompt is high in music, higher than in images. Producing several versions and picking is a better strategy than refining a prompt against a single output, because you may simply have drawn an unrepresentative sample.
When something is close but not right, identify which single element is wrong, the drums, the tempo, the vocal character, and address that alone in the next attempt.