Suno can generate songs quickly, but whether the result meets a project’s requirements depends on the definitions established before generation and the selection that follows it. This article documents how I use Suno to produce themes for game characters. The process begins by defining the character and narrative purpose, then translating them into musical constraints that the model can execute.
Establishing the musical direction from the character
I begin with the character, not the genre. First, I assign the character a dedicated instrument that remains present throughout the song. A violin, for example, can carry one character’s internal voice from beginning to end. I then define an emotional arc and allow that arc to remain unresolved. For a character who cannot remember someone, the ending should not convert the question into a definite answer. An incomplete cadence is more consistent with that state.
Another available structure is a mirrored answer song: two songs present the same relationship from different viewpoints, with the lyrics forming cross-references. The character and narrative must be established first; genre then becomes the outer structure that carries them.
Six components of a Style prompt
Suno’s Style field accepts approximately 1,000 characters and acts as the song’s primary constraint. I divide it into six components instead of accumulating abstract adjectives:
- Anchor parameters: establish tempo, tonality, and overall mood, such as
68 BPM, minor key, tender sorrow that never resolves; - Instrumentation: specify instruments and entry points, and freeze changes that are not wanted, such as
celesta + sparse piano, warm upright bass, brushed snare after the first chorus, no fills; - Vocal specification: state perceived age, breathiness, vibrato, and placement against the beat, such as
youthful breathy female voice, minimal vibrato, slightly behind the beat; - Harmonic constraints: define extended chords and cadence behaviour, such as
extended 7th/9th chords, deceptive cadences, never fully resolves; - Dynamic principle: require tension to grow through texture rather than volume, such as
tension grows through texture, not volume; the final chorus is the quietest; - Bracketed negative list: limit the directions in which the model most often deviates, such as
[no belting, no soaring, no bright pop, no autotune].
The sixth component addresses Suno’s tendency toward high-register display, sweet idol-style timbres, and audible pitch correction. The negative list should target likely deviations rather than expand without limit.
Vocal and harmonic constraints
AI-generated music often defaults to catchy, triumphant, and bright pop expression. I use the following constraints to preserve the restraint required by the character:
- Control vocal dynamics: the song uses no belting and no climax driven solely by volume. Emotion is established through close-microphone presence, and the ending diminishes to an almost whispered level;
- Keep the harmony suspended: use extended chords (maj7/9/11), modal interchange, deceptive cadences, and a final cadence that does not fully resolve. Avoid standard pop sequences such as 4536, 1645, and the Pachelbel progression so that the harmony retains a sense of incomplete arrival;
- Handle tonal movement between verse and chorus carefully: the chorus may move toward another tonal centre to create a displaced layer of memory, but the prompt should specify
a subtle dreamlike non-diatonic shift, never a triumphant liftso that the model does not interpret “key change” as a dramatic upward modulation.
For the intimate and languid vocal direction and chamber texture, my primary reference is Sheena Ringo’s “Shiseikatsu”. For a gentle narrative arc, I have also referenced Zhou Shen’s “化身孤岛的鲸”. The vocal design deliberately avoids a sweet idol-style timbre.
Four Style starting templates
The following templates correspond to different scenes. Their subjects and details can be replaced while the underlying structure remains intact:
- A · Intimate dark ballad (character interior / rest)
72 BPM, minor, intimate chamber pop, breathy young female voice, celesta + sparse piano + warm upright bass, jazz extended chords, deceptive cadences, tension through texture not volume. [no belting, no bright pop, no autotune] - B · Epic orchestral (boss / final battle)
epic cinematic orchestral, minor, driving low strings + brass swells + taiko + choir, relentless build, wide dynamics. [no electronic drums, no cheesy fanfare, no early major-key triumph] - C · Darkwave ambience (exploration / long night)
downtempo darkwave trip-hop, 85 BPM, dusty broken beat, deep sub bass, detuned pads, distant reverbed murmur, film-noir melancholy. [no EDM drop, no bright synth lead] - D · Bardic folk (town / bonfire)
warm acoustic folk ballad, fingerpicked lute or guitar, upright bass, soft brushed percussion, modal folk harmony, candle-lit tavern warmth. [no modern pop production, no big drums, no autotune]
Each template retains a bracketed negative list. Its function is to exclude common deviations in advance and to operate alongside the positive Style description as part of the initial constraint.
Structural design
- Circular structure: use the same hummed phrase in the intro and outro, and add
[Intro: use the exact arrangement of the outro]to the Style prompt so that the song returns to its point of origin; - Suspended ending: do not provide a definite answer at the end. Allow the question to return more quietly before the song closes. For certain characters, incompletion is more consistent with the narrative state than full resolution.
Suno parameter settings
The current settings are Female / Style Influence 70 / Weirdness 30. For an instruction whose effect cannot be predicted reliably, such as movement to another tonal centre, I generate versions both with and without it and compare them as an A/B pair. I may apply the same procedure to genre by generating, for example, a chamber version and a darkwave version before selecting the one that better fits the character and scene.
Common problems and responses
- The violin moves into an excessively high register: add
mid-low register only, never shrillto prevent a sharp high-register performance; - Facial features change when cover art is enlarged: use a non-generative interpolation method such as Lanczos when fidelity matters, so that the face is not reinterpreted by generative upscaling;
- Chinese lyric quality is inconsistent: Suno handles English lyrics reasonably well, but I continue to write the Chinese lyrics myself to avoid extensive revision afterward;
- The negative list weakens the primary style: retain only the four to six most likely deviations. Too many exclusions dilute the positive Style information.
Conclusion
Suno shifts part of composition from writing notes directly to making a sequence of design decisions: defining the character, emotion, instrumentation, dynamics, and every result that must be excluded. The tool provides a large generation space; a distinctive final work depends on precise constraints and a stable selection standard.