100 K-Pop Tracks With YuE2 in 76 Minutes: Five You Can Play, and Two Findings That Held
100 songs in 76 minutes on a 16GB RTX 5080 — 4h36m of audio at 3.63x realtime. Only 15 of 100 reached the requested 195 seconds, and the shortfall differs by style by 44 seconds. 95 of 100 clip. Five tracks embedded with the exact prompt that produced them.
TL;DR — 100 K-pop songs, five production styles, generated in 76 minutes on a 16 GB RTX 5080 — 4 hours 36 minutes of audio at 3.63× realtime. Two findings hold at n=100 that three tracks only hinted at. Only 15 of 100 reached the requested 195 seconds — mean 165.6 s, and the shortfall is systematically different per style, spanning 44 seconds between the shortest and longest. And 95 of 100 clip, above 0 dBTP. Three tracks are embedded below with the exact prompt that produced them.
This is the follow-up to running YuE2 on 16 GB and Windows. That piece established the environment on three takes. This one runs the same setup a hundred times to see which observations survive.
The setup
Identical to the previous run: YuE2-3B int8 (yue2_3b_int8_convrot, 3.69 GB) in ComfyUI on an RTX 5080 16 GB under Windows, 20 steps, cfg 3.0, ABC score stage enabled, 195 seconds requested for every song.
Five style coordinates × 20 songs each. Every song gets its own lyric; the style string is constant within a group.
| Wall clock | Audio produced | Realtime factor | |
|---|---|---|---|
| 100 songs | 76 min (21:05 → 22:21) | 276.0 min | 3.63× |
About 46 seconds per song, unattended, on a consumer card.
What a prompt actually looks like
Two strings go to the model: a style description and a lyric sheet. Here is exactly what produced the third track below, reconstructed by importing the generator rather than paraphrasing it.
Style (670 characters):
Airy minimal K-pop built on a Jersey club / UK garage bed: skipping syncopated
kick pattern, crisp finger-snap and rim clicks, a soft round sub bass, very few
chords held long, and a lot of empty space in the mix. Vocals are light, breathy
and conversational, sung almost without vibrato, layered into loose unison stacks
rather than tight power harmony, sitting close and dry with only a short room
reverb. Nostalgic and warm, slightly lo-fi, tape-soft top end, nothing shouted,
nothing overproduced. 128 to 145 BPM. Reference point for mood and tempo only:
the energy of "Attention" — minimal pulse, pencil-scratch percussion, hushed
lead. The song is about 같이 공부하던 밤.
Note the shape. Almost all of it is production description — drum pattern, bass character, chord density, vocal treatment, reverb, mix density, BPM range. One clause names a released track as a mood and tempo reference, explicitly scoped. The last sentence is the subject, in Korean.
Lyrics (818 characters), twelve sections:
Intro · Verse · Pre-Chorus · Chorus · Post-Chorus · Verse · Pre-Chorus ·
Chorus · Post-Chorus · Bridge · Chorus · Outro
[Chorus]
Stay up, 조금만 더
졸리면 어깨 빌려줄게
Stay up, 조금만 더
한 장 더 넘겨
[Post-Chorus]
Stay, stay, one more page, one more
Stay, stay, one more page, one more
All lyrics were written for this run. Korean and English are mixed line by line, the way the genre does it — English hook line, Korean answer. No lyrics or melodies from released songs were used; the one thing borrowed from a real track is its title, in the mood clause quoted above.
Three tracks
The selection rule was fixed before listening: the track whose duration is the median for its style. Not the best one — cherry-picking from a hundred would misrepresent the batch.
Three of the five styles are represented here. The measurements below cover all five and all 100 tracks; the audio examples are a subset.
These are raw model output. Nothing has been mixed, mastered, edited or trimmed, which is also why they clip.
Bubblegum — springy off-beat bass, plucky stabs, glassy bells, forward sweet vocals doubled in thirds. 118–132 BPM. 141.0 s, +1.04 dBTP
Garage-minimal — the prompt quoted in full above. Jersey club bed, breathy conversational vocals, lots of empty space. 128–145 BPM. 173.4 s, +0.21 dBTP
Glossy-anthemic — lush layered pads, strings, big melodic bass, real belting on the chorus, wide reverb with a long tail. 120–138 BPM. 189.5 s, +0.08 dBTP
These three also happen to bracket the duration finding below: the shortest-running style, the middle, and the longest. Judge the sound yourself — this article deliberately does not tell you whether they are good, for the reason in the last section.
Finding 1: the requested length is a ceiling, and the style decides how far short you fall
Every one of the hundred asked for 195 seconds.
| Style | n | Mean | Range | Reached 195 s |
|---|---|---|---|---|
| Bubblegum | 20 | 140.8 s | 109.8 – 175.6 | 0/20 |
| Hard-chant | 20 | 158.3 s | 126.0 – 193.0 | 0/20 |
| Garage-minimal | 20 | 168.8 s | 122.3 – 195.0 | 2/20 |
| Band-pop | 20 | 175.3 s | 137.2 – 195.0 | 6/20 |
| Glossy-anthemic | 20 | 184.8 s | 163.7 – 195.0 | 7/20 |
| All | 100 | 165.6 s | 109.8 – 195.0 | 15/100 |
Fifteen of a hundred hit the number they were given. Mean shortfall 29.4 seconds — 15.1%.
And it is not noise. The gap between the shortest style's mean and the longest is 44 seconds, and two of the five styles never once reached the target across twenty attempts while another managed it seven times. The style string, which says nothing about duration, is moving the duration by nearly a quarter.
A hypothesis, not a conclusion. The lyric scaffold is twelve sections of short lines. The model appears to stop when the lyric is sung rather than when the clock is met, so anything in the style that slows delivery or extends sections buys length. That fits the ordering: the two longest styles are described with big builds, long reverb tails and belting; the shortest says "the arrangement stays clean and uncluttered."
Testing it properly means holding the lyric constant and varying only the style, which has not been run. Until then this is a pattern with a plausible mechanism, not a demonstrated cause.
The practical consequence stands regardless: if you need a specific runtime, ask for more than you need and trim, or expect to be about 15% short on average and up to 44% short in the worst case observed.
Finding 2: almost everything clips
| Style | Mean true peak | Max | Over 0 dBTP |
|---|---|---|---|
| Bubblegum | — | +1.04 | 19/20 |
| Hard-chant | — | +1.33 | 20/20 |
| Garage-minimal | — | +1.36 | 20/20 |
| Band-pop | — | +0.86 | 18/20 |
| Glossy-anthemic | — | +1.24 | 18/20 |
| All 100 | +0.44 dBTP | +1.36 | 95/100 (95%) |
Three out of three clipped in the previous test. At a hundred it is 95%. This is not a sampling artefact — it is what the model outputs.
Nothing here has been mastered, so this is inter-sample clipping in raw output. Run a limiter or a loudness pass before using any of it. Put these straight into a video timeline and you ship distortion.
Finding 3: the loudness target is remarkably consistent
| Mean | Range | |
|---|---|---|
| Integrated loudness | −12.97 LUFS | −15.24 to −11.17 |
| Loudness range (LRA) | 5.4 LU | 2.7 – 10.1 |
A four-decibel spread across a hundred songs in five different styles is tight. The model is evidently targeting something close to −13 LUFS, which sits between streaming normalisation targets (−14) and a loud commercial master (−9 to −11).
The LRA figure is the more telling one. A mean of 5.4 LU is heavily compressed — squarely in modern pop-master territory rather than anything dynamic. Whatever else the model learned, it learned to sound mastered. Which is also why it clips: it is pushing for loudness with no limiter at the end of the chain.
What this run does not establish
Whether any of it is good. I measured a hundred files and did not listen to them. There is no verdict here on melody, vocal quality, lyric intelligibility, or whether any of these would survive thirty seconds with a listener who likes the genre. The players are above precisely so you can form that judgement yourself instead of taking mine.
Whether the anchor clause does anything. Each prompt names a released track for mood and tempo. Whether the model has any representation of that title, or simply ignores it and works from the production description around it, is untested. Removing the clause and re-running would answer it.
Whether the duration effect is caused by the style or correlated with it. No controlled run — same lyric, different style — has been done.
How int8 compares with full precision. Still untested, as before. Everything here is one quantised build.
Licensing. YuE2's weights are CC BY-NC 4.0. Non-commercial. A hundred tracks is a hundred tracks you may not sell.
Method notes, for anyone reproducing this
- Fix your selection rule before you listen. Median duration per group, decided in advance. A hundred outputs contains a good one by chance, and publishing that one tells the reader nothing.
- Measure the whole batch, not the samples you like. The clipping rate only becomes a fact at n=100; at n=3 it was an anecdote.
- Keep the request constant. Every song asked for 195 seconds, which is what makes the per-style duration differences readable.
- The measurement pass is ffprobe for duration and ffmpeg's
loudnormanalysis pass for integrated loudness, true peak and LRA. Both are a few lines and worth building before you generate, not after.
FAQ
How long does it take to generate 100 songs with YuE2?
76 minutes on an RTX 5080 16 GB under Windows, producing 276 minutes of audio — about 46 seconds per song at 3.63× realtime. That is the int8 ComfyUI build at 20 steps, cfg 3.0, with the ABC score stage on.
Why are my YuE2 songs shorter than I asked for?
Because the requested length is a ceiling, not a target. Across 100 tracks asking for 195 seconds, the mean was 165.6 s and only 15 reached the full length. The shortfall also varies by style — two of five styles never reached the target in twenty attempts.
Does YuE2 output clip?
Yes, almost always. 95 of 100 tracks exceeded 0 dBTP, up to +1.36. The output is loud and unlimited, so apply a limiter or loudness normalisation before using it anywhere.
What loudness does YuE2 target?
About −13 LUFS integrated, with a spread of only four decibels across a hundred tracks in five styles. Loudness range averages 5.4 LU, which is modern-pop-master compressed rather than dynamic.
How do you prompt YuE2 for a specific K-pop style?
Describe production, not artists: drum pattern, bass character, chord density, vocal treatment and layering, reverb, mix density, and a BPM range. In this run that description is roughly 600 characters and does the bulk of the work. Section tags in the lyrics — [Verse], [Pre-Chorus], [Chorus], [Post-Chorus], [Bridge] — control the arrangement.
Can I sell music made with YuE2?
No. The weights are CC BY-NC 4.0, which is non-commercial, regardless of the fact that the code is Apache-2.0 and the lyrics are your own.
← Back to all posts