Measured on-device · English technical report
How fast can Small speak?
Single-stream incremental PCM, synchronized batches, and offline inference on one RTX 3090. This is model throughput, not browser latency.
What “on device” means here
CUDA events bracket TTS prefill, autoregressive global/local decoding and the GPU audio codec. The end event is synchronized before reading elapsed time. This is a device event time span: it includes host launch gaps when the GPU waits, not just a sum of kernel-active durations. Prompt packing and reference encoding are completed before timing. Luna, forced alignment, scoring, CPU PCM copies, MP3 encoding, HTTP and playback are excluded. See PyTorch CUDA timing guidance.
Single stream RTF = device seconds / generated audio seconds Batch aggregate RTF = device seconds / sum(audio seconds across all rows) Per-stream RTF = device seconds / mean(audio seconds per row) Throughput = sum(audio seconds) / device seconds = 1 / aggregate RTF
RTF < 1 means faster than real time. A small aggregate RTF does not automatically mean each individual stream is real-time. We therefore report both.
Single-stream: streaming versus offline
| Mode | GPU span / 8 s audio | RTF | Audio s / GPU s | First PCM ready | Observed RTF range | Trials |
|---|---|---|---|---|---|---|
| Original offline | 7.784 s | 0.9730 | 1.03 | 7784 ms | 0.9689–1.0181 | 3 |
| Original streaming | 8.720 s | 1.0901 | 0.92 | 866 ms | 1.0681–1.1008 | 3 |
| Optimized offline | 1.586 s | 0.1982 | 5.05 | 1586 ms | 0.1969–0.1991 | 3 |
| Optimized streaming | 2.409 s | 0.3011 | 3.32 | 254 ms | 0.3010–0.3017 | 3 |
Streaming really decodes each newly generated ten-frame block into PCM using a persistent causal codec state: 10 frames / 12.5 Hz = 0.8 seconds per chunk. It does not repeatedly decode the growing prefix, and it does not split an already completed waveform. Offline waits for all tokens, then batch-decodes PCM; it saves frequent codec launches and is faster, but offers no early audio. First-PCM time is not network time or time-to-audible playback.
How small should streaming chunks be?
A separate warmed sweep measures the latency-throughput trade-off, with three trials per setting. More frequent codec calls cost substantial GPU work. Chunk sizes below are actual PCM duration, not model time-to-first-audio.
| Streams | Frames / chunk | Audio chunk ms | First PCM ms | Per-stream RTF | Aggregate RTF | Audio s / GPU s |
|---|---|---|---|---|---|---|
| 1 | 1 | 80 | 122 | 1.346 | 1.3459 | 0.74 |
| 1 | 2 | 160 | 138 | 0.771 | 0.7715 | 1.30 |
| 1 | 5 | 400 | 179 | 0.415 | 0.4154 | 2.41 |
| 1 | 10 | 800 | 251 | 0.298 | 0.2976 | 3.36 |
| 1 | 20 | 1600 | 392 | 0.239 | 0.2385 | 4.19 |
| 10 | 1 | 80 | 181 | 1.396 | 0.1396 | 7.16 |
| 10 | 2 | 160 | 198 | 0.810 | 0.0810 | 12.34 |
| 10 | 5 | 400 | 248 | 0.458 | 0.0458 | 21.83 |
| 10 | 10 | 800 | 339 | 0.351 | 0.0351 | 28.46 |
| 10 | 20 | 1600 | 532 | 0.306 | 0.0306 | 32.71 |
Parallel streaming: 1 through 10, then beyond
| Streams | Device s | Aggregate RTF | Per-stream RTF | Audio s / GPU s | First PCM ms | Worst chunk gap s¹ | Peak allocated GiB | Trials |
|---|---|---|---|---|---|---|---|---|
| 1 | 2.409 | 0.3011 | 0.301 | 3.32 | 254 | 0.240 | 5.64 | 3 |
| 2 | 2.443 | 0.1527 | 0.305 | 6.55 | 262 | 0.242 | 5.89 | 3 |
| 3 | 2.470 | 0.1029 | 0.309 | 9.72 | 269 | 0.245 | 6.05 | 3 |
| 4 | 2.484 | 0.0776 | 0.310 | 12.88 | 281 | 0.245 | 6.31 | 3 |
| 5 | 2.501 | 0.0625 | 0.313 | 16.00 | 289 | 0.246 | 6.48 | 3 |
| 6 | 2.510 | 0.0523 | 0.314 | 19.13 | 294 | 0.247 | 6.71 | 3 |
| 7 | 2.549 | 0.0455 | 0.319 | 21.97 | 300 | 0.254 | 6.96 | 3 |
| 8 | 2.628 | 0.0411 | 0.329 | 24.35 | 317 | 0.266 | 7.14 | 3 |
| 9 | 2.734 | 0.0380 | 0.342 | 26.34 | 324 | 0.281 | 7.34 | 3 |
| 10 | 2.824 | 0.0353 | 0.353 | 28.33 | 339 | 0.287 | 7.54 | 3 |
| 12 | 3.054 | 0.0318 | 0.382 | 31.44 | 370 | 0.311 | 7.95 | 1 |
| 16 | 3.426 | 0.0268 | 0.428 | 37.37 | 423 | 0.352 | 8.77 | 1 |
| 24 | 4.285 | 0.0223 | 0.536 | 44.81 | 545 | 0.443 | 10.44 | 1 |
| 32 | 5.142 | 0.0201 | 0.643 | 49.79 | 665 | 0.533 | 12.10 | 1 |
| 48 | 6.971 | 0.0182 | 0.871 | 55.09 | 912 | 0.726 | 15.40 | 1 |
| 64 | Out of memory | — | — | — | — | — | — | 0 completed |
¹ Median across trials of each trial’s maximum gap between successive ready PCM chunks. First-chunk warm-up is separate. B=1–10: three measured trials; B=12–48: one trial each, so no robust high-batch variance estimate.
Longer-cache stress check: 60 seconds per stream
A separate fixed 750-frame run checks sustained decode and memory growth at B=1 and B=12, using the same 800 ms PCM chunk and short prompt recipe. Stop decisions are overridden only for this timing workload. This is not a 60-second spoken-content quality test.
| Streams | Seconds / stream | Trial count | Device s | Per-stream RTF | Aggregate RTF | Peak allocated GiB |
|---|---|---|---|---|---|---|
| 1 | 60 | 2 | 17.727 | 0.295 | 0.2954 | 5.88 |
| 12 | 60 | 2 | 28.440 | 0.474 | 0.0395 | 11.45 |
The demo uses batches up to 12 for best-of-N, with a smaller-batch retry on GPU out-of-memory. The UI’s N limit remains 12. Successful short-prefix stress tests do not guarantee memory capacity for arbitrary long prompt prefixes; cold shape compilation can also increase latency.
What was optimized
| Incremental optimization · offline B=1 | GPU s / 8 s audio | RTF | Peak allocated GiB |
|---|---|---|---|
| Original serial CFG | 7.810 | 0.9762 | 11.63 |
| Fused conditional / neutral batch | 4.343 | 0.5429 | 11.63 |
| Also local CUDA graphs | 3.310 | 0.4137 | 11.65 |
| Also supported BF16 codec weights | 3.316 | 0.4144 | 5.64 |
| Also compiled dynamic global decoder | 1.575 | 0.1968 | 5.64 |
Single-stream CFG cost after optimization
| CFG | Mode | RTF | Device s / 8 s audio |
|---|---|---|---|
| 1 | Offline | 0.1796 | 1.437 |
| 1 | Streaming 800 ms | 0.2837 | 2.270 |
| 1.5 | Offline | 0.1982 | 1.586 |
| 1.5 | Streaming 800 ms | 0.3022 | 2.417 |
CFG=1 uses one conditional branch, CFG=1.5 two fused branches. The measured B=1 overhead is small on this GPU after fusion; this does not imply zero CFG cost at larger batches or on other hardware.
- Fuse the two CFG branches into one global/local batch while retaining shared sampled codebook prefixes, conditional-only stop decisions and the exact equation
u + g(c − u). Left padding preserves text positions and masks. - Capture the fixed 12-codebook local step in reusable CUDA graphs. Captures restore the pre-warm RNG state; graph replay advances RNG normally. The graph cache is bounded and keyed by batch, CFG and sampling parameters.
- Compile the growing-cache global decoder with dynamic shapes. This fuses small pointwise operations and reduces launch overhead without changing model weights.
- Use the codec’s supported BF16 encoder/decoder weight setter, keeping its quantizer FP32. This reduces steady GPU memory by roughly 6 GiB; it did not independently improve measured speed. See the codec inference contract.
- Batch the PCM codec instead of decoding each candidate separately. Streaming maintains persistent row-specific states and masks finished rows.
Cold costs are real: the initial compilation/warm-up in the compile experiment took 139.1 s. Local graph capture costs typically 0.25–0.5 s per new configuration. Steady measurements exclude warm-up; live UI timings include first-use work if it occurs inside generation. Compiler disk caching helps subsequent processes, but new shapes may still compile. This is not a cold-start benchmark.
Numerical and quality checks
Fused eager and local CUDA-graph code sequences were exactly equal for batches 1 and 4 in the numerical test, and across three repeated optimization seeds. Streaming versus full codec decode was tested with rows terminating at 30, 19, 7 and 25 frames, including mid-chunk termination. All output lengths were correct and PCM finite.
| Check | RMS PCM difference | Maximum absolute difference |
|---|---|---|
| Stream vs full · row 0 · 30 frames | 0.000694 | 0.018555 |
| Stream vs full · row 1 · 19 frames | 0.000493 | 0.006012 |
| Stream vs full · row 2 · 7 frames | 0.000837 | 0.010254 |
| Stream vs full · row 3 · 25 frames | 0.001646 | 0.049316 |
| FP32 codec weights vs supported BF16 | 0.000863 | 0.041992 |
Fusing CFG changes BF16 GEMM batch shape, and compiling global operations changes rounding. Stochastic trajectories therefore differ from the original code: optimization is not bit-identical to legacy. Small waveform differences are numerical evidence, not a perceptual equivalence test. The separate emotion/quality study evaluates actual free-running output with ASR and learned proxies.
Reproducibility and limitations
Three trials per main condition use seeds 777, 778 and 779; medians and observed ranges are shown, not an invented precise confidence interval. Each timing workload forces 100 frames per stream (eight seconds) to stop random early EOS from changing the workload. This is scheduling capacity, not a claim that naturally generated speech always lasts eight seconds. The quality study does not force a minimum length.
Reference codec preparation, prompt tensors and all model weights are resident before the timer. The live GPU server was stopped; no scorer or other model ran concurrently on GPU 0 during measurements. Display-server memory remains. CUDA 1 is not used to produce these throughput numbers. Warm-up uses 20 frames; the first measured longer-cache decode may still incur shape-specific work. A small chunk-size sweep and a 60-second cache stress check are included. GPU clocks were not fixed; long-duration thermal equilibrium, latency under request queuing and independent asynchronous multi-user arrivals were not measured.
{
"gpu": "NVIDIA GeForce RTX 3090",
"vram_gib": 23.68426513671875,
"torch": "2.6.0+cu124",
"cuda": "12.4",
"transformers": "5.15.0",
"python": "3.12.13",
"model": "laion/Humaneness-Voice-Small",
"revision": "5de86032771c37a3008d14e78ee7cdb39368e2a8",
"stage": "S3",
"codec_revision": "f6e20e543b33d2c252a7ef71bdf8aa71e5ff9169",
"frames": 100,
"cfg": 1.5,
"temperature": 1.0,
"reference_frames": 36
}Raw trial records: baseline.json · optimizations.json · compile.json · sweep.json · validation.json · latency.json · long.json