Measured on-device · English technical report

How fast can Small speak?

Single-stream incremental PCM, synchronized batches, and offline inference on one RTX 3090. This is model throughput, not browser latency.

0.301Single-stream streaming RTF · lower is faster
0.198Single-stream offline RTF
28.3×Aggregate audio seconds per GPU second · B=10
48Largest tested successful batch · 64 failed
For eight seconds of audio, optimized offline inference takes 1.59 s, versus 7.78 s before optimization (4.9× faster). Incremental streaming takes 2.41 s, versus 8.72 s (3.6× faster). B=10 produces 80 audio seconds in 2.82 s.

What “on device” means here

CUDA events bracket TTS prefill, autoregressive global/local decoding and the GPU audio codec. The end event is synchronized before reading elapsed time. This is a device event time span: it includes host launch gaps when the GPU waits, not just a sum of kernel-active durations. Prompt packing and reference encoding are completed before timing. Luna, forced alignment, scoring, CPU PCM copies, MP3 encoding, HTTP and playback are excluded. See PyTorch CUDA timing guidance.

Single stream RTF = device seconds / generated audio seconds
Batch aggregate RTF = device seconds / sum(audio seconds across all rows)
Per-stream RTF = device seconds / mean(audio seconds per row)
Throughput = sum(audio seconds) / device seconds = 1 / aggregate RTF

RTF < 1 means faster than real time. A small aggregate RTF does not automatically mean each individual stream is real-time. We therefore report both.

Single-stream: streaming versus offline

ModeGPU span / 8 s audioRTFAudio s / GPU sFirst PCM readyObserved RTF rangeTrials
Original offline7.784 s0.97301.037784 ms0.9689–1.01813
Original streaming8.720 s1.09010.92866 ms1.0681–1.10083
Optimized offline1.586 s0.19825.051586 ms0.1969–0.19913
Optimized streaming2.409 s0.30113.32254 ms0.3010–0.30173

Streaming really decodes each newly generated ten-frame block into PCM using a persistent causal codec state: 10 frames / 12.5 Hz = 0.8 seconds per chunk. It does not repeatedly decode the growing prefix, and it does not split an already completed waveform. Offline waits for all tokens, then batch-decodes PCM; it saves frequent codec launches and is faster, but offers no early audio. First-PCM time is not network time or time-to-audible playback.

How small should streaming chunks be?

A separate warmed sweep measures the latency-throughput trade-off, with three trials per setting. More frequent codec calls cost substantial GPU work. Chunk sizes below are actual PCM duration, not model time-to-first-audio.

StreamsFrames / chunkAudio chunk msFirst PCM msPer-stream RTFAggregate RTFAudio s / GPU s
11801221.3461.34590.74
121601380.7710.77151.30
154001790.4150.41542.41
1108002510.2980.29763.36
12016003920.2390.23854.19
101801811.3960.13967.16
1021601980.8100.081012.34
1054002480.4580.045821.83
10108003390.3510.035128.46
102016005320.3060.030632.71
A useful low-latency compromise is 5 frames / 400 ms: first PCM around 179 ms at B=1, RTF around 0.415. The 10-frame / 800 ms default improves throughput (RTF about 0.30), at roughly 251 ms to first PCM. 20 frames / 1.6 s is faster still (RTF about 0.239) but delays the first PCM to about 392 ms. One-frame / 80 ms chunks are slower than real time (RTF about 1.35), despite early first PCM. The UI allows 160, 400, 800 and 1600 ms chunks. Offline remains fastest when early playback is unnecessary.

Parallel streaming: 1 through 10, then beyond

0.0015.8431.6747.5163.35B=1: 3.32111B=2: 6.54992B=3: 9.71843B=4: 12.88284B=5: 15.99615B=6: 19.12676B=7: 21.96927B=8: 24.34968B=9: 26.33809B=10: 28.326410B=12: 31.437812B=16: 37.365316B=24: 44.812324B=32: 49.788532B=48: 55.086148Aggregate throughput · audio s / device sParallel streams in one synchronized batch
StreamsDevice sAggregate RTFPer-stream RTFAudio s / GPU sFirst PCM msWorst chunk gap s¹Peak allocated GiBTrials
12.4090.30110.3013.322540.2405.643
22.4430.15270.3056.552620.2425.893
32.4700.10290.3099.722690.2456.053
42.4840.07760.31012.882810.2456.313
52.5010.06250.31316.002890.2466.483
62.5100.05230.31419.132940.2476.713
72.5490.04550.31921.973000.2546.963
82.6280.04110.32924.353170.2667.143
92.7340.03800.34226.343240.2817.343
102.8240.03530.35328.333390.2877.543
123.0540.03180.38231.443700.3117.951
163.4260.02680.42837.374230.3528.771
244.2850.02230.53644.815450.44310.441
325.1420.02010.64349.796650.53312.101
486.9710.01820.87155.099120.72615.401
64Out of memory——————0 completed

¹ Median across trials of each trial’s maximum gap between successive ready PCM chunks. First-chunk warm-up is separate. B=1–10: three measured trials; B=12–48: one trial each, so no robust high-batch variance estimate.

0.000.290.570.861.15Real-time boundaryB=1: 0.30111B=2: 0.30532B=3: 0.30873B=4: 0.31054B=5: 0.31265B=6: 0.31376B=7: 0.31867B=8: 0.32858B=9: 0.34179B=10: 0.353010B=12: 0.381712B=16: 0.428216B=24: 0.535624B=32: 0.642732B=48: 0.871448Per-stream RTF · lower is fasterParallel streams in one synchronized batch
Capacity is workload-specific. B=48 succeeded with per-stream RTF 0.871 and peak allocated memory 15.40 GiB. B=64 failed during the attempted run. The exact maximum in between was not determined. These are synchronized eight-second streams with short, identical prompt prefixes; arbitrary arrivals, long context, long references, variable stop lengths and a serving scheduler are not covered. The demo still serializes browser jobs; batched best-of-N is not independent multi-user live streaming.

Longer-cache stress check: 60 seconds per stream

A separate fixed 750-frame run checks sustained decode and memory growth at B=1 and B=12, using the same 800 ms PCM chunk and short prompt recipe. Stop decisions are overridden only for this timing workload. This is not a 60-second spoken-content quality test.

StreamsSeconds / streamTrial countDevice sPer-stream RTFAggregate RTFPeak allocated GiB
160217.7270.2950.29545.88
1260228.4400.4740.039511.45

The demo uses batches up to 12 for best-of-N, with a smaller-batch retry on GPU out-of-memory. The UI’s N limit remains 12. Successful short-prefix stress tests do not guarantee memory capacity for arbitrary long prompt prefixes; cold shape compilation can also increase latency.

What was optimized

Incremental optimization · offline B=1GPU s / 8 s audioRTFPeak allocated GiB
Original serial CFG7.8100.976211.63
Fused conditional / neutral batch4.3430.542911.63
Also local CUDA graphs3.3100.413711.65
Also supported BF16 codec weights3.3160.41445.64
Also compiled dynamic global decoder1.5750.19685.64

Single-stream CFG cost after optimization

CFGModeRTFDevice s / 8 s audio
1Offline0.17961.437
1Streaming 800 ms0.28372.270
1.5Offline0.19821.586
1.5Streaming 800 ms0.30222.417

CFG=1 uses one conditional branch, CFG=1.5 two fused branches. The measured B=1 overhead is small on this GPU after fusion; this does not imply zero CFG cost at larger batches or on other hardware.

Cold costs are real: the initial compilation/warm-up in the compile experiment took 139.1 s. Local graph capture costs typically 0.25–0.5 s per new configuration. Steady measurements exclude warm-up; live UI timings include first-use work if it occurs inside generation. Compiler disk caching helps subsequent processes, but new shapes may still compile. This is not a cold-start benchmark.

Numerical and quality checks

Fused eager and local CUDA-graph code sequences were exactly equal for batches 1 and 4 in the numerical test, and across three repeated optimization seeds. Streaming versus full codec decode was tested with rows terminating at 30, 19, 7 and 25 frames, including mid-chunk termination. All output lengths were correct and PCM finite.

CheckRMS PCM differenceMaximum absolute difference
Stream vs full · row 0 · 30 frames0.0006940.018555
Stream vs full · row 1 · 19 frames0.0004930.006012
Stream vs full · row 2 · 7 frames0.0008370.010254
Stream vs full · row 3 · 25 frames0.0016460.049316
FP32 codec weights vs supported BF160.0008630.041992

Fusing CFG changes BF16 GEMM batch shape, and compiling global operations changes rounding. Stochastic trajectories therefore differ from the original code: optimization is not bit-identical to legacy. Small waveform differences are numerical evidence, not a perceptual equivalence test. The separate emotion/quality study evaluates actual free-running output with ASR and learned proxies.

Reproducibility and limitations

Three trials per main condition use seeds 777, 778 and 779; medians and observed ranges are shown, not an invented precise confidence interval. Each timing workload forces 100 frames per stream (eight seconds) to stop random early EOS from changing the workload. This is scheduling capacity, not a claim that naturally generated speech always lasts eight seconds. The quality study does not force a minimum length.

Reference codec preparation, prompt tensors and all model weights are resident before the timer. The live GPU server was stopped; no scorer or other model ran concurrently on GPU 0 during measurements. Display-server memory remains. CUDA 1 is not used to produce these throughput numbers. Warm-up uses 20 frames; the first measured longer-cache decode may still incur shape-specific work. A small chunk-size sweep and a 60-second cache stress check are included. GPU clocks were not fixed; long-duration thermal equilibrium, latency under request queuing and independent asynchronous multi-user arrivals were not measured.

{
  "gpu": "NVIDIA GeForce RTX 3090",
  "vram_gib": 23.68426513671875,
  "torch": "2.6.0+cu124",
  "cuda": "12.4",
  "transformers": "5.15.0",
  "python": "3.12.13",
  "model": "laion/Humaneness-Voice-Small",
  "revision": "5de86032771c37a3008d14e78ee7cdb39368e2a8",
  "stage": "S3",
  "codec_revision": "f6e20e543b33d2c252a7ef71bdf8aa71e5ff9169",
  "frames": 100,
  "cfg": 1.5,
  "temperature": 1.0,
  "reference_frames": 36
}

Raw trial records: baseline.json · optimizations.json · compile.json · sweep.json · validation.json · latency.json · long.json