LAION · Humaneness Voice Small S3 · English research reports

Voice acting, measured.

Two studies explore the trade-offs between fast audio generation, faithful words and emotional delivery. Read the methodology, inspect the measurements, and listen to the unfiltered synthetic takes.

GPU performance

Single-stream offline and incremental PCM generation, parallel streams, batch scaling, first-chunk latency and optimization checks on an RTX 3090.

On-device real-time factor, not end-to-end latency. Warm and first-use costs are distinguished.

Read throughput study →

Emotion × CFG × temperature

660 scored takes: five emotions, two prompt formats, a sampling grid, German/reference-sensitivity checks, and concise structured-cue follow-ups.

Filter and listen to every take, before and after alignment/fades.

Read emotion study →

Recommended starting point

Concise CAPTION: direction plus an exact quoted TRANSCRIPT:; temperature 1.0 and CFG 1.5, with CFG 2.0 as an alternative to audition. Supply a short, clean, distinct speech reference with consent. These are cautious demo defaults, not a universal optimum.

Naturalness and emotion are learned proxies, not human listener ratings. The small prompt set and non-neutral reference provenance limit generalization; the second reference was a hiss-labelled corpus clip containing speech. Read the full caveats before extrapolating.