Controlled pilot · English research report
Emotion without losing the words
A measured CFG × temperature grid for Humaneness Voice Small S3: emotional direction following, intelligibility and naturalness proxies. There is no universal optimum; the goal is to identify useful compromises under explicit constraints.
Practical conclusions
Keep the UI’s requested defaults (CFG 1.5, T=1) unchanged while auditioning the recommended alternatives. If intelligibility is essential, prefer the concise caption surface and a clean short reference. Do not infer that higher CFG must improve emotion: it can also produce repetition, unintelligible speech or flatter expression.
Design saved before scoring
The main grid contains 320 takes: CFG {1, 1.5, 2, 3} × temperature {0, 0.5, 1, 1.2} × five emotions × two prompt surfaces × two seeds. Every condition uses the same English words and reference A. Configuration order was shuffled with a fixed seed; row ordering and batch position are kept fixed. All takes, including high-WER and frame-limit cases, contribute to the means. No best-of-N winner filtering is applied.
You came back. I thought I would never see you again.
Anger, fear, sadness, elation and affection use distinct actor-style descriptions, deliberately avoiding an explicit laugh, scream or sob which could confound speech errors. The structured surface requests one 0.35-second pause; caption uses the same acting direction and exact quoted transcript without the timed cue. Prompt-format comparisons therefore evaluate these complete practical recipes, not an isolated causal effect of label syntax.
The transcript is never rewritten by Luna during the study. Sampling uses top-k 50, top-p 0.95 and repetition penalty 1.0. Audio T=0 is true argmax and also makes stop/continue greedy; at positive audio temperatures stop temperature remains 1.0. Thus the T=0 contrast includes the demo’s associated stop-policy change, not only audio entropy. T=0 performs poorly on this fixed study recipe, not on every possible prompt/reference: a separate German demo smoke test at T=0 produced zero measured WER with different words and reference. CFG affects only audio logits, using an affect-neutral comparison retaining words, reference and requested pause. Generation uses natural EOS with a 5.5-second frame hint plus a 60-frame safety budget, not forced benchmark lengths.
Predeclared preference rule
Within each prompt format: eligible if mean RAW conventional WER <= 0.10 and mean processed genuineness >= same-format CFG=1.5,T=1 baseline minus 0.25. Maximize mean target-emotion rank percentile; ties within 0.01 prefer higher direction CLAP. If no eligible configuration, explicitly report no qualifying optimum and present Pareto trade-offs.
We also mark the Pareto frontier: no other tested configuration simultaneously has lower raw WER, higher genuineness and higher emotion-rank percentile, with at least one strictly better. Changing thresholds or the priority assigned to emotion can change the recommendation.
CFG × temperature maps
WER is measured on untrimmed PCM. Genuineness and emotion describe the processed audio heard in the demo. Green WER ≤10%, yellow ≤30%, red >30%. A green cell alone is not a three-objective win.
Caption / exact transcript
| T = 0 | T = 0.5 | T = 1 | T = 1.2 | |
|---|---|---|---|---|
| CFG 1 | WER 78.2%Genuine 0.77 / 6 Emotion percentile 0.62 outside constraints | WER 47.3%Genuine 1.42 / 6 Emotion percentile 0.62 outside constraints | WER 2.7%Genuine 1.60 / 6 Emotion percentile 0.55 ✓ eligible | WER 4.5%Genuine 1.39 / 6 Emotion percentile 0.55 ✓ eligible |
| CFG 1.5 | WER 80.0%Genuine 0.76 / 6 Emotion percentile 0.47 outside constraints | WER 35.5%Genuine 1.47 / 6 Emotion percentile 0.58 outside constraints | WER 0.9%Genuine 1.64 / 6 Emotion percentile 0.60 ✓ eligible · Pareto | WER 0.0%Genuine 1.35 / 6 Emotion percentile 0.58 outside constraints · Pareto |
| CFG 2 | WER 80.0%Genuine 0.99 / 6 Emotion percentile 0.60 outside constraints | WER 28.2%Genuine 1.52 / 6 Emotion percentile 0.64 outside constraints | WER 3.6%Genuine 1.62 / 6 Emotion percentile 0.59 ✓ eligible | WER 7.3%Genuine 1.54 / 6 Emotion percentile 0.64 ✓ eligible · Pareto |
| CFG 3 | WER 92.7%Genuine 1.37 / 6 Emotion percentile 0.48 outside constraints | WER 31.8%Genuine 1.44 / 6 Emotion percentile 0.57 outside constraints | WER 13.6%Genuine 1.56 / 6 Emotion percentile 0.74 outside constraints · Pareto | WER 10.9%Genuine 1.76 / 6 Emotion percentile 0.65 outside constraints · Pareto |
Structured acting script
| T = 0 | T = 0.5 | T = 1 | T = 1.2 | |
|---|---|---|---|---|
| CFG 1 | WER 100.0%Genuine 1.10 / 6 Emotion percentile 0.71 outside constraints · Pareto | WER 100.0%Genuine 1.35 / 6 Emotion percentile 0.58 outside constraints | WER 101.8%Genuine 1.50 / 6 Emotion percentile 0.62 outside constraints · Pareto | WER 88.2%Genuine 1.64 / 6 Emotion percentile 0.61 outside constraints · Pareto |
| CFG 1.5 | WER 98.2%Genuine 1.06 / 6 Emotion percentile 0.47 outside constraints | WER 128.2%Genuine 1.11 / 6 Emotion percentile 0.49 outside constraints | WER 127.3%Genuine 1.39 / 6 Emotion percentile 0.76 outside constraints · Pareto | WER 130.9%Genuine 1.25 / 6 Emotion percentile 0.64 outside constraints |
| CFG 2 | WER 141.8%Genuine 1.06 / 6 Emotion percentile 0.79 outside constraints · Pareto | WER 130.0%Genuine 1.14 / 6 Emotion percentile 0.65 outside constraints | WER 123.6%Genuine 1.31 / 6 Emotion percentile 0.67 outside constraints · Pareto | WER 168.2%Genuine 1.35 / 6 Emotion percentile 0.70 outside constraints |
| CFG 3 | WER 160.0%Genuine 0.79 / 6 Emotion percentile 0.79 outside constraints | WER 170.9%Genuine 0.98 / 6 Emotion percentile 0.77 outside constraints | WER 200.9%Genuine 0.85 / 6 Emotion percentile 0.76 outside constraints | WER 160.9%Genuine 1.06 / 6 Emotion percentile 0.83 outside constraints · Pareto |
All caption configurations and descriptive 95% cluster-bootstrap WER intervals
| CFG | T | n | Raw WER % [CI] | Post WER % | WER ≤10% | WER >30% | Genuine /6 | Emotion percentile | Target top 5 | CLAP | Frontier |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 0 | 10 | 78.2 [72.7, 89.1] | 78.2 | 0% | 100% | 0.77 | 0.621 | 40% | 0.051 | |
| 1 | 0.5 | 10 | 47.3 [29.1, 67.3] | 47.3 | 40% | 60% | 1.42 | 0.615 | 30% | 0.097 | |
| 1 | 1 | 10 | 2.7 [0.0, 6.4] | 0.0 | 100% | 0% | 1.60 | 0.551 | 10% | 0.137 | |
| 1 | 1.2 | 10 | 4.5 [0.0, 11.8] | 0.9 | 90% | 0% | 1.39 | 0.549 | 30% | 0.140 | |
| 1.5 | 0 | 10 | 80.0 [72.7, 87.3] | 80.0 | 0% | 100% | 0.76 | 0.474 | 0% | 0.063 | |
| 1.5 | 0.5 | 10 | 35.5 [10.0, 60.0] | 35.5 | 60% | 40% | 1.47 | 0.579 | 20% | 0.120 | |
| 1.5 | 1 | 10 | 0.9 [0.0, 2.7] | 0.0 | 100% | 0% | 1.64 | 0.597 | 30% | 0.170 | Pareto |
| 1.5 | 1.2 | 10 | 0.0 [0.0, 0.0] | 0.0 | 100% | 0% | 1.35 | 0.579 | 30% | 0.173 | Pareto |
| 2 | 0 | 10 | 80.0 [72.7, 87.3] | 80.0 | 0% | 100% | 0.99 | 0.600 | 20% | 0.050 | |
| 2 | 0.5 | 10 | 28.2 [11.8, 43.6] | 28.2 | 60% | 30% | 1.52 | 0.644 | 30% | 0.149 | |
| 2 | 1 | 10 | 3.6 [0.0, 7.3] | 3.6 | 80% | 0% | 1.62 | 0.592 | 0% | 0.181 | |
| 2 | 1.2 | 10 | 7.3 [0.0, 16.4] | 5.5 | 80% | 10% | 1.54 | 0.636 | 30% | 0.217 | Pareto |
| 3 | 0 | 10 | 92.7 [90.9, 96.4] | 92.7 | 0% | 100% | 1.37 | 0.482 | 0% | 0.047 | |
| 3 | 0.5 | 10 | 31.8 [1.8, 70.0] | 31.8 | 60% | 30% | 1.44 | 0.567 | 20% | 0.162 | |
| 3 | 1 | 10 | 13.6 [0.0, 39.1] | 13.6 | 80% | 10% | 1.56 | 0.741 | 40% | 0.195 | Pareto |
| 3 | 1.2 | 10 | 10.9 [2.7, 23.6] | 10.0 | 80% | 10% | 1.76 | 0.651 | 30% | 0.203 | Pareto |
All structured configurations and descriptive 95% cluster-bootstrap WER intervals
| CFG | T | n | Raw WER % [CI] | Post WER % | WER ≤10% | WER >30% | Genuine /6 | Emotion percentile | Target top 5 | CLAP | Frontier |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 1 | 0 | 10 | 100.0 [100.0, 100.0] | 100.0 | 0% | 100% | 1.10 | 0.708 | 40% | 0.115 | Pareto |
| 1 | 0.5 | 10 | 100.0 [100.0, 100.0] | 100.0 | 0% | 100% | 1.35 | 0.585 | 20% | 0.172 | |
| 1 | 1 | 10 | 101.8 [95.5, 108.2] | 100.9 | 0% | 100% | 1.50 | 0.615 | 50% | 0.146 | Pareto |
| 1 | 1.2 | 10 | 88.2 [78.2, 98.2] | 88.2 | 0% | 100% | 1.64 | 0.608 | 20% | 0.148 | Pareto |
| 1.5 | 0 | 10 | 98.2 [72.7, 121.8] | 90.9 | 0% | 100% | 1.06 | 0.469 | 20% | 0.069 | |
| 1.5 | 0.5 | 10 | 128.2 [100.0, 166.4] | 100.9 | 0% | 100% | 1.11 | 0.492 | 20% | 0.096 | |
| 1.5 | 1 | 10 | 127.3 [105.5, 150.9] | 98.2 | 0% | 100% | 1.39 | 0.762 | 40% | 0.144 | Pareto |
| 1.5 | 1.2 | 10 | 130.9 [109.1, 150.9] | 103.6 | 0% | 100% | 1.25 | 0.636 | 20% | 0.153 | |
| 2 | 0 | 10 | 141.8 [92.7, 190.9] | 89.1 | 0% | 100% | 1.06 | 0.795 | 40% | 0.119 | Pareto |
| 2 | 0.5 | 10 | 130.0 [85.5, 171.8] | 92.7 | 0% | 90% | 1.14 | 0.651 | 30% | 0.084 | |
| 2 | 1 | 10 | 123.6 [107.3, 146.4] | 103.6 | 0% | 100% | 1.31 | 0.672 | 30% | 0.113 | Pareto |
| 2 | 1.2 | 10 | 168.2 [118.2, 218.2] | 101.8 | 0% | 100% | 1.35 | 0.700 | 30% | 0.092 | |
| 3 | 0 | 10 | 160.0 [112.7, 207.3] | 96.4 | 0% | 100% | 0.79 | 0.790 | 60% | 0.049 | |
| 3 | 0.5 | 10 | 170.9 [102.7, 240.9] | 109.1 | 0% | 90% | 0.98 | 0.767 | 40% | 0.090 | |
| 3 | 1 | 10 | 200.9 [170.9, 240.9] | 104.5 | 0% | 100% | 0.85 | 0.762 | 60% | 0.077 | |
| 3 | 1.2 | 10 | 160.9 [120.9, 198.2] | 96.4 | 0% | 90% | 1.06 | 0.833 | 70% | 0.115 | Pareto |
How emotion-specific are the compromises?
| Emotion | Format | CFG / T | Raw WER | Genuine | Emotion percentile | Exploratory rule |
|---|---|---|---|---|---|---|
| Anger | caption | 3 / 0.5 | 0.0% | 1.39 | 0.77 | Emotion max among WER≤10% |
| Anger | structured | 1 / 1.2 | 72.7% | 1.61 | 0.24 | No low-WER cell; lowest WER |
| Fear | caption | 1.5 / 1 | 0.0% | 1.71 | 0.53 | Emotion max among WER≤10% |
| Fear | structured | 2 / 0 | 81.8% | 0.00 | 0.54 | No low-WER cell; lowest WER |
| Sadness | caption | 1 / 1.2 | 0.0% | 1.41 | 0.97 | Emotion max among WER≤10% |
| Sadness | structured | 2 / 0.5 | 63.6% | 1.89 | 0.92 | No low-WER cell; lowest WER |
| Elation | caption | 1 / 1 | 0.0% | 1.52 | 0.60 | Emotion max among WER≤10% |
| Elation | structured | 1.5 / 0 | 54.5% | 0.82 | 0.47 | No low-WER cell; lowest WER |
| Affection | caption | 1.5 / 1 | 0.0% | 1.62 | 0.91 | Emotion max among WER≤10% |
| Affection | structured | 1 / 1.2 | 77.3% | 2.05 | 0.83 | No low-WER cell; lowest WER |
Only two seeds per emotion-format cell, one utterance, and no per-emotion naturalness constraint in this exploratory table. Do not treat these as reliable personalized presets.
Selected-setting generalization checks
German words, same reference A
| Format | CFG | T | n | Raw WER % | Post WER % | Genuine /6 | Emotion percentile | CLAP |
|---|---|---|---|---|---|---|---|---|
| caption | 1.5 | 1 | 10 | 2.7 | 2.7 | 0.90 | 0.64 | 0.109 |
| caption | 2 | 1 | 10 | 0.0 | 0.0 | 0.97 | 0.60 | 0.113 |
| caption | 2 | 1.2 | 10 | 1.8 | 3.6 | 0.76 | 0.66 | 0.107 |
| caption | 3 | 1.2 | 10 | 1.8 | 1.8 | 0.93 | 0.58 | 0.092 |
| structured | 1.5 | 1 | 10 | 118.2 | 102.7 | 1.47 | 0.63 | 0.102 |
| structured | 2 | 1 | 10 | 150.9 | 101.8 | 1.47 | 0.73 | 0.078 |
| structured | 2 | 1.2 | 10 | 140.0 | 106.4 | 1.12 | 0.78 | 0.116 |
| structured | 3 | 1.2 | 10 | 118.2 | 111.8 | 1.29 | 0.55 | 0.098 |
Same English words, distinct reference B
| Format | CFG | T | n | Raw WER % | Post WER % | Genuine /6 | Emotion percentile | CLAP |
|---|---|---|---|---|---|---|---|---|
| caption | 1.5 | 1 | 10 | 10.0 | 10.0 | 1.46 | 0.73 | 0.123 |
| caption | 2 | 1 | 10 | 9.1 | 8.2 | 1.67 | 0.75 | 0.159 |
| caption | 2 | 1.2 | 10 | 11.8 | 10.9 | 1.67 | 0.63 | 0.134 |
| caption | 3 | 1.2 | 10 | 30.9 | 18.2 | 1.81 | 0.79 | 0.221 |
| structured | 1.5 | 1 | 10 | 148.2 | 103.6 | 1.97 | 0.50 | 0.057 |
| structured | 2 | 1 | 10 | 170.9 | 101.8 | 2.12 | 0.81 | 0.046 |
| structured | 2 | 1.2 | 10 | 173.6 | 97.3 | 1.95 | 0.69 | -0.000 |
| structured | 3 | 1.2 | 10 | 179.1 | 94.5 | 2.50 | 0.54 | 0.021 |
The holdout is a limited generalization check of selected configurations; it is not a complete German or second-speaker grid and cannot establish the optimum there. The German translation has different phonetics, so language comparisons are descriptive.
Caption: equally weighted main / German / second-reference cohorts
| CFG | T | Actual n | Raw WER % | Genuineness /6 | Emotion percentile |
|---|---|---|---|---|---|
| 1.5 | 1 | 30 | 4.55 | 1.332 | 0.654 |
| 2 | 1 | 30 | 4.24 | 1.423 | 0.646 |
| 2 | 1.2 | 30 | 6.97 | 1.326 | 0.642 |
| 3 | 1.2 | 30 | 14.55 | 1.497 | 0.674 |
This descriptive pool is reported after selection; it is not a preregistered selection objective or an independent final test. All three cohorts share emotions and related words. It shows sensitivity to voice/reference and language rather than eliminating it.
First-grid preferred setting minus the existing default, paired by emotion and seed
| Format | Metric | Mean delta | Descriptive 95% cluster-bootstrap interval |
|---|---|---|---|
| caption | raw_wer | +0.0636 | [+0.0000, +0.1545] |
| caption | genuineness | -0.0949 | [-0.2243, +0.0344] |
| caption | emotion_percentile | +0.0385 | [-0.0667, +0.1436] |
| caption | clap | +0.0464 | [-0.0111, +0.1365] |
Negative WER delta is better; positive genuineness/emotion/CLAP delta is better. Intervals overlapping zero do not provide a clear directional result on this small prompt set. No multiple-comparison correction or population significance claim is made.
Exploratory structured-recipe rescue
The initial paragraph-length parenthesized directions were very unreliable. After observing that failure, an additional grid uses terse sentence-level physical cues and explicit 1.4 / 3.3-second sentence duration tags, matching the model card’s recipe style, with the same global descriptions, words, reference and seeds. This is an exploratory follow-up, not an untouched holdout. Cue verbosity, cue placement and authored timing change together, so causal attribution to any one detail is not possible.
| CFG | T | n | Raw WER % | Post WER % | Genuineness /6 | Emotion percentile |
|---|---|---|---|---|---|---|
| 1 | 0 | 10 | 94.5 | 94.5 | 1.56 | 0.48 |
| 1 | 0.5 | 10 | 25.5 | 25.5 | 1.57 | 0.53 |
| 1 | 1 | 10 | 24.5 | 27.3 | 1.58 | 0.73 |
| 1 | 1.2 | 10 | 17.3 | 2.7 | 1.46 | 0.50 |
| 1.5 | 0 | 10 | 58.2 | 56.4 | 1.33 | 0.52 |
| 1.5 | 0.5 | 10 | 30.9 | 29.1 | 1.55 | 0.59 |
| 1.5 | 1 | 10 | 29.1 | 13.6 | 1.41 | 0.68 |
| 1.5 | 1.2 | 10 | 31.8 | 32.7 | 1.66 | 0.62 |
| 2 | 0 | 10 | 83.6 | 83.6 | 1.66 | 0.53 |
| 2 | 0.5 | 10 | 12.7 | 10.0 | 1.74 | 0.57 |
| 2 | 1 | 10 | 21.8 | 9.1 | 1.27 | 0.71 |
| 2 | 1.2 | 10 | 22.7 | 13.6 | 1.39 | 0.70 |
| 3 | 0 | 10 | 80.0 | 80.0 | 1.83 | 0.51 |
| 3 | 0.5 | 10 | 15.5 | 13.6 | 1.50 | 0.71 |
| 3 | 1 | 10 | 9.1 | 5.5 | 1.25 | 0.77 |
| 3 | 1.2 | 10 | 15.5 | 7.3 | 1.44 | 0.68 |
This rescue changes the format-level conclusion: concise structured cues can produce substantially better speech than verbose paragraph cues. The initial failed recipe is not evidence that every GENERAL/SCRIPT prompt fails.
Original-backend quality diagnostic
Twenty additional free-running takes use the original serial-CFG loop and original FP32 codec weights at CFG 1.5 / T=1, with the same batch positions, prompts and seeds. Stochastic trajectories are not bit-identical after optimization, so this is a small distributional diagnostic, not a powered perceptual equivalence or non-inferiority trial.
| Backend | Format | n | Raw WER % | Post WER % | Genuineness /6 | Emotion percentile |
|---|---|---|---|---|---|---|
| Original | caption | 10 | 2.7 | 2.7 | 1.64 | 0.60 |
| Original | structured | 10 | 147.3 | 95.5 | 1.27 | 0.92 |
| Optimized | caption | 10 | 0.9 | 0.0 | 1.64 | 0.60 |
| Optimized | structured | 10 | 127.3 | 98.2 | 1.39 | 0.76 |
Metrics, uncertainty and caveats
- WER: conventional minimum word-edit distance (S+D+I)/N, Unicode-aware case folding and punctuation removal, preserving German letters. Unlike the older demo’s SequenceMatcher distance, one substituted word counts once, not twice. Raw WER is primary; post-fade WER is secondary so trimming cannot hide generation errors. ASR can itself mishear expressive speech; WER is not a human transcript audit. Study ASR uses anti-aliased 48→16 kHz resampling consistently.
- Genuineness: the existing VoiceCLAP-commercial full 0–6 head measures a “lived-in” speech impression. It is a naturalness-related proxy, not listener MOS or a complete audio-quality metric.
- Emotion: the published EIV suite is run with its documented EmoWhisper-AnS-Small-v0.1 encoder and ReLU heads. We report target rank among 40 emotions as percentile (40−rank)/39; percentile 1 means rank 1 and 0 rank 40. Raw scores and ranks are not calibrated probabilities, intensity guarantees or causal proof. Distinct heads can have different scales. The model was trained partly on synthetic voices; proxy optimization may exploit its biases.
- Direction match: VoiceCLAP-commercial centered audio/text cosine gives a second, correlated indication of whether the acting description is audible. Genuineness and direction CLAP share a backbone, so they are not independent confirmations. Blend scores are diagnostic only because no explicit bursts are requested.
- Uncertainty: 2,000 fixed-seed cluster-bootstrap resamples average the repeated seeds within emotion-format blocks. Per-format estimates have only five clusters. Intervals describe variation across this small prompt set, not broad population uncertainty. T=0 seeds may duplicate, and the common text/reference severely limit generalization. Sixteen-way grid selection has winner’s bias; holdouts and auditioning are necessary.
- Scope: S3 only, no LoRA/DPO, reference-conditioned, short utterances, caption and structured surfaces only. Tags/transcript-only, long monologues, vocal-burst onset accuracy, other stages and human preference were not tested. The grid omits intermediate temperatures such as 0.8 and CFG values between tested levels; “best tested” does not imply a continuous global optimum.
The Small model card also distinguishes experimental timing/burst control and caption intelligibility from learned quality proxies; this pilot does not inherit its checkpoint-specific numerical rankings.
Listen to the evidence
All completed takes are available below, not just winners. Players use processed mono MP3 at 96 kbps; metrics use lossless PCM before MP3. Raw playback is available to inspect unwanted lead/tail speech. Filter the table to compare the same emotion and seed across settings.
Full reproducibility
design.json · results.json · summary.json · scorer-environment.json
Re-run using emotion_study.py design, generate, score --follow; each raw take and scored result is persisted separately. The design records exact prompts, seeds, generation budget and predeclared criteria. Reference recordings are local corpus samples, not user uploads, and are not exposed by the report server.