Controlled pilot · English research report

Emotion without losing the words

A measured CFG × temperature grid for Humaneness Voice Small S3: emotional direction following, intelligibility and naturalness proxies. There is no universal optimum; the goal is to identify useful compromises under explicit constraints.

660Actual scored generations
16CFG × temperature configurations
5 × 2Emotions × prompt formats
2Repeated seeds · T=0 duplicates are not independent

Practical conclusions

Practical starting range after generalization checks: CFG 1.5–2, temperature 1.0, concise CAPTION / exact TRANSCRIPT. The first-grid emotion-first winner (CFG 2, T=1.2) did not consistently generalize to the second reference. The existing CFG 1.5 / T=1 default remains a sensible starting point; CFG 2 / T=1 is a measured alternative, not a guaranteed universal improvement.
CAPTION / exact TRANSCRIPT: pilot preference CFG 2, T 1.2. Raw mean WER 7.3%, processed WER 5.5%, genuineness 1.54/6, target emotion rank percentile 0.64. This is the best qualifying configuration under the rule below, not a human-validated global optimum.
Initial verbose GENERAL / SCRIPT recipe: no configuration met both constraints. Lowest raw WER was 88.2% at CFG 1, T 1.2. A low-error, natural and strongly emotional optimum is not established for this format on this pilot.

Keep the UI’s requested defaults (CFG 1.5, T=1) unchanged while auditioning the recommended alternatives. If intelligibility is essential, prefer the concise caption surface and a clean short reference. Do not infer that higher CFG must improve emotion: it can also produce repetition, unintelligible speech or flatter expression.

Relative naturalness, not “good naturalness” proven. Absolute genuineness scores are mostly low in this experiment (roughly 1–2 out of 6, and lower on the German text). The preference rule only limits deterioration from this baseline; it does not certify high-quality natural speech. No human listening panel was conducted.

Design saved before scoring

The main grid contains 320 takes: CFG {1, 1.5, 2, 3} × temperature {0, 0.5, 1, 1.2} × five emotions × two prompt surfaces × two seeds. Every condition uses the same English words and reference A. Configuration order was shuffled with a fixed seed; row ordering and batch position are kept fixed. All takes, including high-WER and frame-limit cases, contribute to the means. No best-of-N winner filtering is applied.

You came back. I thought I would never see you again.

Anger, fear, sadness, elation and affection use distinct actor-style descriptions, deliberately avoiding an explicit laugh, scream or sob which could confound speech errors. The structured surface requests one 0.35-second pause; caption uses the same acting direction and exact quoted transcript without the timed cue. Prompt-format comparisons therefore evaluate these complete practical recipes, not an isolated causal effect of label syntax.

The transcript is never rewritten by Luna during the study. Sampling uses top-k 50, top-p 0.95 and repetition penalty 1.0. Audio T=0 is true argmax and also makes stop/continue greedy; at positive audio temperatures stop temperature remains 1.0. Thus the T=0 contrast includes the demo’s associated stop-policy change, not only audio entropy. T=0 performs poorly on this fixed study recipe, not on every possible prompt/reference: a separate German demo smoke test at T=0 produced zero measured WER with different words and reference. CFG affects only audio logits, using an affect-neutral comparison retaining words, reference and requested pause. Generation uses natural EOS with a 5.5-second frame hint plus a 60-frame safety budget, not forced benchmark lengths.

Predeclared preference rule

Within each prompt format: eligible if mean RAW conventional WER <= 0.10 and mean processed genuineness >= same-format CFG=1.5,T=1 baseline minus 0.25. Maximize mean target-emotion rank percentile; ties within 0.01 prefer higher direction CLAP. If no eligible configuration, explicitly report no qualifying optimum and present Pareto trade-offs.

We also mark the Pareto frontier: no other tested configuration simultaneously has lower raw WER, higher genuineness and higher emotion-rank percentile, with at least one strictly better. Changing thresholds or the priority assigned to emotion can change the recommendation.

CFG × temperature maps

WER is measured on untrimmed PCM. Genuineness and emotion describe the processed audio heard in the demo. Green WER ≤10%, yellow ≤30%, red >30%. A green cell alone is not a three-objective win.

Caption / exact transcript

T = 0T = 0.5T = 1T = 1.2
CFG 1WER 78.2%Genuine 0.77 / 6
Emotion percentile 0.62
outside constraints
WER 47.3%Genuine 1.42 / 6
Emotion percentile 0.62
outside constraints
WER 2.7%Genuine 1.60 / 6
Emotion percentile 0.55
✓ eligible
WER 4.5%Genuine 1.39 / 6
Emotion percentile 0.55
✓ eligible
CFG 1.5WER 80.0%Genuine 0.76 / 6
Emotion percentile 0.47
outside constraints
WER 35.5%Genuine 1.47 / 6
Emotion percentile 0.58
outside constraints
WER 0.9%Genuine 1.64 / 6
Emotion percentile 0.60
✓ eligible · Pareto
WER 0.0%Genuine 1.35 / 6
Emotion percentile 0.58
outside constraints · Pareto
CFG 2WER 80.0%Genuine 0.99 / 6
Emotion percentile 0.60
outside constraints
WER 28.2%Genuine 1.52 / 6
Emotion percentile 0.64
outside constraints
WER 3.6%Genuine 1.62 / 6
Emotion percentile 0.59
✓ eligible
WER 7.3%Genuine 1.54 / 6
Emotion percentile 0.64
✓ eligible · Pareto
CFG 3WER 92.7%Genuine 1.37 / 6
Emotion percentile 0.48
outside constraints
WER 31.8%Genuine 1.44 / 6
Emotion percentile 0.57
outside constraints
WER 13.6%Genuine 1.56 / 6
Emotion percentile 0.74
outside constraints · Pareto
WER 10.9%Genuine 1.76 / 6
Emotion percentile 0.65
outside constraints · Pareto

Structured acting script

T = 0T = 0.5T = 1T = 1.2
CFG 1WER 100.0%Genuine 1.10 / 6
Emotion percentile 0.71
outside constraints · Pareto
WER 100.0%Genuine 1.35 / 6
Emotion percentile 0.58
outside constraints
WER 101.8%Genuine 1.50 / 6
Emotion percentile 0.62
outside constraints · Pareto
WER 88.2%Genuine 1.64 / 6
Emotion percentile 0.61
outside constraints · Pareto
CFG 1.5WER 98.2%Genuine 1.06 / 6
Emotion percentile 0.47
outside constraints
WER 128.2%Genuine 1.11 / 6
Emotion percentile 0.49
outside constraints
WER 127.3%Genuine 1.39 / 6
Emotion percentile 0.76
outside constraints · Pareto
WER 130.9%Genuine 1.25 / 6
Emotion percentile 0.64
outside constraints
CFG 2WER 141.8%Genuine 1.06 / 6
Emotion percentile 0.79
outside constraints · Pareto
WER 130.0%Genuine 1.14 / 6
Emotion percentile 0.65
outside constraints
WER 123.6%Genuine 1.31 / 6
Emotion percentile 0.67
outside constraints · Pareto
WER 168.2%Genuine 1.35 / 6
Emotion percentile 0.70
outside constraints
CFG 3WER 160.0%Genuine 0.79 / 6
Emotion percentile 0.79
outside constraints
WER 170.9%Genuine 0.98 / 6
Emotion percentile 0.77
outside constraints
WER 200.9%Genuine 0.85 / 6
Emotion percentile 0.76
outside constraints
WER 160.9%Genuine 1.06 / 6
Emotion percentile 0.83
outside constraints · Pareto
All caption configurations and descriptive 95% cluster-bootstrap WER intervals
CFGTnRaw WER % [CI]Post WER %WER ≤10%WER >30%Genuine /6Emotion percentileTarget top 5CLAPFrontier
101078.2 [72.7, 89.1]78.20%100%0.770.62140%0.051
10.51047.3 [29.1, 67.3]47.340%60%1.420.61530%0.097
11102.7 [0.0, 6.4]0.0100%0%1.600.55110%0.137
11.2104.5 [0.0, 11.8]0.990%0%1.390.54930%0.140
1.501080.0 [72.7, 87.3]80.00%100%0.760.4740%0.063
1.50.51035.5 [10.0, 60.0]35.560%40%1.470.57920%0.120
1.51100.9 [0.0, 2.7]0.0100%0%1.640.59730%0.170Pareto
1.51.2100.0 [0.0, 0.0]0.0100%0%1.350.57930%0.173Pareto
201080.0 [72.7, 87.3]80.00%100%0.990.60020%0.050
20.51028.2 [11.8, 43.6]28.260%30%1.520.64430%0.149
21103.6 [0.0, 7.3]3.680%0%1.620.5920%0.181
21.2107.3 [0.0, 16.4]5.580%10%1.540.63630%0.217Pareto
301092.7 [90.9, 96.4]92.70%100%1.370.4820%0.047
30.51031.8 [1.8, 70.0]31.860%30%1.440.56720%0.162
311013.6 [0.0, 39.1]13.680%10%1.560.74140%0.195Pareto
31.21010.9 [2.7, 23.6]10.080%10%1.760.65130%0.203Pareto
All structured configurations and descriptive 95% cluster-bootstrap WER intervals
CFGTnRaw WER % [CI]Post WER %WER ≤10%WER >30%Genuine /6Emotion percentileTarget top 5CLAPFrontier
1010100.0 [100.0, 100.0]100.00%100%1.100.70840%0.115Pareto
10.510100.0 [100.0, 100.0]100.00%100%1.350.58520%0.172
1110101.8 [95.5, 108.2]100.90%100%1.500.61550%0.146Pareto
11.21088.2 [78.2, 98.2]88.20%100%1.640.60820%0.148Pareto
1.501098.2 [72.7, 121.8]90.90%100%1.060.46920%0.069
1.50.510128.2 [100.0, 166.4]100.90%100%1.110.49220%0.096
1.5110127.3 [105.5, 150.9]98.20%100%1.390.76240%0.144Pareto
1.51.210130.9 [109.1, 150.9]103.60%100%1.250.63620%0.153
2010141.8 [92.7, 190.9]89.10%100%1.060.79540%0.119Pareto
20.510130.0 [85.5, 171.8]92.70%90%1.140.65130%0.084
2110123.6 [107.3, 146.4]103.60%100%1.310.67230%0.113Pareto
21.210168.2 [118.2, 218.2]101.80%100%1.350.70030%0.092
3010160.0 [112.7, 207.3]96.40%100%0.790.79060%0.049
30.510170.9 [102.7, 240.9]109.10%90%0.980.76740%0.090
3110200.9 [170.9, 240.9]104.50%100%0.850.76260%0.077
31.210160.9 [120.9, 198.2]96.40%90%1.060.83370%0.115Pareto

How emotion-specific are the compromises?

EmotionFormatCFG / TRaw WERGenuineEmotion percentileExploratory rule
Angercaption3 / 0.50.0%1.390.77Emotion max among WER≤10%
Angerstructured1 / 1.272.7%1.610.24No low-WER cell; lowest WER
Fearcaption1.5 / 10.0%1.710.53Emotion max among WER≤10%
Fearstructured2 / 081.8%0.000.54No low-WER cell; lowest WER
Sadnesscaption1 / 1.20.0%1.410.97Emotion max among WER≤10%
Sadnessstructured2 / 0.563.6%1.890.92No low-WER cell; lowest WER
Elationcaption1 / 10.0%1.520.60Emotion max among WER≤10%
Elationstructured1.5 / 054.5%0.820.47No low-WER cell; lowest WER
Affectioncaption1.5 / 10.0%1.620.91Emotion max among WER≤10%
Affectionstructured1 / 1.277.3%2.050.83No low-WER cell; lowest WER

Only two seeds per emotion-format cell, one utterance, and no per-emotion naturalness constraint in this exploratory table. Do not treat these as reliable personalized presets.

Selected-setting generalization checks

Reference provenance matters: reference A is a positive-valence-labelled corpus speech clip; reference B comes from a hiss-labelled take that also contains speech. Neither is a verified neutral-voice control. B may carry burst/timbre contamination, so that cohort is a reference-sensitivity stress check, not a clean-reference replication. These limitations do not invalidate within-reference CFG/temperature comparisons, but prevent a universal clean-reference optimum claim. Exact source filenames and hashes are recorded in the design JSON.

German words, same reference A

FormatCFGTnRaw WER %Post WER %Genuine /6Emotion percentileCLAP
caption1.51102.72.70.900.640.109
caption21100.00.00.970.600.113
caption21.2101.83.60.760.660.107
caption31.2101.81.80.930.580.092
structured1.5110118.2102.71.470.630.102
structured2110150.9101.81.470.730.078
structured21.210140.0106.41.120.780.116
structured31.210118.2111.81.290.550.098

Same English words, distinct reference B

FormatCFGTnRaw WER %Post WER %Genuine /6Emotion percentileCLAP
caption1.511010.010.01.460.730.123
caption21109.18.21.670.750.159
caption21.21011.810.91.670.630.134
caption31.21030.918.21.810.790.221
structured1.5110148.2103.61.970.500.057
structured2110170.9101.82.120.810.046
structured21.210173.697.31.950.69-0.000
structured31.210179.194.52.500.540.021

The holdout is a limited generalization check of selected configurations; it is not a complete German or second-speaker grid and cannot establish the optimum there. The German translation has different phonetics, so language comparisons are descriptive.

Caption: equally weighted main / German / second-reference cohorts

CFGTActual nRaw WER %Genuineness /6Emotion percentile
1.51304.551.3320.654
21304.241.4230.646
21.2306.971.3260.642
31.23014.551.4970.674

This descriptive pool is reported after selection; it is not a preregistered selection objective or an independent final test. All three cohorts share emotions and related words. It shows sensitivity to voice/reference and language rather than eliminating it.

First-grid preferred setting minus the existing default, paired by emotion and seed

FormatMetricMean deltaDescriptive 95% cluster-bootstrap interval
captionraw_wer+0.0636[+0.0000, +0.1545]
captiongenuineness-0.0949[-0.2243, +0.0344]
captionemotion_percentile+0.0385[-0.0667, +0.1436]
captionclap+0.0464[-0.0111, +0.1365]

Negative WER delta is better; positive genuineness/emotion/CLAP delta is better. Intervals overlapping zero do not provide a clear directional result on this small prompt set. No multiple-comparison correction or population significance claim is made.

Exploratory structured-recipe rescue

The initial paragraph-length parenthesized directions were very unreliable. After observing that failure, an additional grid uses terse sentence-level physical cues and explicit 1.4 / 3.3-second sentence duration tags, matching the model card’s recipe style, with the same global descriptions, words, reference and seeds. This is an exploratory follow-up, not an untouched holdout. Cue verbosity, cue placement and authored timing change together, so causal attribution to any one detail is not possible.

CFGTnRaw WER %Post WER %Genuineness /6Emotion percentile
101094.594.51.560.48
10.51025.525.51.570.53
111024.527.31.580.73
11.21017.32.71.460.50
1.501058.256.41.330.52
1.50.51030.929.11.550.59
1.511029.113.61.410.68
1.51.21031.832.71.660.62
201083.683.61.660.53
20.51012.710.01.740.57
211021.89.11.270.71
21.21022.713.61.390.70
301080.080.01.830.51
30.51015.513.61.500.71
31109.15.51.250.77
31.21015.57.31.440.68
Short-cue structured candidate: CFG 3, T=1. Raw WER 9.1%, post-fade WER 5.5%, genuineness 1.25/6 and emotion percentile 0.77. It passes the same relative-baseline rule on this follow-up, but its absolute naturalness proxy is low and it has not been validated on German or reference B. This is an audition candidate, not the recommended universal default.

This rescue changes the format-level conclusion: concise structured cues can produce substantially better speech than verbose paragraph cues. The initial failed recipe is not evidence that every GENERAL/SCRIPT prompt fails.

Original-backend quality diagnostic

Twenty additional free-running takes use the original serial-CFG loop and original FP32 codec weights at CFG 1.5 / T=1, with the same batch positions, prompts and seeds. Stochastic trajectories are not bit-identical after optimization, so this is a small distributional diagnostic, not a powered perceptual equivalence or non-inferiority trial.

BackendFormatnRaw WER %Post WER %Genuineness /6Emotion percentile
Originalcaption102.72.71.640.60
Originalstructured10147.395.51.270.92
Optimizedcaption100.90.01.640.60
Optimizedstructured10127.398.21.390.76

Metrics, uncertainty and caveats

The Small model card also distinguishes experimental timing/burst control and caption intelligibility from learned quality proxies; this pilot does not inherit its checkpoint-specific numerical rankings.

Listen to the evidence

All completed takes are available below, not just winners. Players use processed mono MP3 at 96 kbps; metrics use lossless PCM before MP3. Raw playback is available to inspect unwanted lead/tail speech. Filter the table to compare the same emotion and seed across settings.

Full reproducibility

design.json · results.json · summary.json · scorer-environment.json

Re-run using emotion_study.py design, generate, score --follow; each raw take and scored result is persisted separately. The design records exact prompts, seeds, generation budget and predeclared criteria. Reference recordings are local corpus samples, not user uploads, and are not exposed by the report server.