Training-free inference-time search · measured implementation
Listen, rank, fork, continue.
Reward-guided chunk search for Humaneness Voice Small: can we steer emotional speech while emitting irreversible audio, on one RTX 3090?
English summary of the proposed idea
A pretrained semantic decoder emits one conditioning vector per 80-ms acoustic frame. A local autoregressive Talker then generates the frame’s twelve acoustic tokens. Instead of finishing N independent utterances, generate N continuations for a short chunk, decode them, and score the spoken prefix and delivery. Keep either the single best prefix or the best two, clone their semantic and codec states, and generate fresh children. With a total budget of eight and width two, each retained prefix gets four children at the next step, and the best two are selected globally from the eight continuations.
Winner-of-8: shared prefix → 8 chunk continuations → score → keep 1 → fork 8
Top-2: shared prefix → 8 chunk continuations → score → keep 2
each retained prefix → 4 children → 8 total → global top 2Width one is greedy, reward-guided chunk rollout / sequential best-of-N. Width two is a sampled, reward-guided beam-like search over audio prefixes. It is not conventional token-level beam search ranked by accumulated model log probability. The full prefix reward is recomputed, not summed across overlapping windows; summing would count the same words repeatedly. Future utterance-level reward is unknown, so neither mode is globally optimal.
Architecture caveat: the question describes a roughly 600M semantic backbone plus a 525M local Talker. The checkpoint actually available in this demo has approximately 715.5M total parameters and a 112.8M Talker/bridge, as documented in its model card. These timings therefore do not establish real-time performance for a different 525M-Talker model. The search mechanism applies conceptually to that architecture, but it needs a new benchmark.
What was implemented
- Persistent conditional/unconditional semantic KV caches, current generation row, attention mask, finished flags and all generated acoustic frames.
- Stateful incremental PCM decoding, including the codec’s per-row KV caches and offsets. Retained rows are cloned with
index_select; duplicate children do not share mutable cache storage. The full prefix is never re-generated. - A fixed N-row batch. Width two forks N/2 children per retained prefix, then takes the global top two, not one winner from each parent.
- Batched Parakeet v3 ASR with duration-derived word timestamps, and one shared VoiceCLAP audio embedding pass for style similarity, Genuineness and Burst Blend heads. Audio resampling is anti-aliased.
- Fresh stochastic continuation at each step, CFG 1.5, audio temperature 1.0, top-k 50 and top-p 0.95. Reference A is held fixed. No LoRA, DPO or Luna calls in the experiments.
Reward, normalization and partial WER
style = clip((centered_cosine + 1) / 2, 0, 1) genuine = clip(genuineness / 6, 0, 1) blend = clip(burst_blend / 10, 0, 1) quality = 0.50 * style + 0.25 * genuine + 0.25 * blend reward = quality * max(0, 1 - prefix_WER)
Thus Genuineness plus Burst Blend jointly carry the same nominal weight as style similarity. Fixed scales keep parent/time comparisons meaningful; candidate-set min–max normalization would change the meaning of a score when one bad candidate appears or disappears. Equal nominal weights do not imply equal empirical influence: CLAP has a narrower observed range, and all three quality terms share an encoder. The bounded “inverse WER” factor is 1−WER, not 1/WER, which would be singular at perfect speech.
The style query comes only from delivery instructions and parenthesized acting cues. Known literal transcript text and quoted material are removed; spoken content is never a fallback style query. This is style matching, not a speaker-identity verification score. Voice identity is conditioned by the reference, not independently checked with a speaker recognizer.
At intermediate boundaries, ASR runs on the cumulative audible prefix, retaining words with end timestamps at least 120 ms before the right edge. Word Levenshtein alignment chooses the target-prefix endpoint with minimum edit count; the not-yet-spoken suffix is free, internal deletions/substitutions are not. Empty recognized speech gets a zero intelligibility gate. At EOS or the frame cap, the complete target and complete hypothesis are used. Prefix WER is a surrogate, not a valid final-utterance WER or a confidence estimate. Timestamp filtering helps but cannot guarantee a stable boundary word.
Irreversible streaming versus a real beam
Winner-of-N commits each chosen chunk immediately: discarded histories are gone forever. For top-2, only an identical acoustic-token prefix of both retained beams can be emitted. If the surviving children later share one ancestor, that shared prefix becomes streamable. Otherwise output must wait until the final winner is known. The code never plays the current leader and subsequently swaps it for another voice trajectory.
A bounded-latency variant would retain both beams for one look-ahead chunk and then commit to one ancestor, discarding incompatible histories. That is useful but changes the search into a commitment-constrained beam. It is proposed below, not silently substituted for the exact common-prefix top-2 mode tested here. Complete-utterance Best-of-N could likewise stream only a shared prefix; independent sampled acoustic sequences generally offer none.
One-GPU versus two-GPU performance
One-GPU rows place every active TTS, codec, ASR and quality model on cuda:0, a single RTX 3090 with 24 GiB VRAM. The second card is unused for active computation. Two-GPU rows place TTS/codec on card 0 and scorers on card 1. Scoring precedes the next expansion, so a second card does not remove the sequential decision dependency. No competing GPU jobs ran during measurements; clocks were not locked.
Warm table below uses the second seed (2027) after each shape’s preceding trial. All first trials remain in the complete data, including compilation/capture costs. Each row is a measured observation, not a population average. The generation hint and hard cap were 6.6 s, quantized upward to 83 frames / 6.64 s; natural EOS was enabled and no minimum length was forced. Many outputs reached the cap, and generated trailing silence/noise remains in the RTF denominator.
Single RTX 3090
| N | Beam width | Chunk ms | Search RTF ↓ | First committed PCM, s | TTS/codec GPU s | Scoring wall s | Peak card-0 GiB | RTF <1 |
|---|---|---|---|---|---|---|---|---|
| 4 | 1 | 800 | 0.559 | 0.432 | 2.171 | 1.413 | 8.05 | Yes |
| 4 | 2 | 800 | 0.554 | 1.261 | 2.118 | 1.433 | 8.05 | Yes |
| 4 | 1 | 1600 | 0.398 | 0.591 | 1.764 | 0.808 | 8.05 | Yes |
| 4 | 2 | 1600 | 0.399 | 1.165 | 1.773 | 0.802 | 8.05 | Yes |
| 8 | 1 | 800 | 0.725 | 0.573 | 2.265 | 2.419 | 9.03 | Yes |
| 8 | 2 | 800 | 0.727 | 2.180 | 2.272 | 2.421 | 9.03 | Yes |
| 8 | 1 | 1600 | 0.509 | 0.755 | 1.940 | 1.365 | 9.03 | Yes |
| 8 | 2 | 1600 | 0.505 | 1.456 | 1.928 | 1.353 | 9.03 | Yes |
| 12 | 1 | 800 | 0.935 | 0.729 | 2.609 | 3.451 | 10.02 | Yes |
| 12 | 2 | 800 | 0.930 | 1.410 | 2.608 | 3.417 | 10.02 | Yes |
| 12 | 1 | 1600 | 0.638 | 0.939 | 2.269 | 1.887 | 10.02 | Yes |
| 12 | 2 | 1600 | 0.645 | 1.842 | 2.284 | 1.913 | 10.02 | Yes |
| 16 | 1 | 800 | 1.134 | 0.886 | 2.901 | 4.464 | 11.02 | No |
| 16 | 2 | 800 | 1.127 | 1.687 | 2.899 | 4.418 | 11.02 | No |
| 16 | 1 | 1600 | 0.781 | 1.142 | 2.598 | 2.495 | 11.01 | Yes |
| 16 | 2 | 1600 | 0.780 | 2.225 | 2.597 | 2.493 | 11.01 | Yes |
Two RTX 3090s, separate scorer card
| N | Beam width | Chunk ms | Search RTF ↓ | First committed PCM, s | TTS/codec GPU s | Scoring wall s | Peak card-0 GiB | RTF <1 |
|---|---|---|---|---|---|---|---|---|
| 4 | 1 | 800 | 0.550 | 0.423 | 2.104 | 1.426 | 6.42 | Yes |
| 4 | 2 | 800 | 0.553 | 0.848 | 2.109 | 1.439 | 6.42 | Yes |
| 4 | 1 | 1600 | 0.395 | 0.588 | 1.753 | 0.802 | 6.42 | Yes |
| 4 | 2 | 1600 | 0.396 | 1.157 | 1.753 | 0.805 | 6.42 | Yes |
| 8 | 1 | 800 | 0.715 | 0.566 | 2.246 | 2.371 | 7.41 | Yes |
| 8 | 2 | 800 | 0.715 | 2.141 | 2.239 | 2.377 | 7.41 | Yes |
| 8 | 1 | 1600 | 0.504 | 0.757 | 1.931 | 1.337 | 7.41 | Yes |
| 8 | 2 | 1600 | 0.501 | 1.450 | 1.922 | 1.329 | 7.41 | Yes |
| 12 | 1 | 800 | 0.920 | 0.720 | 2.583 | 3.376 | 8.43 | Yes |
| 12 | 2 | 800 | 0.915 | 1.382 | 2.584 | 3.346 | 8.43 | Yes |
| 12 | 1 | 1600 | 0.638 | 0.947 | 2.281 | 1.874 | 8.41 | Yes |
| 12 | 2 | 1600 | 0.640 | 1.828 | 2.270 | 1.894 | 8.41 | Yes |
| 16 | 1 | 800 | 1.116 | 0.875 | 2.865 | 4.390 | 9.43 | No |
| 16 | 2 | 800 | 1.113 | 1.669 | 2.875 | 4.355 | 9.43 | No |
| 16 | 1 | 1600 | 0.775 | 1.140 | 2.601 | 2.457 | 9.44 | Yes |
| 16 | 2 | 1600 | 0.772 | 2.203 | 2.585 | 2.454 | 9.44 | Yes |
Search RTF = search critical-path wall seconds / final delivered raw PCM seconds. It includes prompt packing/reference coding, TTS/codec, GPU transfers, cumulative-prefix ASR, timestamp processing, quality heads, CPU ranking and cache cloning. It excludes model loading, one-time style text-query embedding, Luna, MP3 encoding, optional final alignment/evaluation, HTTP and browser playback. This differs deliberately from the earlier TTS-only GPU event RTF. An N-way aggregate audio throughput divided by N is not used to claim real-time playback.
First-use costs are real
| Actual run | Search seconds | RTF | Graph capture s |
|---|---|---|---|
| measure-two_gpu-anger_en-search-n4-b1-c10-s777 | 55.07 | 8.29 | 0.74 |
| measure-two_gpu-anger_en-search-n8-b1-c10-s777 | 50.34 | 7.58 | 0.64 |
| measure-two_gpu-anger_en-search-n16-b1-c10-s777 | 7.99 | 1.20 | 0.61 |
| measure-one_gpu-anger_en-search-n16-b1-c10-s2027 | 7.53 | 1.13 | 0.00 |
| measure-one_gpu-anger_en-search-n16-b1-c10-s777 | 7.53 | 1.13 | 0.00 |
Fresh-seed speech/language stress checks
After profiling one anger text, winner search was checked on the fear and German-affection texts with a fresh seed 12345, again on one card. These are additional systems checks, not many independent speakers.
| Text | N | Chunk ms | RTF | First commit s | Raw WER % |
|---|---|---|---|---|---|
| affection_de | 12 | 800 | 0.927 | 0.732 | 0.0 |
| affection_de | 12 | 1600 | 0.647 | 0.952 | 0.0 |
| affection_de | 16 | 800 | 1.126 | 0.881 | 0.0 |
| affection_de | 16 | 1600 | 0.783 | 1.149 | 0.0 |
| affection_de | 4 | 800 | 0.551 | 0.428 | 0.0 |
| affection_de | 4 | 1600 | 0.397 | 0.596 | 60.0 |
| affection_de | 8 | 800 | 0.719 | 0.569 | 0.0 |
| affection_de | 8 | 1600 | 0.503 | 0.743 | 0.0 |
| fear_en | 12 | 800 | 1.077 | 1.699 | 0.0 |
| fear_en | 12 | 1600 | 0.644 | 0.944 | 0.0 |
| fear_en | 16 | 800 | 1.199 | 1.351 | 0.0 |
| fear_en | 16 | 1600 | 0.777 | 1.131 | 0.0 |
| fear_en | 4 | 800 | 8.305 | 51.886 | 0.0 |
| fear_en | 4 | 1600 | 0.392 | 0.585 | 40.0 |
| fear_en | 8 | 800 | 7.554 | 45.926 | 0.0 |
| fear_en | 8 | 1600 | 0.502 | 0.757 | 0.0 |
Quality: single sample versus whole-utterance BoN versus search
The primary comparison is three fixed texts (anger English, fear English, affection German) × two seeds × six methods = 36 selected outputs. Every method uses the same text, reference, CFG, temperature and 83-frame cap. Best-of-8 is ranked by the same fixed-scale reward, not the legacy UI’s candidate-min–max normalization. Equal N does not mean identical random samples or equal total compute: search evaluates many prefix sets, while BoN evaluates the finished set once. No human listening study, independent ASR or speaker-identity evaluation was performed.
| Method | Takes | Raw WER % ↓ | Style CLAP ↑ | Genuine /6 ↑ | Blend /10 ↑ | Selection reward ↑ | Held-out EIV rank percentile ↑ |
|---|---|---|---|---|---|---|---|
| Full-utterance Best-of-8 | 6 | 0.00 | 0.201 | 1.310 | 8.309 | 0.563 | 0.842 |
| Single sample | 6 | 7.73 | 0.122 | 1.195 | 7.599 | 0.483 | 0.807 |
| Top-2 / budget 8 · 1600 ms | 6 | 0.00 | 0.151 | 1.339 | 9.229 | 0.574 | 0.829 |
| Top-2 / budget 8 · 800 ms | 6 | 14.55 | 0.155 | 1.580 | 8.630 | 0.489 | 0.882 |
| Winner-of-8 · 1600 ms | 6 | 13.03 | 0.157 | 1.614 | 7.911 | 0.486 | 0.772 |
| Winner-of-8 · 800 ms | 6 | 6.06 | 0.131 | 1.338 | 9.386 | 0.539 | 0.754 |
The direction is mixed even under automatic evaluation: compared with Best-of-8, top-2/1600 ms lowered style CLAP (0.151 vs 0.201) and had slightly lower independently measured emotion rank percentile (0.829 vs 0.842), despite a higher combined search reward (0.574 vs 0.563). Top-2/800 ms had higher emotion rank percentile (0.882) but mean WER 14.55%. Winner-of-8/800 ms had mean WER 6.06% versus 7.73% for one sample, with an error concentrated in one take. These few observations do not establish a generally better decoder.
Paired effects versus Best-of-8
Seeds are paired within text; bootstrap resamples the three text clusters after averaging each seed pair. With only three clusters, intervals are descriptive and very weak evidence. Negative WER deltas are better; positive quality deltas are better. This analysis is not corrected for multiple comparisons.
| Method − Best-of-8 | Metric | Mean paired delta | Descriptive 95% text-cluster interval |
|---|---|---|---|
| Single sample | prefix_wer | +0.0773 | [+0.0000, +0.1818] |
| Single sample | clap | -0.0798 | [-0.1304, +0.0195] |
| Single sample | genuineness | -0.1153 | [-0.2891, +0.1497] |
| Single sample | reward | -0.0792 | [-0.1776, -0.0219] |
| Top-2 / budget 8 · 1600 ms | prefix_wer | +0.0000 | [+0.0000, +0.0000] |
| Top-2 / budget 8 · 1600 ms | clap | -0.0503 | [-0.0790, -0.0280] |
| Top-2 / budget 8 · 1600 ms | genuineness | +0.0289 | [-0.3437, +0.3140] |
| Top-2 / budget 8 · 1600 ms | reward | +0.0116 | [-0.0038, +0.0281] |
| Top-2 / budget 8 · 800 ms | prefix_wer | +0.1455 | [+0.0000, +0.3000] |
| Top-2 / budget 8 · 800 ms | clap | -0.0460 | [-0.0826, -0.0275] |
| Top-2 / budget 8 · 800 ms | genuineness | +0.2696 | [+0.0341, +0.5155] |
| Top-2 / budget 8 · 800 ms | reward | -0.0734 | [-0.1553, +0.0335] |
| Winner-of-8 · 1600 ms | prefix_wer | +0.1303 | [+0.0909, +0.1500] |
| Winner-of-8 · 1600 ms | clap | -0.0444 | [-0.0659, -0.0285] |
| Winner-of-8 · 1600 ms | genuineness | +0.3032 | [+0.0556, +0.5074] |
| Winner-of-8 · 1600 ms | reward | -0.0768 | [-0.0985, -0.0481] |
| Winner-of-8 · 800 ms | prefix_wer | +0.0606 | [+0.0000, +0.1818] |
| Winner-of-8 · 800 ms | clap | -0.0709 | [-0.0836, -0.0471] |
| Winner-of-8 · 800 ms | genuineness | +0.0276 | [-0.2592, +0.3579] |
| Winner-of-8 · 800 ms | reward | -0.0234 | [-0.1021, +0.0244] |
Exploratory real burst / shout recipe
After the speech-only comparison, a separate structured prompt requested a 0.3-second frustrated groan followed by angry shouted speech. No settings were re-tuned. Two seeds per method are only a diagnostic. This recipe is not comparable to the caption-only cases as an isolated causal emotion effect.
| Method | Takes | Raw WER % ↓ | Style CLAP ↑ | Genuine /6 ↑ | Blend /10 ↑ | Selection reward ↑ | Held-out EIV rank percentile ↑ |
|---|---|---|---|---|---|---|---|
| Full-utterance Best-of-8 | 2 | 6.25 | 0.222 | 2.046 | 7.554 | 0.542 | 0.868 |
| Single sample | 2 | 62.50 | 0.269 | 1.730 | 2.214 | 0.158 | 0.526 |
| Top-2 / budget 8 · 1600 ms | 2 | 0.00 | 0.234 | 2.351 | 6.758 | 0.575 | 0.947 |
| Top-2 / budget 8 · 800 ms | 2 | 0.00 | 0.211 | 1.777 | 8.488 | 0.589 | 0.566 |
| Winner-of-8 · 1600 ms | 2 | 0.00 | 0.255 | 2.142 | 7.172 | 0.582 | 0.750 |
| Winner-of-8 · 800 ms | 2 | 6.25 | 0.232 | 1.832 | 8.559 | 0.563 | 0.750 |
Independent emotion-model check
The post-search check uses the Empathic-Insight-Voice-Small heads with the documented mkrausio/EmoWhisper-AnS-Small-v0.1 encoder and ReLU activation, on raw output. This encoder/head set was not used for selection. Actual loaded head count: 39. Rank percentile is normalized by this actual count, not interpreted as a probability or an intensity guarantee. It provides a differently modeled diagnostic, still not a human rating. ASR, Genuineness, Blend and CLAP endpoint scores reuse the ranking models and are vulnerable to selection/reward hacking.
Correctness and practical UI behavior
{
"continuation": {
"same_tokens": true,
"token_agreement": 1.0,
"same_lengths": true,
"pcm_rms_difference": 0.0,
"commits_equal_final": true,
"chunk_count": 2
},
"shared_audio_tower": {
"genuineness_embedding_max_abs": 0.0,
"blend_embedding_max_abs": 0.0
},
"notes": "Numerical/causal validation on one prompt, four candidates, twenty frames; not a human perceptual-equivalence study."
}The validation checks one four-row / twenty-frame trajectory. With a constant selector, chunk-search continuation produced exactly the same acoustic tokens and PCM as the optimized normal stream’s first row. Concatenated emitted chunks exactly equaled the final raw PCM. Independently loaded packaged head towers matched the shared VoiceCLAP embedding numerically. These checks do not prove all future cache revisions or human perceptual quality; unknown codec state shapes fail closed.
Actual HTTPS Chromium integration checks passed at BS 1, 3, 4, 8 and 16 across both chunk sizes. The checks verified all 1–16 UI options, effective single-path behavior at BS 1, audio arriving before the completed result, decoded MP3 segment durations matching raw PCM duration without encoder-padding gaps, mono 96-kbps encoding, session-private audio access and mobile layout. These developer smoke tests used the live two-card deployment and are separate from the isolated one-card throughput/quality measurements; they do not establish performance for every intermediate BS.
The live demo adds two generation modes, every batch size from 1 through 16, 800/1600-ms search chunks and a configurable duration cap. Top-2 divides N children between the retained prefixes; for odd N the higher-ranked prefix receives the extra child (e.g. 4+3 for BS 7). BS 1 uses one path in either mode. The isolated throughput measurements above cover N=4/8/12/16, not every intermediate size; a newly used shape may require first-use compilation. The demo sends mono 96-kbps MP3 segments as soon as PCM is committed, schedules decoded segments using Web Audio, and keeps the final complete take downloadable. Live audio is raw; forced-alignment fades apply only to the finished take and cannot retroactively alter audio already heard. Private reference uploads and generated chunks keep the existing session isolation and one-hour cleanup. The browser stream may underrun despite model RTF <1 if compilation, HTTP/MP3 decoding or a delayed beam commitment exhausts its buffer.
Assessment and improvements worth pursuing
- Use winner-of-8 at 800 ms for a low-delay starting point; use 1600 ms for more throughput headroom. Sixteen strands need 1600 ms in these measurements. Keep search optional: it is not proven to beat BoN for overall voice quality.
- Protect semantic progress. Prefix-WER can prefer slowing down or an easy partial phrase. Add monotonic alignment/progress constraints, expected speech/non-speech masks and calibrated uncertainty for short ASR spans; use complete-target WER at completion.
- Control burst timing and identity separately. Add a localized burst-presence term or hard eligibility check for actual requested events, and a consented reference-speaker embedding constraint. Blend and style CLAP alone cannot verify either.
- Bound beam latency explicitly. Offer a one-chunk look-ahead commitment-constrained beam as a distinct mode, or keep an exact top-2 mode and show when output must wait. Do not splice unrelated beam PCM or crossfade to hide incompatible histories.
- Preserve diversity. Two near-identical children of one parent can eliminate the other promising beam immediately. Consider per-parent diversity quotas, distinct acoustic-prefix deduplication or soft/particle resampling, with a clearly changed selection rule.
- Calibrate and audit the reward. Learn fixed percentile scales on an external calibration set, validate held-out emotions/texts/voices and use blind human pairwise listening. More search can exploit a proxy rather than improve speech; a CLAP/Genuine/Blend shared backbone is not three independent judges.
- Reduce repeated scoring work. The biggest implemented saving is one batch embedding instead of three encodes per candidate. Prefix rescoring still re-encodes cumulative audio; a truly cached streaming ASR/quality encoder, bounded rolling windows plus transcript progress, and shortlist-first scoring could help. They require new accuracy/cost checks, not merely smaller windows.
- Keep model likelihood as a safety signal. Track per-frame probability/entropy and apply a calibrated plausibility constraint or KL-style regularizer, rather than selecting arbitrarily off-model prefixes solely by a noisy short-window proxy. This is a future extension, not a measured result here.
All measured output takes
Each row is a selected final output, not every discarded branch. Complete per-chunk candidate scores and retained row indices are in the JSON; discarded branch audio was not archived. Filter by phase, method, case, N, beam or chunk. Playback files are raw generated mono MP3s, not source references.
Reproducibility and sources
Design/environment · All measured results and chunk traces · Cache/embedding validation · Unselected emotion-model evaluation · Implementation and reproduction notes
NVIDIA Parakeet v3 model card documents multilingual support and timestamp capabilities. Official Transformers Parakeet implementation/docs describe duration/timestamp handling. Inference-Time Reward Hacking in Large Language Models motivates a proxy-hacking caution by analogy, not direct evidence for this TTS model.