Training-free inference-time search · measured implementation

Listen, rank, fork, continue.

Reward-guided chunk search for Humaneness Voice Small: can we steer emotional speech while emitting irreversible audio, on one RTX 3090?

129Actual output takes, including first-use checks
4 / 8 / 12 / 16Fixed parallel candidate budgets
800 / 1600 ms10 / 20 acoustic frames per search decision
1 / 2Retained prefixes; not token-level probability beams
Systems answer: yes, within the measured envelope. On one RTX 3090, warm eight-strand search measured about RTF 0.73 at 800 ms and 0.51 at 1600 ms, including ASR and quality selection. Sixteen strands measured about 1.13 at 800 ms (slower than real time), but 0.78 at 1600 ms. This is output-duration RTF, not aggregate candidate RTF or an end-to-end browser guarantee. First-use compilation can take tens of seconds.
Quality answer: a trade-off, not a universal win. In the three speech-only texts × two seeds pilot, top-2/1600 ms and Best-of-8 both had zero mean measured WER. Top-2/1600 ms had a slightly higher combined selection reward but lower style CLAP. Other search variants introduced word errors, including substantial errors in a German take. Learned reward gains are not evidence of improved human naturalness.

English summary of the proposed idea

A pretrained semantic decoder emits one conditioning vector per 80-ms acoustic frame. A local autoregressive Talker then generates the frame’s twelve acoustic tokens. Instead of finishing N independent utterances, generate N continuations for a short chunk, decode them, and score the spoken prefix and delivery. Keep either the single best prefix or the best two, clone their semantic and codec states, and generate fresh children. With a total budget of eight and width two, each retained prefix gets four children at the next step, and the best two are selected globally from the eight continuations.

Winner-of-8: shared prefix → 8 chunk continuations → score → keep 1 → fork 8
Top-2:       shared prefix → 8 chunk continuations → score → keep 2
             each retained prefix → 4 children → 8 total → global top 2

Width one is greedy, reward-guided chunk rollout / sequential best-of-N. Width two is a sampled, reward-guided beam-like search over audio prefixes. It is not conventional token-level beam search ranked by accumulated model log probability. The full prefix reward is recomputed, not summed across overlapping windows; summing would count the same words repeatedly. Future utterance-level reward is unknown, so neither mode is globally optimal.

Architecture caveat: the question describes a roughly 600M semantic backbone plus a 525M local Talker. The checkpoint actually available in this demo has approximately 715.5M total parameters and a 112.8M Talker/bridge, as documented in its model card. These timings therefore do not establish real-time performance for a different 525M-Talker model. The search mechanism applies conceptually to that architecture, but it needs a new benchmark.

What was implemented

Reward, normalization and partial WER

style = clip((centered_cosine + 1) / 2, 0, 1)
genuine = clip(genuineness / 6, 0, 1)
blend = clip(burst_blend / 10, 0, 1)
quality = 0.50 * style + 0.25 * genuine + 0.25 * blend
reward = quality * max(0, 1 - prefix_WER)

Thus Genuineness plus Burst Blend jointly carry the same nominal weight as style similarity. Fixed scales keep parent/time comparisons meaningful; candidate-set min–max normalization would change the meaning of a score when one bad candidate appears or disappears. Equal nominal weights do not imply equal empirical influence: CLAP has a narrower observed range, and all three quality terms share an encoder. The bounded “inverse WER” factor is 1−WER, not 1/WER, which would be singular at perfect speech.

The style query comes only from delivery instructions and parenthesized acting cues. Known literal transcript text and quoted material are removed; spoken content is never a fallback style query. This is style matching, not a speaker-identity verification score. Voice identity is conditioned by the reference, not independently checked with a speaker recognizer.

At intermediate boundaries, ASR runs on the cumulative audible prefix, retaining words with end timestamps at least 120 ms before the right edge. Word Levenshtein alignment chooses the target-prefix endpoint with minimum edit count; the not-yet-spoken suffix is free, internal deletions/substitutions are not. Empty recognized speech gets a zero intelligibility gate. At EOS or the frame cap, the complete target and complete hypothesis are used. Prefix WER is a surrogate, not a valid final-utterance WER or a confidence estimate. Timestamp filtering helps but cannot guarantee a stable boundary word.

Intentional pauses and bursts need a better gate. A chunk containing only an intended groan, laugh or silence can have empty ASR and hence zero reward in this implementation. It then falls back to stable row-order tie breaking. The exploratory groan case tests this limitation; no burst-presence detector is included in the search reward. A high Blend score does not demonstrate that a requested burst exists or occurred at the right time.

Irreversible streaming versus a real beam

Winner-of-N commits each chosen chunk immediately: discarded histories are gone forever. For top-2, only an identical acoustic-token prefix of both retained beams can be emitted. If the surviving children later share one ancestor, that shared prefix becomes streamable. Otherwise output must wait until the final winner is known. The code never plays the current leader and subsequently swaps it for another voice trajectory.

A bounded-latency variant would retain both beams for one look-ahead chunk and then commit to one ancestor, discarding incompatible histories. That is useful but changes the search into a commitment-constrained beam. It is proposed below, not silently substituted for the exact common-prefix top-2 mode tested here. Complete-utterance Best-of-N could likewise stream only a shared prefix; independent sampled acoustic sequences generally offer none.

One-GPU versus two-GPU performance

One-GPU rows place every active TTS, codec, ASR and quality model on cuda:0, a single RTX 3090 with 24 GiB VRAM. The second card is unused for active computation. Two-GPU rows place TTS/codec on card 0 and scorers on card 1. Scoring precedes the next expansion, so a second card does not remove the sequential decision dependency. No competing GPU jobs ran during measurements; clocks were not locked.

Warm table below uses the second seed (2027) after each shape’s preceding trial. All first trials remain in the complete data, including compilation/capture costs. Each row is a measured observation, not a population average. The generation hint and hard cap were 6.6 s, quantized upward to 83 frames / 6.64 s; natural EOS was enabled and no minimum length was forced. Many outputs reached the cap, and generated trailing silence/noise remains in the RTF denominator.

Single RTX 3090

NBeam widthChunk msSearch RTF ↓First committed PCM, sTTS/codec GPU sScoring wall sPeak card-0 GiBRTF <1
418000.5590.4322.1711.4138.05Yes
428000.5541.2612.1181.4338.05Yes
4116000.3980.5911.7640.8088.05Yes
4216000.3991.1651.7730.8028.05Yes
818000.7250.5732.2652.4199.03Yes
828000.7272.1802.2722.4219.03Yes
8116000.5090.7551.9401.3659.03Yes
8216000.5051.4561.9281.3539.03Yes
1218000.9350.7292.6093.45110.02Yes
1228000.9301.4102.6083.41710.02Yes
12116000.6380.9392.2691.88710.02Yes
12216000.6451.8422.2841.91310.02Yes
1618001.1340.8862.9014.46411.02No
1628001.1271.6872.8994.41811.02No
16116000.7811.1422.5982.49511.01Yes
16216000.7802.2252.5972.49311.01Yes

Two RTX 3090s, separate scorer card

NBeam widthChunk msSearch RTF ↓First committed PCM, sTTS/codec GPU sScoring wall sPeak card-0 GiBRTF <1
418000.5500.4232.1041.4266.42Yes
428000.5530.8482.1091.4396.42Yes
4116000.3950.5881.7530.8026.42Yes
4216000.3961.1571.7530.8056.42Yes
818000.7150.5662.2462.3717.41Yes
828000.7152.1412.2392.3777.41Yes
8116000.5040.7571.9311.3377.41Yes
8216000.5011.4501.9221.3297.41Yes
1218000.9200.7202.5833.3768.43Yes
1228000.9151.3822.5843.3468.43Yes
12116000.6380.9472.2811.8748.41Yes
12216000.6401.8282.2701.8948.41Yes
1618001.1160.8752.8654.3909.43No
1628001.1131.6692.8754.3559.43No
16116000.7751.1402.6012.4579.44Yes
16216000.7722.2032.5852.4549.44Yes

Search RTF = search critical-path wall seconds / final delivered raw PCM seconds. It includes prompt packing/reference coding, TTS/codec, GPU transfers, cumulative-prefix ASR, timestamp processing, quality heads, CPU ranking and cache cloning. It excludes model loading, one-time style text-query embedding, Luna, MP3 encoding, optional final alignment/evaluation, HTTP and browser playback. This differs deliberately from the earlier TTS-only GPU event RTF. An N-way aggregate audio throughput divided by N is not used to claim real-time playback.

First-use costs are real

Actual runSearch secondsRTFGraph capture s
measure-two_gpu-anger_en-search-n4-b1-c10-s77755.078.290.74
measure-two_gpu-anger_en-search-n8-b1-c10-s77750.347.580.64
measure-two_gpu-anger_en-search-n16-b1-c10-s7777.991.200.61
measure-one_gpu-anger_en-search-n16-b1-c10-s20277.531.130.00
measure-one_gpu-anger_en-search-n16-b1-c10-s7777.531.130.00

Fresh-seed speech/language stress checks

After profiling one anger text, winner search was checked on the fear and German-affection texts with a fresh seed 12345, again on one card. These are additional systems checks, not many independent speakers.

TextNChunk msRTFFirst commit sRaw WER %
affection_de128000.9270.7320.0
affection_de1216000.6470.9520.0
affection_de168001.1260.8810.0
affection_de1616000.7831.1490.0
affection_de48000.5510.4280.0
affection_de416000.3970.59660.0
affection_de88000.7190.5690.0
affection_de816000.5030.7430.0
fear_en128001.0771.6990.0
fear_en1216000.6440.9440.0
fear_en168001.1991.3510.0
fear_en1616000.7771.1310.0
fear_en48008.30551.8860.0
fear_en416000.3920.58540.0
fear_en88007.55445.9260.0
fear_en816000.5020.7570.0

Quality: single sample versus whole-utterance BoN versus search

The primary comparison is three fixed texts (anger English, fear English, affection German) × two seeds × six methods = 36 selected outputs. Every method uses the same text, reference, CFG, temperature and 83-frame cap. Best-of-8 is ranked by the same fixed-scale reward, not the legacy UI’s candidate-min–max normalization. Equal N does not mean identical random samples or equal total compute: search evaluates many prefix sets, while BoN evaluates the finished set once. No human listening study, independent ASR or speaker-identity evaluation was performed.

MethodTakesRaw WER % ↓Style CLAP ↑Genuine /6 ↑Blend /10 ↑Selection reward ↑Held-out EIV rank percentile ↑
Full-utterance Best-of-860.000.2011.3108.3090.5630.842
Single sample67.730.1221.1957.5990.4830.807
Top-2 / budget 8 · 1600 ms60.000.1511.3399.2290.5740.829
Top-2 / budget 8 · 800 ms614.550.1551.5808.6300.4890.882
Winner-of-8 · 1600 ms613.030.1571.6147.9110.4860.772
Winner-of-8 · 800 ms66.060.1311.3389.3860.5390.754

The direction is mixed even under automatic evaluation: compared with Best-of-8, top-2/1600 ms lowered style CLAP (0.151 vs 0.201) and had slightly lower independently measured emotion rank percentile (0.829 vs 0.842), despite a higher combined search reward (0.574 vs 0.563). Top-2/800 ms had higher emotion rank percentile (0.882) but mean WER 14.55%. Winner-of-8/800 ms had mean WER 6.06% versus 7.73% for one sample, with an error concentrated in one take. These few observations do not establish a generally better decoder.

Paired effects versus Best-of-8

Seeds are paired within text; bootstrap resamples the three text clusters after averaging each seed pair. With only three clusters, intervals are descriptive and very weak evidence. Negative WER deltas are better; positive quality deltas are better. This analysis is not corrected for multiple comparisons.

Method − Best-of-8MetricMean paired deltaDescriptive 95% text-cluster interval
Single sampleprefix_wer+0.0773[+0.0000, +0.1818]
Single sampleclap-0.0798[-0.1304, +0.0195]
Single samplegenuineness-0.1153[-0.2891, +0.1497]
Single samplereward-0.0792[-0.1776, -0.0219]
Top-2 / budget 8 · 1600 msprefix_wer+0.0000[+0.0000, +0.0000]
Top-2 / budget 8 · 1600 msclap-0.0503[-0.0790, -0.0280]
Top-2 / budget 8 · 1600 msgenuineness+0.0289[-0.3437, +0.3140]
Top-2 / budget 8 · 1600 msreward+0.0116[-0.0038, +0.0281]
Top-2 / budget 8 · 800 msprefix_wer+0.1455[+0.0000, +0.3000]
Top-2 / budget 8 · 800 msclap-0.0460[-0.0826, -0.0275]
Top-2 / budget 8 · 800 msgenuineness+0.2696[+0.0341, +0.5155]
Top-2 / budget 8 · 800 msreward-0.0734[-0.1553, +0.0335]
Winner-of-8 · 1600 msprefix_wer+0.1303[+0.0909, +0.1500]
Winner-of-8 · 1600 msclap-0.0444[-0.0659, -0.0285]
Winner-of-8 · 1600 msgenuineness+0.3032[+0.0556, +0.5074]
Winner-of-8 · 1600 msreward-0.0768[-0.0985, -0.0481]
Winner-of-8 · 800 msprefix_wer+0.0606[+0.0000, +0.1818]
Winner-of-8 · 800 msclap-0.0709[-0.0836, -0.0471]
Winner-of-8 · 800 msgenuineness+0.0276[-0.2592, +0.3579]
Winner-of-8 · 800 msreward-0.0234[-0.1021, +0.0244]

Exploratory real burst / shout recipe

After the speech-only comparison, a separate structured prompt requested a 0.3-second frustrated groan followed by angry shouted speech. No settings were re-tuned. Two seeds per method are only a diagnostic. This recipe is not comparable to the caption-only cases as an isolated causal emotion effect.

MethodTakesRaw WER % ↓Style CLAP ↑Genuine /6 ↑Blend /10 ↑Selection reward ↑Held-out EIV rank percentile ↑
Full-utterance Best-of-826.250.2222.0467.5540.5420.868
Single sample262.500.2691.7302.2140.1580.526
Top-2 / budget 8 · 1600 ms20.000.2342.3516.7580.5750.947
Top-2 / budget 8 · 800 ms20.000.2111.7778.4880.5890.566
Winner-of-8 · 1600 ms20.000.2552.1427.1720.5820.750
Winner-of-8 · 800 ms26.250.2321.8328.5590.5630.750

Independent emotion-model check

The post-search check uses the Empathic-Insight-Voice-Small heads with the documented mkrausio/EmoWhisper-AnS-Small-v0.1 encoder and ReLU activation, on raw output. This encoder/head set was not used for selection. Actual loaded head count: 39. Rank percentile is normalized by this actual count, not interpreted as a probability or an intensity guarantee. It provides a differently modeled diagnostic, still not a human rating. ASR, Genuineness, Blend and CLAP endpoint scores reuse the ranking models and are vulnerable to selection/reward hacking.

Correctness and practical UI behavior

{
  "continuation": {
    "same_tokens": true,
    "token_agreement": 1.0,
    "same_lengths": true,
    "pcm_rms_difference": 0.0,
    "commits_equal_final": true,
    "chunk_count": 2
  },
  "shared_audio_tower": {
    "genuineness_embedding_max_abs": 0.0,
    "blend_embedding_max_abs": 0.0
  },
  "notes": "Numerical/causal validation on one prompt, four candidates, twenty frames; not a human perceptual-equivalence study."
}

The validation checks one four-row / twenty-frame trajectory. With a constant selector, chunk-search continuation produced exactly the same acoustic tokens and PCM as the optimized normal stream’s first row. Concatenated emitted chunks exactly equaled the final raw PCM. Independently loaded packaged head towers matched the shared VoiceCLAP embedding numerically. These checks do not prove all future cache revisions or human perceptual quality; unknown codec state shapes fail closed.

Actual HTTPS Chromium integration checks passed at BS 1, 3, 4, 8 and 16 across both chunk sizes. The checks verified all 1–16 UI options, effective single-path behavior at BS 1, audio arriving before the completed result, decoded MP3 segment durations matching raw PCM duration without encoder-padding gaps, mono 96-kbps encoding, session-private audio access and mobile layout. These developer smoke tests used the live two-card deployment and are separate from the isolated one-card throughput/quality measurements; they do not establish performance for every intermediate BS.

The live demo adds two generation modes, every batch size from 1 through 16, 800/1600-ms search chunks and a configurable duration cap. Top-2 divides N children between the retained prefixes; for odd N the higher-ranked prefix receives the extra child (e.g. 4+3 for BS 7). BS 1 uses one path in either mode. The isolated throughput measurements above cover N=4/8/12/16, not every intermediate size; a newly used shape may require first-use compilation. The demo sends mono 96-kbps MP3 segments as soon as PCM is committed, schedules decoded segments using Web Audio, and keeps the final complete take downloadable. Live audio is raw; forced-alignment fades apply only to the finished take and cannot retroactively alter audio already heard. Private reference uploads and generated chunks keep the existing session isolation and one-hour cleanup. The browser stream may underrun despite model RTF <1 if compilation, HTTP/MP3 decoding or a delayed beam commitment exhausts its buffer.

Assessment and improvements worth pursuing

  1. Use winner-of-8 at 800 ms for a low-delay starting point; use 1600 ms for more throughput headroom. Sixteen strands need 1600 ms in these measurements. Keep search optional: it is not proven to beat BoN for overall voice quality.
  2. Protect semantic progress. Prefix-WER can prefer slowing down or an easy partial phrase. Add monotonic alignment/progress constraints, expected speech/non-speech masks and calibrated uncertainty for short ASR spans; use complete-target WER at completion.
  3. Control burst timing and identity separately. Add a localized burst-presence term or hard eligibility check for actual requested events, and a consented reference-speaker embedding constraint. Blend and style CLAP alone cannot verify either.
  4. Bound beam latency explicitly. Offer a one-chunk look-ahead commitment-constrained beam as a distinct mode, or keep an exact top-2 mode and show when output must wait. Do not splice unrelated beam PCM or crossfade to hide incompatible histories.
  5. Preserve diversity. Two near-identical children of one parent can eliminate the other promising beam immediately. Consider per-parent diversity quotas, distinct acoustic-prefix deduplication or soft/particle resampling, with a clearly changed selection rule.
  6. Calibrate and audit the reward. Learn fixed percentile scales on an external calibration set, validate held-out emotions/texts/voices and use blind human pairwise listening. More search can exploit a proxy rather than improve speech; a CLAP/Genuine/Blend shared backbone is not three independent judges.
  7. Reduce repeated scoring work. The biggest implemented saving is one batch embedding instead of three encodes per candidate. Prefix rescoring still re-encodes cumulative audio; a truly cached streaming ASR/quality encoder, bounded rolling windows plus transcript progress, and shortlist-first scoring could help. They require new accuracy/cost checks, not merely smaller windows.
  8. Keep model likelihood as a safety signal. Track per-frame probability/entropy and apply a calibrated plausibility constraint or KL-style regularizer, rather than selecting arbitrarily off-model prefixes solely by a noisy short-window proxy. This is a future extension, not a measured result here.

All measured output takes

Each row is a selected final output, not every discarded branch. Complete per-chunk candidate scores and retained row indices are in the JSON; discarded branch audio was not archived. Filter by phase, method, case, N, beam or chunk. Playback files are raw generated mono MP3s, not source references.

Reproducibility and sources

Design/environment · All measured results and chunk traces · Cache/embedding validation · Unselected emotion-model evaluation · Implementation and reproduction notes

NVIDIA Parakeet v3 model card documents multilingual support and timestamp capabilities. Official Transformers Parakeet implementation/docs describe duration/timestamp handling. Inference-Time Reward Hacking in Large Language Models motivates a proxy-hacking caution by analogy, not direct evidence for this TTS model.