# Tested chunk-search integration source

This folder contains the actual new search implementation and its integration
into the existing local Humaneness Voice Small demo. The Hugging Face Space is
**static**: it hosts reports, JSON, code and synthetic listening examples; it
does not execute these Python files or offer free GPU inference.

## Core files

- `beam_search.py`: batched continuations, persistent semantic KV state,
  cloning/reordering of conditional/unconditional rows and codec caches,
  winner-of-N and global top-2 with common-prefix audio commitment.
- `search_scoring.py`: batched Parakeet v3, duration-derived word timestamps,
  prefix Levenshtein gate, fixed-scale reward and one shared commercial
  VoiceCLAP audio encode plus Genuineness/Blend heads.
- `performance_core.py`, `metrics.py`: fast local CUDA graphs, fused CFG,
  stateful streaming codec and conventional Unicode-aware word errors.
- `small_app.py`, `index.html`: session-private SSE/MP3 chunks, Web Audio
  scheduling, search controls and performance inspection in the live demo.
- `engine.py` and adjacent demo modules: Small S3 loader and integration
  context. This wrapper still imports the pre-existing parent demo modules.

## Reproduce in the original demo environment

Stop the live app before isolated GPU measurements, but keep its HTTPS tunnel.
Use the demo's existing CUDA environment and source/reward assets:

```bash
python search_benchmark.py --phase measure
python search_benchmark.py --phase quality
python search_benchmark.py --phase audit
python search_validation.py
python search_emotion_check.py
python search_report.py
```

The benchmark writes resumable per-output JSON, synthetic mono 96-kbps MP3s,
and local float PCM for scoring. Source reference WAVs and local PCM WAVs are
not included in the public Space. Reproduction needs a consented/distinct
speech reference; the exact original reference provenance and hashes are in
`../search/design.json`, but its recording is deliberately not uploaded.

This is **integration source, not a self-contained installation package**.
Dependencies include the pinned model repository's custom MOSS/Qwen3 code,
MOSS Audio Tokenizer v2, CUDA PyTorch/torchaudio, Transformers with Parakeet TDT,
FastAPI, FFmpeg, and the existing parent demo's `tts_engine`, `bestofn`,
`asr_engine`, `retrieval`, `timed_script`, `score_engine`, `reports`, and local
`settings`/`config` modules. The commercial VoiceCLAP, Genuineness, Blend and
Empathic Insight head assets must be provisioned separately with their
upstream licenses. Machine-specific asset paths in these integration files
are provenance/working-environment defaults, not portable download locations.
No API credential, private reference audio or runtime directory is published.

`SMALL_SCORE_DEVICE=cuda:0` puts the live demo's scorers on the TTS card; the
measured benchmark also explicitly tests moving every active search model to
one RTX 3090. The default live deployment keeps scorers on `cuda:1`. Search is
not supported by the fallback legacy generator, and the tested optimized
path expects 12 codebooks and this pinned codec's per-row streaming state.

## Important semantics

The UI/API accepts every integer batch size 1–16 and 800/1600-ms chunks.
Top-2 divides the batch between its retained prefixes, giving the extra child
to the higher-ranked prefix for odd N. BS 1 uses one path in either mode.
The isolated throughput study measured 4/8/12/16, not all intermediate sizes;
new shapes may require first-use compilation.

The duration field is a cap, rounded to 80-ms frames, not forced minimum
speech duration. Winner search streams irreversible chunks immediately.
Exact top-2 can emit only common acoustic-token prefixes, so its first audio
may be delayed until the end. Live MP3 audio is raw; final optional alignment
fades cannot retroactively change it. Quality scores are proxies and search
does not guarantee an improvement over full-utterance Best-of-N.

Original study/integration material: CC BY 4.0, attribute LAION / Humaneness
Voice Small. Upstream dependencies retain their own licenses. Refer to the
[complete measured report](../search/index.html) before extrapolating the
timings to a 525M-Talker model or other hardware.
