Case Study · Research Work · Speech Synthesis Engine

Zhiyin

A model-agnostic local Chinese TTS engine — emotion-controllable, commercially usable, running on a consumer 8GB GPU

Role:Solo design & implementation
Timeline:2026
Status:Studio v0.2.0 released · Technical report v0.3
Test platform:RTX 2080 Super Max-Q (8GB)
4 open-source engines evaluated
3.3% held-out pinyin CER after fine-tuning
0.892 speaker cosine (0.423 cross-speaker)
4 speech engines compared

The name comes from a classical Chinese story: Yu Boya played the zither, and Zhong Ziqi was the one who truly heard what he meant. Zhiyin = the one who truly understands your voice.

Research questions

Audiobooks and interactive narration need Chinese TTS that is emotionally rich, voice-clonable, and commercially usable. Every existing option falls short somewhere: commercial cloud APIs raise cost and data-sovereignty concerns, while most open-source models have weak emotion control or insufficient Chinese quality. And the "best engine" keeps moving as the open-source ecosystem evolves.

So instead of picking one model, I first asked: can swapping models be made cheap? That question defined the entire architecture.

  1. RQ1 (architecture): Can a single abstraction layer make mutually incompatible TTS engines hot-swappable, so that evaluating and replacing engines becomes a low-cost operation?
  2. RQ2 (engine selection): Under the joint constraint of cloning × emotion × Chinese quality × commercial licensing × 8GB local hardware, where are the real capability limits of open-source engines? Cascaded voice conversion or end-to-end?
  3. RQ3 (controllability): Can voice-acting-level control (pronunciation, pauses, breaths, sound effects, emotion) be provided independently of any single engine — including engines that lack the corresponding input channel?
  4. RQ4 (reliability): A single output from a sampling-based TTS is unreliable (dropped or truncated words). What mechanism can guarantee production-grade correctness?
  5. RQ5 (data efficiency): How much audio must an amateur record before cloning quality is "good enough"? Are timbre, pronunciation, and emotion equally sensitive to data volume?

System design

Four design decisions, addressing RQ1 through RQ4 in order.

  1. 1 · Backend abstraction (the swappable seam)

    A Backend interface (is_up(), synthesize(voice, text, lang, params), capabilities()). Each voice declares its engine; the core knows nothing about any specific model. Adding an engine = implement the interface + register one line. This paid off twice: a four-engine evaluation and a later switch of the primary engine, both at zero migration cost.

  2. 2 · Multi-environment isolation

    The engines have mutually exclusive dependencies (torch 2.0/cu118 + Python 3.9 vs torch 2.8/cu126 vs onnxruntime ≥ 1.20). Rather than forcing coexistence, each heavy engine runs in its own venv, bridged by subprocess (JSON over stdin, results over stdout) — like calling an external tool, with zero cross-contamination. Side benefit: not a single line of upstream code enters my repo, keeping the licensing boundary clean.

  3. 3 · Engine-level markup system (capability lifting)

    Inline markup for emotion, timed pauses, real human breath sounds, sound effects, and pronunciation overrides is implemented at the engine layer (parse → split → synthesize per segment → insert assets → resample and concatenate), so every backend gets it for free. Principle: engine-layer features are inherently shared across models; a model's intrinsic input channels cannot be transplanted — but they can often be imitated.

  4. 4 · ASR-verified retry pipeline (production loop)

    A resident Whisper worker (loaded once, cutting verification from 30–60s to seconds) plus engine-level QC: synthesize(qc=True) runs generate → compute pinyin CER → reroll with a new seed if over threshold → keep the best. This turns reliability of sampling-based generation from luck into a guarantee.

Evaluation methodology

All engines were tested through the same Backend interface with the same reference audio, following the mainstream objective protocol for zero-shot TTS (Seed-TTS eval):

  1. Content accuracy — pinyin character error rate (CER)

    Synthesized audio is transcribed by faster-whisper (large-v3, CPU int8, beam=5); both transcript and target are converted to toneless pinyin syllable sequences (pypinyin), then compared by syllable-level Levenshtein distance normalized by target length. Working at the pinyin layer avoids traditional/simplified mismatches and measures pronunciation directly.

  2. Timbre fidelity — speaker embedding cosine

    ERes2NetV2 (16 kHz fbank) cosine similarity between the synthesized audio and the centroid of that speaker's real recordings. Intra-speaker similarity and a cross-speaker baseline provide upper and lower anchors, so a score like 0.8 can actually be interpreted.

  3. Subjective check — native-speaker blind listening against the transcript

    Every key decision required objective and subjective agreement. The two are complementary and neither suffices: ASR catches dropped words but its language model silently "repairs" real mispronunciations, and it cannot hear noise floor or rhythm; human listening is subject to single-impression and expectation bias.

Engine evaluation results

EngineApproachChineseEmotion controlOutcome
GPT-SoVITS v2Profew-shot cloninggoodno native emotion knobbackup → promoted to daily driver after fine-tuning
CosyVoice2instruction + cloningdialect-leaning / unstableinstruction texteliminated (Chinese quality)
seed-vc (cascade)emotion conversion on any TTSdepends on sourceconversion-basedeliminated (emotion dilution)
IndexTTS-2zero-shot cloning + decoupled emotionnatively goodemo_vector / emo_text / emo_audio + intensity αadopted as specialist
Evaluation outcomes under the joint constraint of cloning × emotion × Chinese × licensing × 8GB.

No single open-source model wins on all three axes (emotion, fine control, naturalness). The system's value is exactly that the architecture makes "pick the best engine per character" a zero-cost decision — and lets me reverse course as the ecosystem moves. Re-testing after real human recordings arrived overturned my initial judgement: IndexTTS-2's weaker Chinese articulation proved intrinsic, while a fine-tuned GPT-SoVITS was near-perfect and 2–8 s per line. The primary engine swapped — at zero migration cost, because the architecture was ready.

Two negative results (transferable knowledge)

Negative results rarely make it into a portfolio, but they are often worth more than the successes — they save someone else from walking down a dead end.

  1. ① Cascaded voice conversion dilutes emotion layer by layer

    The route "generate neutral speech with any TTS, then convert emotion with seed-vc" measurably loses emotion at each stage: to preserve timbre and content, VC suppresses emotional displacement, leaving intensity too weak to differentiate. Conclusion: emotion must be generated natively at synthesis time, not converted afterwards. This eliminated the entire cascade route.

  2. ② Synthetic speech cannot be zero-shot re-cloned ("a copy of a copy")

    MOSS-VoiceGenerator designs a voice from a text description and synthesizes it at 0% CER. But feeding that synthetic audio as a reference into IndexTTS-2 for zero-shot cloning degrades CER to 11–57% (35–39% with strong emotion), and lowering temperature or emotion intensity or changing seeds does not rescue it. → A designed voice must be synthesized by its own generating model, never used as a cloning reference.

  3. ③ But fine-tuning solves it — the way out of the negative result

    24 training lines generated from one designed voice (fixed seed to lock timbre) → ASR-QC cleaning (21 kept) → local GPT-SoVITS fine-tune (8GB, ~15 min). Held-out results: 3.3% pinyin CER and 0.892 speaker cosine (0.423 cross-speaker) — higher than the mutual consistency of the source clips (0.716). Fine-tuning not only preserves timbre, it corrects the cross-utterance drift of generative voices: zero-shot cloning imitates one clip (inheriting its flaws), while fine-tuning learns the distribution (averaging them out).

Performance: the hard 8GB boundary

I initially measured "~40 minutes per line" for IndexTTS-2. Item-by-item measurement showed latency has three separable sources: external GPU contention, model loading, and inference itself.

ItemMeasuredMeaning
Model load (one-time)~183 spaid once with a resident worker
Pure inference (per line)~84–89 sthe real marginal cost
100 lines / chapter (resident)≈ 2.6 hvs 7.5 h reloading each time
fp32 per diffusion step52 → 105 → 350 scatastrophic slowdown after VRAM overflow
Measured with fp16 on an idle GPU. Loading, not inference, dominates small-batch latency; fp32 needs ~7.8GB and spills into shared system memory.
  • The resident worker is the biggest lever: load once, loop on requests, reducing per-line cost to pure inference.
  • fp16 is a requirement on 8GB, not an option: fp32 overflows VRAM and degrades catastrophically. Speedups must come from fewer diffusion steps or more VRAM — not higher precision.
  • Latency is extremely sensitive to machine idleness: under contention, per-step time jumped from ~3.5 s to ~240 s (~70×). Production runs belong on an idle machine or off-peak batch.

Data efficiency: how much recording is enough?

"How long do I have to record?" is the first question an amateur asks. I trained one version per subset of the same speaker's data (zero-shot / 5 / 10 / 20 / 28 lines) and measured on the same held-out lines:

Training dataSpeaker cosine (vs real centroid)Held-out CER (single roll)
Zero-shot (1 reference line)0.767~9%
5 lines (~30 s)0.885~9%
10 / 20 / 28 lines0.886 / 0.889 / 0.878 (flat)7%–32% (high variance)
Timbre similarity saturates at roughly 5 lines; pronunciation accuracy is independent of data volume.

Finding 1: timbre similarity saturates at ~5 lines (30 seconds) — "sounding like yourself" is cheap. Additional material buys emotion references, breath assets, and stability, not timbre.

Finding 2: pronunciation accuracy is unrelated to data volume and is dominated by sampling randomness — the same model on the same line varies between 3% and 58% CER across seeds. Methodological warning: a single CER measurement cannot judge a model (two bad rolls once led me to wrongly conclude a training version was broken; the third seed was nearly perfect). → Production stability comes from per-segment ASR verification and rerolling, not more data.

This curve translated directly into a tiered recording script for users: 1 line to play with, 5–10 lines to sound like yourself, a full set for emotion and breaths.

Methodology: two self-corrections

What this project taught me matters more than its conclusions. Twice I was wrong, and measurement is what corrected me.

  1. The "40 minutes per line" mystery = external contention, not the engine

    Measured at 40 min/line while a game held the GPU at ~100%; re-tested on an idle GPU, the same line took 270 s. Had I concluded "this engine is too slow", I would have cut the model that later became my specialist engine. Lesson: rule out environmental variables before attributing cause — measure, don't guess.

  2. The "28-line version is broken" = sampling variance, not the model

    Two consecutive bad samples led me to declare a training version broken. A third seed was nearly perfect. That correction directly motivated the ASR-QC retry pipeline — turning the insight "a single result is untrustworthy" into a component of the system.

  3. Attributing an intelligibility problem: text case + reference audio, not precision

    A native listener judged one line unclear and I suspected fp16 or an engine defect. Controlled elimination found the real causes: (1) that line was a short, high-emotion exclamation — a specific text case, not systemic; (2) I was using synthetic speech as the cloning reference — a copy of a copy again. Rule out controlled variables before blaming the model or precision, or you will cut fp16, which this hardware requires.

Threats to validity

This is a single-machine case study with small samples. I state the limits here rather than waiting to be asked — the results should be read as engineering decision evidence plus a reproducible experimental design, not population-level statistical inference.

Threats to Validity

  • Sample size: n = 2 speakers, 3 held-out lines per condition, 1 listener; the data-efficiency curve is a single-speaker case. Scaling up (speakers, sentences, listeners, multi-seed statistics) is the first task in turning this into a formal paper.
  • Judge bias: a single ASR judge (Whisper large-v3) whose language-model correction systematically underestimates mispronunciation errors. Cross-checking with a second judge (e.g. Paraformer) would mitigate this.
  • Single-sample variance: CER under identical conditions ranges 3%–58%; every single-run number here is a point estimate.
  • Semantic side effects of homophone substitution: substituted text is read by the engine's BERT front-end, and the prosodic impact of the semantic change is unquantified; the mapping is based on primary readings and does not model tone sandhi.
  • Licensing analysis is not legal advice: engine licence terms were read from an engineering perspective and need legal review before commercial release.
  • Hardware extrapolation: all performance conclusions are bound to a single 8GB Turing GPU and may not hold on larger or newer hardware.

Security posture

  • No eval / exec / pickle / os.system / shell=True; every subprocess call uses argument lists (no shell injection surface).
  • The resident worker communicates over stdin/stdout pipes as a child process, not a network service.
  • HTTP services bind 127.0.0.1 only, never 0.0.0.0 — the upstream API's /tts endpoint can read arbitrary files and carries RCE risk.
  • Only load models from trusted sources: .pth / .ckpt files are pickle deserialization and can execute arbitrary code.
  • The commercial risk is not engine licensing (MIT / permissive) but whose voice is cloned — real voices require the person's consent.

Tech stack

PythonGPT-SoVITSIndexTTS-2MOSS-VoiceGeneratorKokorofaster-whisperpypinyinERes2NetV2PyTorchsubprocess bridging / multi-venv isolationFastAPIEBU R128 loudnorm

Future work

  1. Larger samples and multi-judge evaluation — closing the statistical gaps listed under Threats to Validity is the first step toward turning this report into a formal paper.
  2. A library of acting presets — pre-recorded professional deliveries so users can pick a performance instead of acting it themselves: record 30 seconds, choose a delivery, get professional output.
  3. Cloud fine-tuning for IndexTTS-2 — so designed voices can also win on all three axes with the specialist engine; no off-the-shelf tooling exists, so this is DIY training engineering.
  4. Real-time voice conversion and singing synthesis — reusing the existing recording assets for interactive scenarios.

Further reading