Antalia Mini Web: Offline Turkish TTS That Reads Numbers, Dates and Abbreviations Aloud
🇬🇧 Offline Turkish TTS running entirely in the browser, with a smart normalizer that correctly reads numbers, dates and abbreviations.
Almost every Turkish text-to-speech tool depends on a paid API or sends your voice to a server. If you want one-click speech synthesis with privacy, offline support, and zero setup, the options are bleak. To close that gap, I ported the Antalia-2 Mini model to run fully in the browser.
Meet Antalia Mini Web: an open-source web app that turns your text into speech on-device with WebGPU/WASM — nothing ever leaves your machine. | 🇹🇷 Türkçe
▶️ Live Demo: fr0stb1rd.github.io/antalia-mini-web | 💻 Repo: github.com/fr0stb1rd/antalia-mini-web
What Is Antalia Mini Web?
Antalia Mini Web is a browser port of cloud0day3/antalia-mini (PatientDesk AI, Antalia-2 Mini): 7.62M parameters, 48 kHz output, one synthetic male voice, Apache-2.0. The PyTorch pipeline is converted to ONNX and executed with onnxruntime-web + WebAudio. No server, no API keys.
The first launch downloads ~50 MB once (3 ONNX models + 4 dictionaries); afterwards it works fully offline via Cache Storage.
✨ Highlights
- 1:1 Turkish normalizer (
normalize.js): a line-by-line port of the original Python normalizer. Reads numbers (1250), decimals (68,5), ordinals (2.), dates (15.10.2026), times (14:30), percentages (%18), money (1.250,75 TL→ lira + kuruş), phones, e-mails/URLs, abbreviations (Dr.→doktor,SGK→segeka,THY→teheye), Roman numerals, and English/brand words (GitHub→git hab). Matches Python output on 9,000+ inputs, verified by tests on every CI run. - Identical tone & pauses (
audio.js): release tone path (+6 dB low-shelf at 210 Hz, −9 dB high-shelf at 4 kHz, soft limiter) and inter-chunk pauses (0.12 s after sentences, 0.06 s otherwise), sample-identical to the original within float32 tolerance. - 49 showcase examples in 9 categories (numbers, money, phones, brands, abbreviations…).
- Streaming synthesis: sentences are produced piece by piece; the first one plays immediately while the rest generate in the background, with live word highlighting.
- Flow-Matching + Shortcut Self-Consistency: 8-step DiT inference with CFG 2.0, baked into the ONNX model at export. Re-export with
--steps 4or--steps 2for faster variants. - Zero bandwidth on repeat visits: models and dictionaries cached via the Cache Storage API (1-day TTL).
- Privacy: 100% client-side; text and audio never leave the device.
🏗️ Architecture
The PyTorch pipeline is split into three ONNX stages for browser streaming:
flowchart LR
A["Raw Turkish Text"] --> B["normalize.js\nNumbers, dates, brands"]
B --> C["text_stage.onnx\nLetters → Features + Durations"]
C --> D["Frame Planning (JS)\nDurations → Timeline"]
D --> E["sound_stage.onnx\nDiT Flow-Matching → Mel"]
E --> F["decoder.onnx\nConvNeXt Vocoder + iSTFT → 48 kHz"]
F --> G["audio.js\nTone EQ + Pauses → WebAudio + .wav"]
| Stage | Input | Output |
|---|---|---|
text_stage.onnx | ids [B,L], mask [B,L] | h [B,L,192], logd [B,L] |
| Frame planning (JS) | durations | frame timeline |
sound_stage.onnx | cond, fmask, noise | mel [B,128,T] (5-block DiT, 8 steps, CFG 2.0) |
decoder.onnx | mel | audio [B,S] @48 kHz |
Browser dictionaries built at export and served from HuggingFace: foreign_dict.json (80k phonetic respellings), foreign_parts.json (37k CamelCase parts), known_words.json (6k anti-respell shield), english_i.json (14k dotted/dotless-I decisions).
🚀 Usage
- Open the live site.
- Type text, press Speak. The first sentence plays instantly.
- Optionally download the result as
.wav.
No installation; any modern browser works (faster with WebGPU, WASM fallback otherwise).
🛠️ Developer Notes: Automated CI/CD
Everything is automated with GitHub Actions: export-onnx.yml produces 3 ONNX models + dictionaries from safetensors, validates with ONNX checker + smoke test + normalizer/audio differential tests, uploads via push_hf.py to HuggingFace Hub (fr0stb1rd/antalia-mini-web-onnx), and pages.yml deploys to GitHub Pages.
1
2
python3 tests/make_fixtures.py # regenerate oracle fixtures from Python
node --test tests/test_normalize.mjs tests/test_audio.mjs
One-time setup is creating the antalia-mini-web-onnx Hub repo and adding HF_TOKEN; afterwards it is one click on Run workflow.
📄 License and Credits
- Model weights and original code: Apache-2.0 (cloud0day3/antalia-mini, PatientDesk AI — NOTICE attribution included verbatim).
- Web app code in this repo (
app.js,normalize.js,audio.js, …): Apache-2.0 © 2026 fr0stb1rd.
Inspect the code and contribute: fr0stb1rd/antalia-mini-web — sibling project: EMA Lightning Web
