Post

EMA Lightning Web: Lightning-Fast Offline Turkish TTS in Just 36 MB

🇬🇧 Lightning-fast offline Turkish TTS in just ~36 MB — serverless and privacy-first, running entirely in the browser.

EMA Lightning Web: Lightning-Fast Offline Turkish TTS in Just 36 MB

Continuing my in-browser Turkish TTS work: as a sibling to Antalia Mini Web, this time I ported a lighter and faster model — EMA Lightning — to the web.

Meet EMA Lightning Web: an open-source app that converts text to speech fully in the browser with onnxruntime-web + WebAudio — no server involved. | 🇹🇷 Türkçe

▶️ Live Demo: fr0stb1rd.github.io/ema-lightning-web | 💻 Repo: github.com/fr0stb1rd/ema-lightning-web

What Is EMA Lightning Web?

EMA Lightning Web is a browser port of canberk7/ema-lightning (Canberk Aslan): 8.6M parameters (~34 MB), Apache-2.0, 0.92% WER on Freya-TR-Eval. The original PyTorch pipeline (ema-lightning package) is split into three ONNX stages for streaming: playback starts with the first piece while the rest still generate.

Models download once on first launch (~36 MB), then run with zero bandwidth from the browser cache for a day. Your text and voice never leave the device.

✨ Highlights

  • Light and fast: ~36 MB total; desktop CPU Python is ~6× realtime, ~1–2× on short sentences in WASM. Long text streams piece by piece while the first piece plays.
  • Streaming pipeline: text_stage → plan (JS) → sound_stage (4-step DiT → 25 Hz latents) → decoder (HiFi-GAN → 48 kHz).
  • No UI freeze: inference runs in a worker via ort.env.wasm.proxy = true, with main-thread fallback. Providers ['webgpu', 'wasm'].
  • Smart cache: models are cache-first in Cache Storage (ema-lightning-web-v1, 1-day TTL). vocab.json (778 bytes) ships same-origin.
  • History + lock screen: generations in memory, metadata in localStorage; lock-screen controls via Media Session API.
  • HuggingFace Hub insight: why Hub for weights? GitHub Releases sends no Access-Control-Allow-Origin, so browsers refuse the download (measured); embedding into the repo adds ~36 MB to git history on every re-export. *.cdn.hf.co sends Access-Control-Allow-Origin: *.

🏗️ Architecture

flowchart LR
    A["Turkish text"] --> B["app.js frontend\nchunk + alphabet"]
    B --> C["text_stage.onnx\nletters → features + durations"]
    C --> D["plan (JS)\ndurations → frame timeline"]
    D --> E["sound_stage.onnx\n4-step DiT → 25 Hz latents"]
    E --> F["decoder.onnx\nHiFi-GAN → 48 kHz audio"]
    F --> G["WebAudio playback\n.wav download"]
StageInputOutput
text_stage.onnxids [B,L], mask [B,L]h [B,L,224], dur [B,L]
plan (JS, ported from engine.py)dur, word mapfw [T], fp [T]
sound_stage.onnxh, dur, masks, maps, noiselatents [B,T,64]
decoder.onnxz [B,64,T]audio [B,S] @48 kHz

Frame-planning math (duration rounding, repeat_interleave, intra-word positions) and the windowed decoder (1 s first window with 8-frame context, then 4 s windows) are ported 1:1 from engine.py.

🚀 Usage

  1. Open the live site.
  2. Type text, press Speak.
  3. Download as .wav if you like.

🔢 Note: no automatic number/date reading in this version — the original normalizer-tr package is not in the browser; app.js only lowercases + filters the alphabet. Inputs like 1.250.000 TL are read digit by digit, so write numbers out (bin iki yüz elli lira). If you need number reading, see Antalia Mini Web.

🛠️ Developer Notes

Fully automated build: Actions → export-onnx → Run workflow. export_onnx.py (torch.export, dynamo=True) produces the models, --check numerically verifies every stage against onnxruntime (text+durations 1e-4, latents 1e-3, audio 1e-4), push_hf.py uploads to Hub, pages.yml publishes via workflow_run.

Prerequisite (once): an ema-lightning-web-onnx Hub repo + HF_TOKEN under repo Settings → Secrets → Actions.

⚠️ Limitations

  • No number/date reading (see above).
  • No seed parity: the browser uses its own PRNG, so the same text yields valid but not bit-identical audio.
  • Single voice, Turkish only — the original model’s limits apply as-is.

📄 License and Credits

Inspect the code and contribute: fr0stb1rd/ema-lightning-web

This post is licensed under CC BY 4.0 by the author.