Skip to content
Browser TTS workspace

Kokoro TTS

Free AI text to speech with Kokoro TTS. Browser local TTS audio synthesis with 54 voices across 9 languages and WebGPU or WASM inference.

Private generation WAV + MP3 export Default TTS workspace
305-326
q4 / fp32
MB model
54
9 languages
voices
82M
StyleTTS 2
params
WebGPU
+WASM fallback
GPU+CPU

TTS works best on desktop

You can still try lightweight engines on mobile, but desktop Chrome or Edge remains the most reliable setup for large model downloads and long-form generation.

About Kokoro TTS

Kokoro TTS is the default engine on OfflineTTS, offering 54 voices across 9 language variants including English, Japanese, Chinese, Spanish, French, Hindi, Italian, and Portuguese. It makes OfflineTTS a practical Local TTS workspace for browser audio synthesis without requiring an account or hosted synthesis API.

English uses the local kokoro-js phonemizer after required assets download. Non-English Kokoro sends entered text to the documented phonemization service, receives pronunciation tokens, and then generates the waveform on your device.

It supports two model sizes: q4 (~305MB, recommended) and fp32 (~326MB, full precision). q8 quantization is not available as it produces garbled audio with this model.

Compare engines: Kitten TTS (8 voices, 24MB, lightest) ยท Piper TTS (25 voices, fastest CPU) ยท Supertonic TTS (31 languages, local inference) ยท Pocket TTS (8 voices, cloning)

Kokoro Data Path and Practical Limits

Language determines the network path

Model and voice files must be downloaded before synthesis. English text is phonemized in the browser after those assets are available. Japanese, Chinese, Spanish, French, Hindi, Italian, and Brazilian Portuguese text is sent to api.offlinetts.com for phonemization; audio inference and export remain local. Hosting, analytics, and model delivery are documented separately in the Privacy Policy.

Review before a long or public export

Generate a sample containing the final script's names, numbers, abbreviations, quotations, and longest sentence. Compare q4 and fp32 only if the audible result justifies the larger precision choice. Listen for pronunciation, clipped chunk boundaries, repeated words, and pacing, then keep the voice ID, precision, speed, browser, and reviewed source script with an important production file.

Getting Started with Kokoro TTS

New to AI text to speech? Use this sequence to create and review a representative Kokoro sample before generating a longer file.

1. Choose Your Model Size

Start with q4 (~305MB) when download size matters. FP32 (~326MB) retains full numerical precision. Compare both on the same text before assuming an audible benefit.

2. Pick a Voice

Filter by language and compare the same passage in several voices. Catalog grades and traits are discovery labels, not a guarantee for every script.

3. Write Your Script

Use proper punctuation โ€” commas add pauses, periods create full stops, question marks raise pitch. Well-punctuated text produces the most natural speech.

4. Generate & Download

Click generate, wait for the audio to play, then download as WAV (lossless) or MP3 (compressed). WAV is recommended for further editing.

Tips for Best TTS Quality

1.

Compare the available backend. WebGPU can be faster on compatible hardware, but results vary by browser, model, GPU, driver, and memory. Start with automatic selection and record the environment when performance matters.

2.

Punctuate properly. This is the single most important factor for natural-sounding speech. Commas, periods, question marks, and exclamation marks all create distinct prosodic effects.

3.

Break long text into paragraphs. The tool handles up to 50,000 characters, but shorter paragraphs with clear punctuation produce better rhythm and pacing.

4.

Try multiple voices. Different voices suit different content types. Heart excels at warm narration, Bella at energetic delivery, Michael at professional reviews.

5.

Use WAV for production. WAV preserves full audio quality for editing. MP3 is fine for quick sharing, but use WAV if you plan to mix, master, or further process the audio.

Kokoro TTS Browser Architecture & Best Practices

OfflineTTS Kokoro workspace runs cutting-edge speech synthesis models client-side. Unlike legacy cloud text-to-speech providers that bill per character and transmit sensitive text to remote computing clusters, client-side inference keeps your drafts strictly private on your workstation.

Local Device Execution

Neural inference executes inside a dedicated Web Worker using ONNX Runtime Web. Once model weights (approx. 305MB for quantized q4 or 326MB for FP32) are stored in your browser's IndexedDB, subsequent synthesis operations require zero Internet connectivity for English voices.

Chunking & Memory Management

To prevent browser tab crashes during long-form narration, the system automatically splits scripts exceeding 500 characters into natural clause and sentence chunks. Chunks are synthesized sequentially, then stitched smoothly in audio buffers before playback.

When generating voice-overs for commercial multimedia, keep track of your selected voice preset, speed factor, and punctuation tweaks. Because browser-based generation is deterministic for identical seeds and settings, saving your source script ensures you can reproduce matching audio pickups at any time.

Frequently Asked Questions about Kokoro TTS

How does Kokoro TTS run in my browser without a server?

Kokoro TTS is powered by ONNX Runtime Web compiled to WebAssembly (WASM) and WebGPU. On your first visit, the AI model weights are downloaded from a content delivery network and stored in your browser's local IndexedDB cache. Once cached, all audio synthesis calculations run locally on your device's CPU or GPU.

Is my text or generated audio uploaded to any server?

No. For English text-to-speech, both text phonemization and speech synthesis occur entirely inside your browser. For non-English languages (Japanese, Chinese, Spanish, French, Hindi, Italian, Portuguese), the text string is converted to pronunciation phonemes via an API endpoint, but no audio is uploaded or stored on our servers.

What is the difference between WebGPU and WebAssembly (WASM)?

WebGPU leverages your computer's graphics hardware (GPU) to accelerate neural network tensor operations, often generating speech 3 to 10 times faster than real-time on modern laptops and desktops. WebAssembly (WASM) runs on your computer's CPU and serves as a universal, reliable fallback across all modern browsers.

Can I use generated speech audio for commercial projects?

Kokoro TTS is an open-weight model licensed under the Apache 2.0 license. In general, audio synthesized with Apache 2.0 open-weight models may be used for personal and commercial projects including YouTube videos, podcasts, and audiobooks, subject to compliance with upstream terms and applicable laws.

How can I improve speech naturalness and pronunciation?

Punctuation is critical for neural TTS prosody. Use commas to introduce brief breathing pauses, periods for complete cadence drops, and question marks for rising terminal pitch. For difficult acronyms or proper nouns, spelling them out phonetically or separating syllables with hyphens often yields superior naturalness.

Should I export my generated speech as WAV or MP3?

WAV exports provide uncompressed 24kHz 16-bit PCM audio, making it the ideal format for video editing in Premiere or Final Cut, podcast production in Audacity, or audio mastering. MP3 encoding occurs locally in your browser using lamejs and offers a compact file size ideal for quick listening and web distribution.