WORDS INTO VOICE

Text to Speech

Create English narration with 28 voices. Listen to instant voice samples, shape the pacing, then generate and download audio on your device.

Free · no API keyLocal speech generationWAV & MP3
01 · FIND YOUR VOICE

Listen before you generate

Samples are short pre-generated clips at 1× speed. No model download needed.

Start here: click ▶ Preview to listen, then Use voice → to choose.

Choose Preview on a voice card.
02 · WRITE YOUR SCRIPT

What should your voice say?

0 / 5,000 characters

Ready. Voice previews play instantly; your script requires the speech model.

First CPU use downloads about 88 MiB of model data, plus runtime and voice files. Internet is needed for uncached files. Generation speed depends on your device; short tests help.

04 · LISTEN & KEEP IT

Your generated audio

No audio generated yet

Generate the script or its first part. Downloads become available after processing.

Subtitles use generated segment boundaries, not word alignment. A ZIP includes audio, the generated script, settings and timing. MP3 resamples to 48 kHz for encoding and adds no extra speech detail.

Generated parts & timing

    Using Text to Speech

    Generate English speech on your device with Kokoro voices. Listen to pre-generated samples before loading the speech model.

    1. Preview voices, filter by American or British accent and choose a voice. Voice types describe the published model voices, not gender detection.
    2. Write or open a script up to 5,000 characters. Choose speaking speed, extra sentence/paragraph pauses and output level. Try the first part before generating a long script.
    3. Generate, listen and download WAV or MP3. A project ZIP includes audio, the generated script, settings and segment timing; save a script project separately for later editing.

    Example

    Choose Heart, keep speed at 1× and use the Welcome demo.
    Preview the voice instantly, then choose Try first part to check how it reads your own text.

    Questions & answers

    Is Text to Speech free?

    Yes. There is no speech API key or paid inference service. The model runs on your device; downloading model files still uses your connection and bandwidth.

    Is my script sent to a speech server?

    Our tool sends the script to a local browser worker. It downloads model and voice files from Hugging Face, but does not send the script to a speech inference API. Advertising and other website scripts are separate from the speech processing.

    Why is the first generation slow?

    The CPU model is about 88 MiB plus runtime and voice files. Your browser may cache downloaded files. Speech generation itself uses your device, and can take longer than the final audio duration. Preview clips do not load the model.

    Does it work offline?

    Already loaded model and voice files can be reused in the open page without internet. Uncached voices or model files need a connection. Reopening or refreshing offline is not guaranteed; this is not an installed offline app.

    Which languages can I use?

    This version supports English with American and British voices. The underlying model has other variants, but this browser integration does not claim those languages.

    Will GPU mode be faster?

    It may be faster with a compatible WebGPU device. It uses a larger full-precision model download. CPU mode works on more devices. If GPU initialization fails, select CPU and retry; no server GPU is used.

    Can I clone a voice or choose emotion?

    No. These are fixed model voices. Speed and pauses change delivery, but the tool does not clone a person or guarantee happy, sad or dramatic emotion.

    Are subtitles word-aligned?

    No. SRT and JSON describe generated segment start/end times. They include inserted gaps in the timeline, but do not provide word-level forced alignment.

    What do downloads contain?

    WAV is 24 kHz mono 16-bit PCM. MP3 is encoded at 192 kbps using audio resampled to 48 kHz; resampling adds no speech detail. MP3 may have encoder padding.

    Are my scripts saved automatically?

    No. Favorites and explicitly saved settings use browser storage. Script projects are JSON downloads you choose to save; audio and script content are otherwise kept in memory until replaced or the page closes.

    Help improve this tool

    Report a problem or suggest an improvement

    Describe the issue without pasting private tool input. Feedback goes to our admin inbox.

    Find another tool · Read practical guides