sonixdocs

Voices

GET /v2/voices

Lists the voices of the active speech model in the paginated ElevenLabs shape:

curl "$SONIX_URL/v2/voices" -H "xi-api-key: $SONIX_API_KEY"
{
  "voices": [
    {
      "voice_id": "af_alloy",
      "name": "Alloy",
      "category": "premade",
      "labels": { "gender": "female", "language": "en" }
    }
  ],
  "has_more": false,
  "total_count": 1,
  "next_page_token": null
}

The legacy GET /v1/voices returns the same list as { "voices": [...] }.

The catalog reflects what the deployment you target actually serves; list voices first to discover what your endpoint offers.

Voice resolution

Wherever a voice is named (the voice field on /v1/audio/speech or the {voice_id} path segment on the ElevenLabs surface), sonix resolves it in order:

  1. An exact match against the model's voice catalog (af_alloy), then a case-insensitive match, then a catalog display-name match (Alloy).
  2. The literal default selects the model's default voice.
  3. An OpenAI-style name (alloy) maps to the catalog ID with the same suffix (af_alloy) when one exists.

So OpenAI's stock voice names work out of the box on models whose catalogs follow that convention. Exact catalog IDs always win.

Default voices

Every voice below reads the same line, so what you hear between two rows is the voice and nothing else. Press ↑ and ↓ to step through a catalog, playing each voice as it lands.

kokoro-82m

Style-vector voices covering American and British English.

kokoro-82m 24 voices · default af_heart

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderqualitystylelen
af_heartdefaultHearten-USfemaleAneutral4.9s
af_alloyAlloyen-USfemaleCneutral4.9s
af_aoedeAoedeen-USfemaleC+neutral5.2s
af_bellaBellaen-USfemaleA-expressive5.0s
af_jessicaJessicaen-USfemaleDneutral4.2s
af_koreKoreen-USfemaleC+neutral5.0s
af_nicoleNicoleen-USfemaleB-whisper7.4s
af_novaNovaen-USfemaleCneutral4.6s
af_riverRiveren-USfemaleDneutral4.2s
af_sarahSarahen-USfemaleC+neutral4.9s
af_skySkyen-USfemaleC-neutral4.8s
am_adamAdamen-USmaleF+neutral4.8s
am_echoEchoen-USmaleDneutral4.9s
am_ericEricen-USmaleDneutral4.2s
am_fenrirFenriren-USmaleC+neutral5.1s
am_liamLiamen-USmaleDneutral4.4s
am_michaelMichaelen-USmaleC+neutral5.5s
am_onyxOnyxen-USmaleDneutral4.5s
am_puckPucken-USmaleC+neutral4.9s
bf_emmaEmmaen-GBfemaleB-neutral5.5s
bf_isabellaIsabellaen-GBfemaleCneutral4.9s
bm_georgeGeorgeen-GBmaleCneutral4.7s
bm_lewisLewisen-GBmaleD+neutral4.7s
bm_danielDanielen-GBmaleDneutral4.8s

↑↓ steps through voices and plays each · esc stops

OpenAI's alloy, echo, onyx and nova resolve by suffix. fable, shimmer, ash, ballad, coral, sage and verse have no counterpart here; they return a validation error rather than silently remapping to a different voice.

kitten-tts-nano-0.8, kitten-tts-micro-0.8 and kitten-tts-mini-0.8

All three sizes share the same eight voice ids and display names, but they are different models and do not sound the same — mini reads the same line more slowly than nano, with micro between them. Compare the tables before picking a size.

kitten-tts-nano-0.8 8 voices · default expr-voice-2-f

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
expr-voice-2-fdefaultBellaenfemaleexpressive5.8s
expr-voice-2-mJasperenmaleexpressive3.0s
expr-voice-3-fLunaenfemaleexpressive3.3s
expr-voice-3-mBrunoenmaleexpressive3.5s
expr-voice-4-fRosieenfemaleexpressive3.8s
expr-voice-4-mHugoenmaleexpressive3.3s
expr-voice-5-fKikienfemaleexpressive3.0s
expr-voice-5-mLeoenmaleexpressive4.4s

↑↓ steps through voices and plays each · esc stops

kitten-tts-micro-0.8 8 voices · default expr-voice-2-f

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
–expr-voice-2-fdefaultBellaenfemaleexpressive—
–expr-voice-2-mJasperenmaleexpressive—
–expr-voice-3-fLunaenfemaleexpressive—
–expr-voice-3-mBrunoenmaleexpressive—
–expr-voice-4-fRosieenfemaleexpressive—
–expr-voice-4-mHugoenmaleexpressive—
–expr-voice-5-fKikienfemaleexpressive—
–expr-voice-5-mLeoenmaleexpressive—

↑↓ steps through voices and plays each · esc stops · 8 without a sample

kitten-tts-mini-0.8 8 voices · default expr-voice-2-f

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
expr-voice-2-fdefaultBellaenfemaleexpressive7.4s
expr-voice-2-mJasperenmaleexpressive4.2s
expr-voice-3-fLunaenfemaleexpressive5.2s
expr-voice-3-mBrunoenmaleexpressive4.7s
expr-voice-4-fRosieenfemaleexpressive5.1s
expr-voice-4-mHugoenmaleexpressive4.5s
expr-voice-5-fKikienfemaleexpressive4.6s
expr-voice-5-mLeoenmaleexpressive6.1s

↑↓ steps through voices and plays each · esc stops

qwen3-tts-0.6b

Nine preset speakers baked into the model. Voice cloning is not available. Address each speaker in its native language for best results — which is why the Chinese, Japanese and Korean speakers below are auditioned in their own language rather than in English.

qwen3-tts-0.6b 9 voices · default aiden

en “Every voice reads the same line, so the only thing that changes is how it sounds.”zh “每个声音都朗读同一句话,所以唯一改变的是它听起来的感觉。”ja “すべての声が同じ文章を読みます。変わるのは響きだけです。”ko “모든 목소리가 같은 문장을 읽습니다. 달라지는 것은 소리뿐입니다.”

idnamelanguagegenderaccentstylelen
aidendefaultAidenen-USmale—sunny American5.5s
ryanRyanenmale—dynamic, rhythmic8.0s
vivianVivianzhfemale—bright, young8.1s
serenaSerenazhfemale—warm, gentle7.7s
uncle_fuUncle Fuzhmale—seasoned, mellow11.7s
dylanDylanzhmaleBeijing dialectyouthful10.1s
ericEriczhmaleSichuan (Chengdu) dialectlively8.8s
ono_annaOno Annajafemale—playful6.1s
soheeSoheekofemale—warm8.5s

↑↓ steps through voices and plays each · esc stops

chatterbox

The catalog's single Default voice uses the checkpoint's pinned conds.pt conditioning. The upstream checkpoint does not document whose voice it is, so it is not presented as a named person; the OpenAI alias alloy resolves to the same conditioning. This model is English only. It returns a complete 24 kHz mono WAV rather than live streaming, and every output carries the upstream Perth watermark.

chatterbox 1 voices · default default

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagestylelen
defaultdefaultDefaultenpinned conditioning4.0s

↑↓ steps through voices and plays each · esc stops

pocket-tts

Ten predefined voices baked into the model. This model is English only. Voice cloning is not available: the deployed bundle is upstream's no-cloning release, so these ten voice states are the entire surface. The default, Alba, is built from Alba MacKenna's recordings (CC BY 4.0); the other nine come from CC0 sources (Kyutai voice donations, LibriVox) and CC-BY-4.0 VCTK recordings.

pocket-tts 10 voices · default alba

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
albadefaultAlbaenfemalecasual, conversational4.5s
azelmaAzelmaen—VCTK speaker5.7s
bill_boerstBill BoerstenmaleLibriVox reader4.8s
caro_davyCaro Davyen—LibriVox reader5.8s
eponineEponineen—VCTK speaker5.3s
fantineFantineen—VCTK speaker4.8s
javertJaverten—Kyutai voice donation (Butter)6.5s
mariusMariusen—Kyutai voice donation (Selfie)3.8s
peter_yearsleyPeter YearsleyenmaleLibriVox reader4.4s
stuart_bellStuart BellenmaleLibriVox reader5.9s

↑↓ steps through voices and plays each · esc stops

neutts-nano-q4

Voices are cloned from reference recordings. This model is English only. Upstream ships multilingual NeuTTS as separate per-language checkpoints; the one served here is the English base, so send it English text.

Upstream publishes nine reference speakers; the six English ones are offered here. The other three are cloned from German, French and Spanish recordings, and they are omitted rather than sold as accented English: NeuTTS phonemizes a voice's reference transcript with the same English-only phonemizer it uses for your text, so a non-English reference degrades the clone no matter what you ask it to say.

neutts-nano-q4 6 voices · default jo

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
daveDaveenmaleneutral3.9s
emilyEmilyenfemaleneutral4.4s
jodefaultJoenfemaleneutral4.3s
paulPaulenmaleneutral4.2s
sophieSophieenfemaleneutral5.7s
stevenStevenenmaleneutral4.1s

↑↓ steps through voices and plays each · esc stops

neutts-air-q4

The same reference-cloning mechanism as nano, on a larger model. This model is English only. Two reference voices are offered: jo, Nano's default, and dave, another Nano catalog voice.

neutts-air-q4 2 voices · default jo

“Every voice reads the same line, so the only thing that changes is how it sounds.”

idnamelanguagegenderstylelen
–jodefaultJoenfemaleneutral—
–daveDaveenmaleneutral—

↑↓ steps through voices and plays each · esc stops · 2 without a sample

Resold vendor models

The three vendor models are resold through the gateway rather than hosted on the fleet, so their voice surfaces are the vendors' own and no auditioning samples are committed for them.

elevenlabs/eleven_multilingual_v2

The voice field is the vendor's own voice id, passed through unchanged; the vendor's catalog is authoritative. The six premade stock voices (Bella, Sarah, Roger, Laura, Charlie, George) are what the playground bench lists — a sample of the vendor's full catalog rather than the whole surface.

fish/s2.1-pro-free

The voice field is a Fish reference id, passed through unvalidated. There is no voice listing route for this model.

gradium/default

The voice_id passes through to the vendor. There is no ElevenLabs-shaped voice listing route for this model.