VocalLab APIBase URL https://api.vocallab.netOpenAPI

Introduction

The VocalLab API follows the GenMax API: the same paths, request fields, response fields, status codes and pagination. Code written for GenMax works after changing the base URL to https://api.vocallab.net and using a VocalLab API key. Create keys in the studio under API.

Authentication

Send your key in the xi-api-key header on every request. Keys start with vl_live_. Missing or invalid keys return 401 {"error":"invalid api key"}.

Asynchronous tasks

Generation endpoints return 202 with an id or task_id and "status":"pending". Poll the matching history endpoint every 2 seconds until status is completed or failed. Statuses are pending, processing, completed and failed.

To make a create request safe to retry after a network error, send an Idempotency-Key header (12–120 characters). The same key returns the original task instead of creating and charging a second one.

Credits

Credits are reserved when a task is accepted. A failed task is refunded in full; a completed task is charged its actual cost, at most the reserved amount. credits_deducted shows the reserved amount while running and the final charge afterwards. Check your balance with Get user, credits.

Errors

Errors use {"error": "…"}, sometimes with message for detail. 402 responses include required, the credits the request needs. Common codes: 400 invalid request, 401 invalid key, 402 insufficient credits, 404 not found or owned by another account, 409 conflicting state, 422 unusable source, 429 rate limit, 503 tool not enabled or provider unavailable.

Pagination

History lists take page (zero-based) and page_size (default 30, maximum 100) and return has_more. MiniMax and CapCut voice libraries use one-based page and also return total.

Result files

audio_url, srt_url, preview_url and similar fields are signed links that work without headers, for example in an <audio> element. They expire after about 24 hours; read the task again for a fresh link. Results are kept for 30 days.

Rate limits

Each key allows 120 requests per minute and has a daily request quota (see quota in Get user). Successful responses carry X-RateLimit-Limit, X-RateLimit-Remaining and X-RateLimit-Reset; 429 responses carry Retry-After.

Differences from GenMax

  • GET /v1/api-keys returns a masked key: VocalLab stores only a hash. The full key is shown once, on creation or regeneration.
  • Retry endpoints create a new task and return its ID. Poll the returned ID.
  • Tools whose pricing is not enabled return 503.
  • Extensions: optional Idempotency-Key, output_format for text to speech, and the VocalLab extension endpoints for quotes, batch scenes and SRT voice-over.

Auth

Get user, credits...

GET/v1/auth/me

Returns the account that owns the API key, including its remaining credit balance. key and quota describe the calling key and its daily request quota.

curl -X GET "https://api.vocallab.net/v1/auth/me" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "id": "550e8400-e29b-41d4-a716-446655440000",
  "email": "user@example.com",
  "name": "John Doe",
  "avatar_url": null,
  "credit_balance": 4000,
  "role": "user",
  "key": {
    "id": "0c6b1a2e-…",
    "prefix": "vl_live_AbCdEfGh",
    "name": "Production"
  },
  "quota": {
    "dailyLimit": 1000,
    "used": 12,
    "remaining": 988,
    "resetsAt": "2026-10-12T00:00:00.000Z"
  }
}

API Keys

Get API key

GET/v1/api-keys

Describes the calling API key. VocalLab stores only a hash of each key, so api_key is masked; the full key is shown once when it is created or regenerated.

curl -X GET "https://api.vocallab.net/v1/api-keys" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "api_key": "vl_live_AbCdEfGh…",
  "id": "0c6b1a2e-…",
  "name": "Production",
  "created_at": "2026-10-11T07:00:00.000Z",
  "last_used_at": "2026-10-11T08:00:00.000Z"
}

Regenerate API key

POST/v1/api-keys/regenerate

Revokes the calling key immediately and returns its replacement. Store the new key; it is not shown again.

curl -X POST "https://api.vocallab.net/v1/api-keys/regenerate" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "api_key": "vl_live_…",
  "message": "API key regenerated successfully"
}

Models

Get models

GET/v1/models

Text to speech models enabled on VocalLab. Defaults to ElevenLabs; pass provider=minimax or provider=capcut for other providers.

Query parameters

NameTypeDescription
providerstringelevenlabs (default), minimax or capcut.
curl -X GET "https://api.vocallab.net/v1/models" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 ElevenLabs models (default)

[
  {
    "model_id": "eleven_multilingual_v2",
    "name": "Eleven Multilingual v2",
    "description": "",
    "can_do_text_to_speech": true,
    "can_do_voice_conversion": false,
    "maximum_text_length_per_request": 5000,
    "token_cost_factor": 1,
    "model_rates": {
      "character_cost_multiplier": 1
    },
    "languages": [
      {
        "language_id": "en",
        "name": "English"
      }
    ]
  }
]

200 MiniMax models (provider=minimax)

[
  {
    "model_id": "speech-2.8-hd",
    "name": "MiniMax Speech 2.8 HD",
    "description": "",
    "credit_ratio": 1,
    "max_chars": 5000
  }
]

Languages

Get languages

GET/v1/languages

Languages accepted for text to speech. ElevenLabs and CapCut use ISO codes (e.g. vi); MiniMax uses full names (e.g. Vietnamese). Either form is accepted in requests.

Query parameters

NameTypeDescription
providerstringelevenlabs (default), minimax or capcut.
curl -X GET "https://api.vocallab.net/v1/languages" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 ElevenLabs languages (default)

[
  {
    "code": "en",
    "name": "English"
  },
  {
    "code": "vi",
    "name": "Vietnamese"
  }
]

200 MiniMax languages (provider=minimax)

[
  {
    "code": "English",
    "name": "English"
  },
  {
    "code": "Vietnamese",
    "name": "Vietnamese"
  }
]

Elevenlabs Voices

List default voices

GET/v1/default-voices

ElevenLabs premade voices. preview_url is a short-lived signed link served by VocalLab.

Query parameters

NameTypeDescription
page_sizeintegerVoices per page. Default: 30.
searchstringFilter by voice name.
curl -X GET "https://api.vocallab.net/v1/default-voices?page_size=30" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "voices": [
    {
      "voice_id": "rv30Fd6w5bnbL0kHzWlr",
      "name": "Aria",
      "description": "A young, expressive female voice.",
      "preview_url": "https://api.vocallab.net/audio/voices/elevenlabs/rv30Fd6w5bnbL0kHzWlr.mp3?exp=1767225600&sig=…",
      "category": "premade",
      "gender": "female",
      "age": "young",
      "accent": "american",
      "language": "en",
      "use_case": "social_media",
      "descriptive": "expressive",
      "free_users_allowed": true,
      "verified_languages": [
        {
          "language_id": "en",
          "language": "English",
          "model_id": "eleven_multilingual_v2",
          "accent": "American",
          "locale": "en-US"
        }
      ]
    }
  ],
  "has_more": true,
  "last_sort_id": "next_page_token_value"
}

List shared voices

GET/v1/shared-voices

The ElevenLabs shared voice library with filters and sorting.

Query parameters

NameTypeDescription
page_sizeintegerVoices per page, max 100. Default: 30.
pageintegerZero-based page number. Default: 0.
searchstringSearch name or description.
sortstringtrending, created_date, cloned_by_count or usage_character_count_1y. Default: "trending".
categorystringhigh_quality, or omit for all.
genderstringmale, female or neutral.
agestringyoung, middle_aged or old.
accentstringE.g. american, british.
required_languagesstringLanguage code, e.g. en.
use_cases[]arrayRepeatable. conversational, narrative_and_story, characters_and_animation, social_media, entertainment_and_tv, advertisement, informative_and_educational.
curl -X GET "https://api.vocallab.net/v1/shared-voices?page_size=30&page=0" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "voices": [
    {
      "voice_id": "pNInz6obpgDQGcFmaJgB",
      "name": "Adam",
      "description": "A warm, friendly male voice.",
      "preview_url": "https://api.vocallab.net/audio/voices/elevenlabs/pNInz6obpgDQGcFmaJgB.mp3?exp=1767225600&sig=…",
      "category": "professional",
      "gender": "male",
      "age": "middle_aged",
      "accent": "american",
      "language": "en",
      "use_case": "narration",
      "descriptive": "warm",
      "free_users_allowed": true,
      "usage_character_count_1y": 1500000,
      "usage_character_count_7d": 25000,
      "cloned_by_count": 1200,
      "featured": false,
      "rate": 0,
      "live_moderation_enabled": false,
      "date_unix": 1700000000,
      "verified_languages": [
        {
          "language_id": "en",
          "language": "English",
          "model_id": "eleven_multilingual_v2",
          "accent": "American",
          "locale": "en-US"
        }
      ],
      "image_url": "https://storage.googleapis.com/images/adam.jpg"
    }
  ],
  "has_more": true,
  "last_sort_id": "abc123def456"
}

MiniMax Voices

List system voices

GET/v1/minimax/system-voices

The MiniMax system voice library. Use voice_id or uniq_id with Text to speech (provider minimax).

Query parameters

NameTypeDescription
pageintegerOne-based page number. Default: 1.
page_sizeintegerVoices per page. Default: 30.
searchstringFilter by name.
genderstringE.g. Male, Female.
languagestringE.g. English, Chinese (Mandarin).
accentstringE.g. EN-British.
agestringComma-separated, e.g. Young,Middle-aged.
use_casesstringComma-separated, e.g. Audiobook,Podcast.
curl -X GET "https://api.vocallab.net/v1/minimax/system-voices?page=1&page_size=30" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "voice_list": [
    {
      "voice_id": "226905123659939",
      "voice_name": "Aussie Bloke",
      "uniq_id": "English_Aussie_Bloke",
      "tag_list": [
        "English",
        "Male",
        "Young",
        "EN-Australian"
      ],
      "cover_url": "https://cdn.hailuoai.video/moss/voice_cover/xxx.webp",
      "sample_audio": "https://api.vocallab.net/audio/voices/minimax/226905123659939.mp3?exp=1767225600&sig=…",
      "description": "A warm Australian male voice."
    }
  ],
  "total": 150,
  "has_more": true
}

List cloned voices

GET/v1/minimax/voices

Your cloned voices with status processing, done or failed. Use id as the voice_id for MiniMax text to speech once the status is done.

curl -X GET "https://api.vocallab.net/v1/minimax/voices" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "voices": [
    {
      "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
      "voice_name": "My Custom Voice",
      "sample_audio_url": "https://api.vocallab.net/audio/5d4c3b2a-1f0e-4d9c-8b7a-6f5e4d3c2b1a.mp3?exp=1767225600&sig=…",
      "cover_url": null,
      "language_tag": "English",
      "gender": "Male",
      "status": "done",
      "created_at": "2026-10-11T10:30:00.000Z"
    },
    {
      "id": "b2c3d4e5-f6a7-8901-bcde-f12345678901",
      "voice_name": "Processing Voice",
      "language_tag": "English",
      "gender": "Female",
      "status": "processing",
      "created_at": "2026-10-11T10:30:00.000Z"
    }
  ]
}

Clone voice

POST/v1/minimax/voices/clone

Clone a voice from an audio sample (10 s – 5 min, max 20 MB). Processing is asynchronous: poll List cloned voices until the status is done. Credits are reserved when the request is accepted and refunded if cloning fails.

Request body multipart/form-data

NameTypeDescription
file requiredfileAudio sample, max 20 MB.
voice_name requiredstringName for the cloned voice.
language_tagstringLanguage of the sample, e.g. English, Vietnamese. Default: "English".
genderstringMale or Female. Default: "Male".
need_noise_reductionbooleanApply noise reduction to the sample. Default: true.
preview_textstringText for the preview sample, max 200 characters. Default: "Hello".
curl -X POST "https://api.vocallab.net/v1/minimax/voices/clone" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "file=@sample.mp3" \
  -F "voice_name=My Voice" \
  -F "language_tag=English" \
  -F "gender=Male"

Responses

200 Clone accepted (processing)

{
  "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890",
  "voice_name": "My Voice",
  "language_tag": "English",
  "gender": "Male",
  "status": "processing",
  "created_at": "2026-10-11T10:30:00.000Z"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

Delete cloned voice

DELETE/v1/minimax/voices/{id}

Delete a cloned voice. It no longer appears in the list and can no longer be used.

Path parameters

NameTypeDescription
id requiredstringCloned voice ID.
curl -X DELETE "https://api.vocallab.net/v1/minimax/voices/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 Deleted

{
  "success": true,
  "id": "a1b2c3d4-e5f6-7890-abcd-ef1234567890"
}

404 The task does not exist or belongs to another account.

{
  "error": "voice not found"
}

Retry clone

POST/v1/minimax/voices/{id}/retry

Re-run a failed or completed clone with the original sample and settings. Credits are charged again. The response carries the ID of the new clone.

Path parameters

NameTypeDescription
id requiredstringCloned voice ID.
curl -X POST "https://api.vocallab.net/v1/minimax/voices/{id}/retry" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 Clone re-initiated (processing)

{
  "id": "c3d4e5f6-a7b8-9012-cdef-123456789012",
  "voice_name": "My Voice",
  "language_tag": "English",
  "gender": "Male",
  "status": "processing",
  "created_at": "2026-10-11T10:30:00.000Z"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

404 The task does not exist or belongs to another account.

{
  "error": "voice not found"
}

409 Only completed or failed tasks can be retried.

{
  "error": "only completed or failed tasks can be retried"
}

422 The stored source file is no longer available. Nothing is charged.

{
  "error": "cannot retry: source file no longer available"
}

CapCut Voices

List system voices

GET/v1/capcut/system-voices

The CapCut voice library. Each voice_id can be used with Text to speech (provider capcut).

Query parameters

NameTypeDescription
pageintegerOne-based page number. Default: 1.
page_sizeintegerVoices per page. Default: 30.
searchstringCase-insensitive match on the name.
languagestringen, vi, zh, id, es, pt, ja or th.
genderstringMale, Female or Gender Diversity.
agestringE.g. Elderly, Teenager.
emotionstringE.g. Excited, Plain.
accentstringE.g. Gentle / Kind, Deep.
curl -X GET "https://api.vocallab.net/v1/capcut/system-voices?page=1&page_size=30" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "voice_list": [
    {
      "voice_id": "7036989371643923969",
      "name": "Energetic Narrator",
      "language": "en",
      "gender": "Male",
      "platform": "sami",
      "image_url": null,
      "tags": [
        "male",
        "steady",
        "narration"
      ],
      "preview_url": "https://api.vocallab.net/audio/voices/capcut/7036989371643923969.mp3?exp=1767225600&sig=…"
    }
  ],
  "total": 480,
  "has_more": true
}

Preview voice

GET/v1/capcut/voices/{voice_id}/preview

Stream the audio sample of a CapCut voice.

Path parameters

NameTypeDescription
voice_id requiredstringCapCut voice ID.
curl -X GET "https://api.vocallab.net/v1/capcut/voices/{voice_id}/preview" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 Audio sample (audio/mpeg)

404 Unknown voice.

{
  "error": "voice not found"
}

Text to speech

Text to speech

POST/v1/text-to-speech/{voice_id}

Convert text to speech with ElevenLabs, MiniMax or CapCut. The task runs asynchronously: poll Get history detail with the returned id until status is completed or failed. Credits are reserved upfront and the unused or failed part is refunded.

Path parameters

NameTypeDescription
voice_id requiredstringElevenLabs voice ID; MiniMax voice_id, uniq_id or cloned voice ID; or CapCut voice ID.

Request body application/json

NameTypeDescription
text requiredstringText to convert.
model_idstringModel from Get models. For CapCut use capcut. Default: "eleven_multilingual_v2".
providerstringelevenlabs, minimax or capcut. Default: "elevenlabs".
language_code requiredstringISO code (ElevenLabs, CapCut) or language name (MiniMax). See Get languages.
voice_settingsobjectVoice settings. Accepted fields depend on the provider.
export_transcriptbooleanAlso generate an SRT transcript. Adds about 15% to the credit cost. Default: false.
output_formatstringVocalLab extension: source, mp3_44100_128, mp3_44100_192 or wav_8000…wav_48000. Default: "source".
curl -X POST "https://api.vocallab.net/v1/text-to-speech/{voice_id}" \
  -H "xi-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Hello world",
    "model_id": "eleven_multilingual_v2",
    "language_code": "en",
    "voice_settings": {
      "stability": 0.5,
      "similarity_boost": 0.75,
      "speed": 1
    }
  }'

Responses

202 Accepted — the task is queued.

{
  "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

400 Invalid request, e.g. a missing language code.

{
  "error": "missing_language_code"
}

Get history list

GET/v1/history

Text to speech and dialogue tasks, newest first.

Query parameters

NameTypeDescription
page_sizeintegerTasks per page, 1–100. Default: 30.
pageintegerZero-based page number. Default: 0.
curl -X GET "https://api.vocallab.net/v1/history?page_size=30&page=0" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "tasks": [
    {
      "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
      "user_id": "user_123",
      "status": "completed",
      "progress": 100,
      "provider": "elevenlabs",
      "text": "Hello world!",
      "voice_id": "UgBBYS2sOqTuMpoF3BR0",
      "model_id": "eleven_multilingual_v2",
      "metadata": {
        "language_code": "en",
        "voice_settings": {
          "stability": 0.5,
          "similarity_boost": 0.75,
          "style": 0,
          "use_speaker_boost": true,
          "speed": 1
        }
      },
      "result": {
        "audio_url": "https://api.vocallab.net/audio/9b19d3e2-7c68-4eca-be48-e9917682d73e.mp3?exp=1767225600&sig=…"
      },
      "error": null,
      "characters_used": 12,
      "credits_deducted": 12,
      "created_at": "2026-10-11T07:42:13.641Z",
      "updated_at": "2026-10-11T07:42:16.432Z"
    },
    {
      "id": "id_of_another_task",
      "user_id": "user_123",
      "status": "completed",
      "progress": 100,
      "provider": "minimax",
      "text": "Hello from MiniMax!",
      "voice_id": "English_Aussie_Bloke",
      "model_id": "speech-2.8-turbo",
      "metadata": {
        "language_code": "English",
        "voice_settings": {
          "speed": 1,
          "pitch": 0,
          "vol": 1
        }
      },
      "result": {
        "audio_url": "https://api.vocallab.net/audio/9b19d3e2-7c68-4eca-be48-e9917682d73e.mp3?exp=1767225600&sig=…"
      },
      "error": null,
      "characters_used": 19,
      "credits_deducted": 19,
      "created_at": "2026-10-11T07:42:13.641Z",
      "updated_at": "2026-10-11T07:42:16.432Z"
    }
  ],
  "has_more": false
}

Get history detail

GET/v1/history/{id}

One text to speech or dialogue task, for polling. status is pending, processing, completed or failed. result.audio_url is a signed link valid for 24 hours; read the task again for a fresh link.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X GET "https://api.vocallab.net/v1/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "user_id": "user_123",
  "status": "processing",
  "progress": 60,
  "provider": "elevenlabs",
  "text": "Hello world!",
  "voice_id": "UgBBYS2sOqTuMpoF3BR0",
  "model_id": "eleven_multilingual_v2",
  "metadata": {
    "language_code": "en",
    "voice_settings": {
      "stability": 0.5,
      "similarity_boost": 0.75,
      "style": 0,
      "use_speaker_boost": true,
      "speed": 1
    }
  },
  "result": {
    "audio_url": null
  },
  "error": null,
  "characters_used": 12,
  "credits_deducted": 12,
  "created_at": "2026-10-11T07:42:13.641Z",
  "updated_at": "2026-10-11T07:42:16.432Z"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

Delete history

DELETE/v1/history/{id}

Delete a finished text to speech or dialogue task and its files.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X DELETE "https://api.vocallab.net/v1/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "message": "deleted"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 The task is still pending or processing.

{
  "error": "task_still_running",
  "message": "This task is still running. Wait for it to finish before deleting it."
}

Retry task

POST/v1/history/{id}/retry

Re-run a completed or failed text to speech, dialogue or SRT task with its original text and settings. Credits are charged again. Poll the new task id that is returned.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X POST "https://api.vocallab.net/v1/history/{id}/retry" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

202 Accepted — a new task is queued.

{
  "id": "b7c1d2e3-f4a5-4b6c-8d7e-9f0a1b2c3d4e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 Only completed or failed tasks can be retried.

{
  "error": "only completed or failed tasks can be retried"
}

Dialogue

Create dialogue

POST/v1/dialogue

Convert a multi-speaker script to speech. Each line starting with Name: begins a new turn; following lines without a prefix continue that turn. Every speaker has its own provider, voice, model and settings. Poll Get history detail with the returned id.

Request body application/json

NameTypeDescription
text requiredstringScript in Speaker: text lines, with at least two different speakers.
speakers requiredobjectMap of speaker name to {provider, voice_id, model_id, language_code?, voice_settings?}. Providers: elevenlabs or minimax.
pause_between_turnsnumberSilence between turns in seconds, 0–3. Default: 0.
curl -X POST "https://api.vocallab.net/v1/dialogue" \
  -H "xi-api-key: YOUR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "text": "Alice: Hello, how are you?\nBob: I am doing great, thanks for asking!",
    "speakers": {
      "Alice": {
        "provider": "elevenlabs",
        "voice_id": "rv30Fd6w5bnbL0kHzWlr",
        "model_id": "eleven_multilingual_v2",
        "language_code": "en"
      },
      "Bob": {
        "provider": "minimax",
        "voice_id": "English_Aussie_Bloke",
        "model_id": "speech-2.8-hd",
        "language_code": "English"
      }
    },
    "pause_between_turns": 0.3
  }'

Responses

202 Accepted — the dialogue task is queued.

{
  "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

Speech to text

Speech to text

POST/v1/speech-to-text

Transcribe audio or video with ElevenLabs Scribe v2. Asynchronous: returns a task_id; poll Get transcription every couple of seconds. Credits are charged from the audio duration and refunded in full if the job fails. Requests without a valid language_code are rejected before any charge.

Request body multipart/form-data

NameTypeDescription
file requiredfileAudio or video, max 90 MB (AAC, AIFF, OGG, MP3, Opus, WAV, WebM, FLAC, M4A, MP4).
language_code requiredstringISO 639-1 or 639-3 code, e.g. en, vie.
timestamps_granularitystringnone, word or character. Default: "word".
diarizebooleanLabel each word with a speaker_id. Default: false.
num_speakersintegerMaximum number of speakers, 1–32.
diarization_thresholdnumber0.0–1.5. Used when diarize=true and num_speakers is not set.
tag_audio_eventsbooleanTag non-speech events such as (laughter). Default: true.
no_verbatimbooleanRemove filler words and false starts. Default: false.
temperaturenumber0.0–2.0.
seedinteger0–2147483647, for repeatable output.
file_formatstringother or pcm_s16le_16. Default: "other".
use_multi_channelbooleanTranscribe up to 5 channels separately. Not with diarize. Default: false.
keytermsarrayRepeatable. Terms to bias towards, max 1000, each under 50 characters.
entity_detectionarrayRepeatable. all, pii, phi, pci, other, offensive_language.
entity_redactionarrayRepeatable. Subset of entity_detection to redact.
entity_redaction_modestringredacted, entity_type or enumerated_entity_type. Default: "enumerated_entity_type".
additional_formatsarrayRepeatable. srt, txt, docx, html, pdf, segmented_json.
curl -X POST "https://api.vocallab.net/v1/speech-to-text" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "file=@recording.mp3" \
  -F "language_code=en" \
  -F "diarize=true"

Responses

202 Accepted — poll Get transcription.

{
  "task_id": "f534211b-b414-4697-bb58-a2a7a688a49f",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

400 Missing or unrecognised language code. Nothing is charged.

{
  "error": "missing_language_code"
}

Get transcription

GET/v1/speech-to-text/{task_id}

Poll a transcription until status is completed or failed. result appears once completed; on failure detail_error explains why and the credits have been refunded.

Path parameters

NameTypeDescription
task_id requiredstringThe task_id returned on submission.
curl -X GET "https://api.vocallab.net/v1/speech-to-text/{task_id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "task_id": "f534211b-b414-4697-bb58-a2a7a688a49f",
  "status": "completed",
  "progress": 100,
  "filename": "recording.mp3",
  "audio_duration_secs": 56.15,
  "credits_deducted": 17,
  "created_at": "2026-10-11T13:50:48.000Z",
  "updated_at": "2026-10-11T13:50:56.000Z",
  "result": {
    "language_code": "eng",
    "language_probability": 0.97,
    "text": "Hello everyone.",
    "words": [
      {
        "text": "Hello",
        "start": 0,
        "end": 0.32,
        "type": "word",
        "logprob": -0.002,
        "speaker_id": "speaker_0"
      },
      {
        "text": " ",
        "start": 0.32,
        "end": 0.4,
        "type": "spacing",
        "logprob": -0.001,
        "speaker_id": "speaker_0"
      },
      {
        "text": "everyone.",
        "start": 0.4,
        "end": 0.92,
        "type": "word",
        "logprob": -0.04,
        "speaker_id": "speaker_0"
      }
    ],
    "additional_formats": [],
    "transcription_id": "Qe71uRUkuArC2JNlRsGy",
    "audio_duration_secs": 56.15
  }
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

List transcriptions

GET/v1/speech-to-text

Your transcriptions, newest first. Transcripts are not included; read a single task to get its result.

Query parameters

NameTypeDescription
pageintegerZero-based page number. Default: 0.
page_sizeintegerItems per page, max 100. Default: 30.
curl -X GET "https://api.vocallab.net/v1/speech-to-text?page=0&page_size=30" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "tasks": [
    {
      "task_id": "f534211b-b414-4697-bb58-a2a7a688a49f",
      "status": "completed",
      "progress": 100,
      "filename": "recording.mp3",
      "audio_duration_secs": 56.15,
      "credits_deducted": 17,
      "created_at": "2026-10-11T13:50:48.000Z"
    }
  ],
  "has_more": false
}

Delete transcription

DELETE/v1/speech-to-text/{task_id}

Permanently delete a finished transcription and its stored audio. Deleting does not refund credits.

Path parameters

NameTypeDescription
task_id requiredstringThe transcription to delete.
curl -X DELETE "https://api.vocallab.net/v1/speech-to-text/{task_id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 Deleted

{
  "deleted": true
}

409 The task is still pending or processing.

{
  "error": "task_still_running",
  "message": "This transcription is still running. Wait for it to finish before deleting it."
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

Voice Changer

Voice changer

POST/v1/voice-changer

Speech to speech: re-voice an audio file with an ElevenLabs voice (max 50 MB, 5 minutes). Asynchronous; poll Get history detail. Credits are based on the audio duration.

Request body multipart/form-data

NameTypeDescription
audio requiredfileSource audio, max 50 MB and 5 minutes.
voice_id requiredstringTarget ElevenLabs voice ID.
model_idstringElevenLabs speech-to-speech model. Default: "eleven_multilingual_sts_v2".
remove_background_noisestringtrue to remove background noise. Default: "false".
voice_settingsstringJSON string with stability, similarity_boost, style, use_speaker_boost.
curl -X POST "https://api.vocallab.net/v1/voice-changer" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "audio=@speech.mp3" \
  -F "voice_id=rv30Fd6w5bnbL0kHzWlr" \
  -F "voice_settings={\"stability\":0.5,\"similarity_boost\":0.75}"

Responses

202 Accepted

{
  "task_id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

Get history list

GET/v1/voice-changer/history

Voice changer tasks, newest first.

Query parameters

NameTypeDescription
page_sizeintegerTasks per page, 1–100. Default: 30.
pageintegerZero-based page number. Default: 0.
curl -X GET "https://api.vocallab.net/v1/voice-changer/history?page_size=30&page=0" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "tasks": [
    {
      "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
      "user_id": "user_123",
      "status": "completed",
      "progress": 100,
      "voice_id": "rv30Fd6w5bnbL0kHzWlr",
      "model_id": "eleven_multilingual_sts_v2",
      "source_audio_url": "https://api.vocallab.net/audio/5f0c2a4e-1d1b-4c8e-9a51-2a3b4c5d6e7f.mp3?exp=1767225600&sig=…",
      "result": {
        "audio_url": "https://api.vocallab.net/audio/9b19d3e2-7c68-4eca-be48-e9917682d73e.mp3?exp=1767225600&sig=…"
      },
      "metadata": {
        "voice_settings": {
          "stability": 0.5,
          "similarity_boost": 0.75
        },
        "remove_noise": false
      },
      "credits_deducted": 500,
      "error": null,
      "detail_error": null,
      "created_at": "2026-10-11T10:30:00.000Z",
      "updated_at": "2026-10-11T10:30:15.000Z"
    }
  ],
  "has_more": false
}

Get history detail

GET/v1/voice-changer/history/{id}

One voice changer task, for polling.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X GET "https://api.vocallab.net/v1/voice-changer/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "user_id": "user_123",
  "status": "completed",
  "progress": 100,
  "voice_id": "rv30Fd6w5bnbL0kHzWlr",
  "model_id": "eleven_multilingual_sts_v2",
  "source_audio_url": "https://api.vocallab.net/audio/5f0c2a4e-1d1b-4c8e-9a51-2a3b4c5d6e7f.mp3?exp=1767225600&sig=…",
  "result": {
    "audio_url": "https://api.vocallab.net/audio/9b19d3e2-7c68-4eca-be48-e9917682d73e.mp3?exp=1767225600&sig=…"
  },
  "metadata": {
    "voice_settings": {
      "stability": 0.5,
      "similarity_boost": 0.75
    },
    "remove_noise": false
  },
  "credits_deducted": 500,
  "error": null,
  "detail_error": null,
  "created_at": "2026-10-11T10:30:00.000Z",
  "updated_at": "2026-10-11T10:30:15.000Z"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

Retry task

POST/v1/voice-changer/history/{id}/retry

Re-run a completed or failed voice changer task with its original audio. Credits are charged again from the audio duration.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X POST "https://api.vocallab.net/v1/voice-changer/history/{id}/retry" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

202 Accepted — a new task is queued.

{
  "task_id": "b7c1d2e3-f4a5-4b6c-8d7e-9f0a1b2c3d4e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 Only completed or failed tasks can be retried.

{
  "error": "only completed or failed tasks can be retried"
}

422 The stored source file is no longer available. Nothing is charged.

{
  "error": "cannot retry: source file no longer available"
}

Delete task

DELETE/v1/voice-changer/history/{id}

Delete a finished voice changer task with its source and result audio.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X DELETE "https://api.vocallab.net/v1/voice-changer/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "message": "deleted"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 The task is still pending or processing.

{
  "error": "task_still_running",
  "message": "This task is still running. Wait for it to finish before deleting it."
}

Dubbing

1. Init upload

POST/v1/aidubbing/upload/init

Step 1 of an AI dubbing job: open an upload session. Sources up to 2 GB and 2.5 hours (MP3, WAV, AAC, M4A, MP4, MOV). Flow: init → chunk (×N) → complete.

curl -X POST "https://api.vocallab.net/v1/aidubbing/upload/init" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 Upload session created

{
  "upload_id": "0f8fad5b-d9cb-469f-a165-70867728950e",
  "chunk_size": 52428800
}

2. Upload chunks

POST/v1/aidubbing/upload/chunk

Step 2: upload each part of the file at its zero-based index, slicing at chunk_size. Parts may arrive in any order or in parallel, and re-posting an index replaces it. Max 100 MB per part and 2 GB per session.

Request body multipart/form-data

NameTypeDescription
upload_id requiredstringSession ID from init.
index requiredintegerZero-based part index, max 511.
chunk requiredfileThe file slice for this index.
curl -X POST "https://api.vocallab.net/v1/aidubbing/upload/chunk" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "upload_id=0f8fad5b-d9cb-469f-a165-70867728950e" \
  -F "index=0" \
  -F "chunk=@part0.bin"

Responses

200 Part stored

{
  "index": 0,
  "received": true
}

400 Invalid index, missing data, part too large or session over 2 GB.

{
  "error": "chunk too large"
}

404 Unknown upload session.

{
  "error": "upload session not found"
}

3. Complete & start dubbing

POST/v1/aidubbing/upload/complete

Final step: reassemble parts 0..total_chunks-1 and start the dubbing job. Credits are reserved upfront from the source duration. Poll Get history detail with the returned task_id. Once assembly starts the session is consumed, whether or not the job is accepted.

Request body multipart/form-data

NameTypeDescription
upload_id requiredstringSession ID from init.
total_chunks requiredintegerNumber of parts uploaded, 1–512.
filenamestringOriginal file name, shown in history. Default: "upload".
content_typestringMIME type of the source, e.g. video/mp4.
target_language requiredstringISO 639-1 code to dub into: af, ar, hy, as, az, be, bn, bs, bg, ca, ceb, ny, zh, hr, cs, da, nl, en, et, fil, fi, fr, gl, ka, de, el, gu, ha, he, hi, hu, is, id, ga, it, ja, jv, kn, kk, ky, ko, lv, ln, lt, lb, mk, ms, ml, mr, ne, no, ps, fa, pl, pt, pa, ro, ru, sr, sd, sk, sl, so, es, sw, sv, ta, te, th, tr, uk, ur, vi, cy, nn, yue. CapCut supports en, zh, es, id, pt, fr, de, vi, ja, th, tr, it.
source_language requiredstringSource language code (same list plus mi, tl). auto is not accepted.
providerstringminimax (speech-2.8-hd), elevenlabs (eleven_v3) or capcut. Default: "minimax".
voice_id requiredstringMiniMax uniq_id or cloned voice ID, ElevenLabs voice ID, or CapCut voice ID matching the target language.
curl -X POST "https://api.vocallab.net/v1/aidubbing/upload/complete" \
  -H "xi-api-key: YOUR_API_KEY" \
  -F "upload_id=0f8fad5b-d9cb-469f-a165-70867728950e" \
  -F "total_chunks=3" \
  -F "filename=interview.mp4" \
  -F "content_type=video/mp4" \
  -F "source_language=en" \
  -F "target_language=vi" \
  -F "provider=minimax" \
  -F "voice_id=Wise_Woman"

Responses

202 Accepted — the dubbing task is queued.

{
  "task_id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "status": "pending"
}

400 Missing part or invalid parameters.

{
  "error": "missing chunk 2"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

404 Unknown upload session.

{
  "error": "upload session not found"
}

Get history list

GET/v1/aidubbing/history

AI dubbing tasks, newest first.

Query parameters

NameTypeDescription
page_sizeintegerTasks per page, 1–100. Default: 30.
pageintegerZero-based page number. Default: 0.
curl -X GET "https://api.vocallab.net/v1/aidubbing/history?page_size=30&page=0" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "tasks": [
    {
      "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
      "user_id": "user_123",
      "status": "completed",
      "progress": 100,
      "source_lang": "en",
      "target_lang": "vi",
      "provider": "minimax",
      "voice_id": "Wise_Woman",
      "model_id": "speech-2.8-hd",
      "source_audio_url": "https://api.vocallab.net/audio/0d7f1c9e-6a1b-4f3e-8c2d-1e2f3a4b5c6d.mp3?exp=1767225600&sig=…",
      "result": {
        "audio_url": "https://api.vocallab.net/audio/7a1e9c2b-3d4f-4a5b-8c6d-9e0f1a2b3c4d.mp4?exp=1767225600&sig=…",
        "srt_url": "https://api.vocallab.net/audio/3c2b1a0f-9e8d-4c7b-a6f5-e4d3c2b1a0f9.srt?exp=1767225600&sig=…"
      },
      "metadata": {
        "source_file_name": "interview.mp4",
        "original_video_url": null,
        "voice_name": "Wise Woman"
      },
      "credits_deducted": 18000,
      "error": null,
      "detail_error": null,
      "created_at": "2026-10-11T10:30:00.000Z",
      "updated_at": "2026-10-11T10:35:15.000Z"
    }
  ],
  "has_more": false
}

Get history detail

GET/v1/aidubbing/history/{id}

One dubbing task, for polling. result.audio_url is the dubbed video (MP4) for video sources, otherwise audio; result.srt_url holds translated subtitles.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X GET "https://api.vocallab.net/v1/aidubbing/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK

{
  "id": "9b19d3e2-7c68-4eca-be48-e9917682d73e",
  "user_id": "user_123",
  "status": "completed",
  "progress": 100,
  "source_lang": "en",
  "target_lang": "vi",
  "provider": "minimax",
  "voice_id": "Wise_Woman",
  "model_id": "speech-2.8-hd",
  "source_audio_url": "https://api.vocallab.net/audio/0d7f1c9e-6a1b-4f3e-8c2d-1e2f3a4b5c6d.mp3?exp=1767225600&sig=…",
  "result": {
    "audio_url": "https://api.vocallab.net/audio/7a1e9c2b-3d4f-4a5b-8c6d-9e0f1a2b3c4d.mp4?exp=1767225600&sig=…",
    "srt_url": "https://api.vocallab.net/audio/3c2b1a0f-9e8d-4c7b-a6f5-e4d3c2b1a0f9.srt?exp=1767225600&sig=…"
  },
  "metadata": {
    "source_file_name": "interview.mp4",
    "original_video_url": null,
    "voice_name": "Wise Woman"
  },
  "credits_deducted": 18000,
  "error": null,
  "detail_error": null,
  "created_at": "2026-10-11T10:30:00.000Z",
  "updated_at": "2026-10-11T10:35:15.000Z"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

Retry task

POST/v1/aidubbing/history/{id}/retry

Re-run a completed or failed dubbing task from its stored source, without uploading again. The original credit cost is charged again.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X POST "https://api.vocallab.net/v1/aidubbing/history/{id}/retry" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

202 Accepted — a new task is queued.

{
  "task_id": "b7c1d2e3-f4a5-4b6c-8d7e-9f0a1b2c3d4e",
  "status": "pending"
}

402 Not enough credits. required is the cost of the request.

{
  "error": "insufficient credits",
  "required": 18000
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 Only completed or failed tasks can be retried.

{
  "error": "only completed or failed tasks can be retried"
}

422 The stored source file is no longer available. Nothing is charged.

{
  "error": "cannot retry: source file no longer available"
}

Delete task

DELETE/v1/aidubbing/history/{id}

Delete a finished dubbing task and its stored files.

Path parameters

NameTypeDescription
id requiredstringTask ID returned when the task was created.
curl -X DELETE "https://api.vocallab.net/v1/aidubbing/history/{id}" \
  -H "xi-api-key: YOUR_API_KEY"

Responses

200 OK - task deleted

{
  "status": "deleted"
}

404 The task does not exist or belongs to another account.

{
  "error": "task not found"
}

409 The task is still pending or processing.

{
  "error": "task_still_running",
  "message": "This task is still running. Wait for it to finish before deleting it."
}

VocalLab extensions

VocalLab-only endpoints: explicit quotes, one task API for every tool (including batch scenes and srt voice-over), resumable uploads and authenticated downloads. See the OpenAPI document for their schemas.

  • POST/v1/quotes — Quote one task before submitting it
  • POST/v1/tasks — Submit any configured tool using kind, native payload and private upload_id
  • GET/v1/tasks — Paginated private task history
  • GET/v1/tasks/{id}/retry-input — Read canonical studio input for an owned settled task
  • POST/v1/history/{id}/cancel — Cancel an owned task
  • POST/v1/uploads — Upload audio/video into private VocalLab storage
  • POST/v1/uploads/init — Begin resumable dubbing upload
  • POST/v1/uploads/{id}/chunks/{index} — Upload one binary chunk with replay validation
  • POST/v1/uploads/{id}/complete — Assemble and probe uploaded media
  • GET/v1/voices — List public or user-owned library voices
  • GET/v1/audio/{id} — Download owned result with VocalLab API key
  • GET/v1/artifacts/{id} — Download owned media or transcript artifact