# Sawtak Arabi documentation > Complete public English documentation for agents integrating the Sawtak Arabi API. Endpoint guides below describe the supported API. No public OpenAPI schema is currently served. Base URL: `https://api.sawtakarabi.ai/v1` Authentication: `Authorization: Bearer ` ## Sawtak Arabi documentation > Build Arabic speech experiences with the Sawtak Arabi API. Canonical page: https://sawtakarabi.ai/docs/index.md Sawtak Arabi provides an OpenAI-compatible API for Arabic text-to-speech and reusable voices. ## Generate with your AI agent Use the [official Sawtak Arabi skill](/docs/agent-skill) with Claude Code or Codex. Your agent can install it, guide you through creating an API key, find a voice, and generate the audio. > Use Sawtak Arabi to generate a Saudi Najdi Arabic voiceover. Follow https://sawtakarabi.ai/docs/agent-skill and reply in English. Agents can start from [llms.txt](/llms.txt), read [all guides](/llms-full.txt), or open the [skill instructions](https://raw.githubusercontent.com/SawtakArabi/skill/main/skills/arabic-voiceover/SKILL.md). ## Start here - [Quickstart](/docs/quickstart) — send a first synthesis request. - [Authentication](/docs/authentication) — create and use an API key. - [Text to speech](/docs/text-to-speech) — generate streamed audio. ## Build with confidence Use the [voice catalog](/docs/voices) to choose a voice, or [clone a voice](/docs/voice-cloning) for a reusable private voice. The [Errors](/docs/errors) and [Request limits](/docs/limits) guides explain the response envelope and safe retry behaviour. ## Public API endpoints Use `https://api.sawtakarabi.ai/v1` for integrations. Website `/api/*` routes are not the public API. No public OpenAPI schema is currently served. | Endpoint | Purpose | | --- | --- | | `POST /audio/speech` | [Generate streamed speech](/docs/text-to-speech). | | `GET /voices` and `GET /voices/{id}` | [Find voices and inspect metadata](/docs/voices). | | `GET /voices/{id}/preview` | Audition an existing sample; public samples need no key. | | `POST /voices`, `PATCH /voices/{id}`, `DELETE /voices/{id}` | [Create and manage owned voices](/docs/voice-cloning). | | `GET /operations/{id}`, `POST /operations/{id}/cancel` | [Inspect status/billing or cancel work](/docs/operations). | | `GET /me`, `GET /usage` | [Account details and daily usage](/docs/usage). | | `GET /models`, `GET /models/arabic-tts-1` | Authenticated OpenAI-compatible model discovery. The served model is `arabic-tts-1`. | All listed routes require authentication except public voice previews. `/v1/auth/tokens` is reserved for the platform's service credential; customer integrations use dashboard API keys directly. ## Use with an AI agent > Generate Arabic voiceovers with Claude Code, Codex, and the official Sawtak Arabi skill. Canonical page: https://sawtakarabi.ai/docs/agent-skill.md Ask your agent to create Arabic audio with Sawtak Arabi. The official **arabic-voiceover** skill selects voices, generates a WAV, and saves it in your workspace. It is free to install; generation uses your Sawtak Arabi balance. > Use Sawtak Arabi to generate a 20-second Saudi Najdi Arabic ad for my coffee shop. Read https://sawtakarabi.ai/docs/agent-skill and use the official skill. Reply in English. ## 1. Install the official skill In Claude Code or Codex, ask your agent to install [SawtakArabi/skill](https://github.com/SawtakArabi/skill), or run: ```bash npx skills add SawtakArabi/skill --skill arabic-voiceover -g -a claude-code -a codex ``` The installer requires Node.js/npm. Generation requires Python 3.10+ and the skill's Python dependencies; your agent can set those up. The skill uses the official OpenAI Python SDK with Sawtak Arabi's API. Agents can read the [skill instructions directly](https://raw.githubusercontent.com/SawtakArabi/skill/main/skills/arabic-voiceover/SKILL.md). The [latest release](https://github.com/SawtakArabi/skill/releases/latest) includes a downloadable ZIP with the helper and its bundled dialect list. Resolve helper paths relative to the installed `SKILL.md`; do not guess an installation directory. ## 2. Connect your API key Open [API Keys](/dashboard/api-keys), sign in or create an account, click **Create key**, enter a name such as `Claude voiceovers`, click **Create**, then **Copy** the key shown once. Configure `SAWTAK_API_KEY` in the environment used by your agent. [Authentication](/docs/authentication#configure-the-key-locally) provides exact terminal instructions. Check [Billing](/dashboard/billing) if your balance needs funding. If the key is missing, the agent should ask: > Open https://sawtakarabi.ai/dashboard/api-keys and sign in. Click Create key, give it a name, and copy the key shown once. Set it using the terminal or secret-field instructions below—not in this chat. Tell me when it’s ready and I’ll continue your voiceover. The agent should provide the appropriate Bash/Zsh command from [Authentication](/docs/authentication#configure-the-key-locally), or instructions for its supported secret field. Set the environment variable before launching a local agent; an already-running session does not inherit changes in another terminal. Do not paste the key into ordinary chat, source files, or a public repository. No permission-selection step is needed for dashboard-created keys. A browsing-only chat can read the instructions but cannot run the helper. Use Claude Code, Codex, or another environment that can execute Python and reach the API. While setup is pending, the agent can still prepare your script and find dialect names. ## 3. Generate your narration Ask naturally, or invoke `/arabic-voiceover` in Claude Code or `$arabic-voiceover` in Codex. The agent should follow this sequence: 1. Preserve your script and requested dialect, or draft the narration when asked. Reply in the language of your prompt unless you request another reply language. 2. Check setup with `doctor`. This makes no paid synthesis request. 3. Resolve dialect names with `dialects`, then query `voices` and choose a returned voice with `status: ready`. Voice discovery requires your API key; the bundled dialect list does not. 4. Write the narration to a UTF-8 file and run `generate`. Include `--enhance-pronunciation` unless you ask to disable it. 5. Return the WAV link and measured duration. Distinguish file validation from a content check or listening review. For a video request, integrate the audio into the requested video. For example, with `skill-dir` replaced by the actual installed skill directory: ```bash python3 /scripts/generate.py doctor python3 /scripts/generate.py dialects --search Najdi python3 /scripts/generate.py voices --dialect saudi-najdi --use-case advertisement --limit 5 python3 /scripts/generate.py generate \ --voice --text-file narration.txt \ --output narration.wav --enhance-pronunciation ``` If the use-case filter finds no voices, try the same dialect without that filter; do not silently switch dialects. Use `voices --details` for descriptions and preview links. `--sharing-status public` or `private` limits visibility. The bundled dialect labels are a snapshot; live catalog results determine current availability. ## If generation is interrupted The helper saves IDs and settings in `narration.json` and keeps received audio as `narration.partial.wav`. A partial WAV is incomplete narration even when it plays. Inspect the saved `operation_id` with: ```bash python3 /scripts/generate.py inspect ``` Inspection reads status and billing; it does not generate or download audio. A missing operation is unknown. The agent must not automatically start another paid generation. Reuse successful files and ask before regenerating uncertain work. For direct API integrations, see [Quickstart](/docs/quickstart), [Text to speech](/docs/text-to-speech), [Voices](/docs/voices), and [Operations and retries](/docs/operations). Machine-readable guides are indexed in [llms.txt](/llms.txt) and collected in [llms-full.txt](/llms-full.txt). ## Quickstart > Make your first Arabic speech request. Canonical page: https://sawtakarabi.ai/docs/quickstart.md Using Claude Code or Codex? Start with the [official agent skill](/docs/agent-skill) to have your agent select a voice and save a finalized WAV. Generate Arabic speech from text using the `/v1/audio/speech` endpoint. This guide walks through authentication, selecting a voice, sending a synthesis request, saving the audio, and handling supported errors. Set these variables once before running the examples: ```bash export SAWTAK_API_BASE_URL="https://api.sawtakarabi.ai/v1" export SAWTAK_API_KEY="" ``` ## 1. Create an API Key Open [API Keys](/dashboard/api-keys), sign in, click **Create key**, name it, and copy the key shown once. Dashboard-created keys need no permission selection. Follow [Authentication](/docs/authentication) to configure `SAWTAK_API_KEY`. Requests send it in the `Authorization` header. Pass the key using the `SAWTAK_API_KEY` environment variable: ```bash -H "Authorization: Bearer $SAWTAK_API_KEY" ``` ## 2. Choose a Voice Query the voice catalog using `GET $SAWTAK_API_BASE_URL/voices`: ```bash curl --fail --show-error "$SAWTAK_API_BASE_URL/voices" \ -H "Authorization: Bearer $SAWTAK_API_KEY" ``` The response returns a list envelope (`object: "list"`) containing voice objects in `data`. Each voice item includes an `id`, `name`, `status`, and `labels` (such as `dialect` and `gender`). Pass a returned `id` in the required `voice` field of your synthesis request. There is no default voice. ## 3. Issue a Speech Request Send a `POST` request to `$SAWTAK_API_BASE_URL/audio/speech` with the model, text input, and desired response format: ```bash curl --fail --show-error "$SAWTAK_API_BASE_URL/audio/speech" \ -H "Authorization: Bearer $SAWTAK_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "arabic-tts-1", "voice": "", "input": "مرحبا بك", "response_format": "wav", "enhance_pronunciation": true }' \ --output speech.wav ``` Replace `` with a returned voice whose `status` is `ready`. `enhance_pronunciation: true` enables dialect-conditioned diacritization; omit it or set it to `false` to disable it. ## 4. Play or Save the Audio The speech endpoint streams audio bytes directly into the file specified by `--output`: - **WAV (`response_format: "wav"`)**: The gateway prepends a RIFF/WAV header with streaming length markers (`audio/wav`). For a finalized file with exact lengths, use the [skill](/docs/agent-skill) or follow [Streaming audio](/docs/streaming#playing-streaming-wav-audiowav). You can play it immediately with command-line tools: ```bash afplay speech.wav ``` or on Linux: ```bash aplay speech.wav ``` - **PCM (`response_format: "pcm"`)**: The gateway streams raw 16-bit PCM audio (`audio/pcm`). Save it to `speech.pcm` when integrating with streaming decoders or pipelines that consume headerless PCM. ## 5. Supported Failure Responses Error responses follow the OpenAI-compatible error envelope containing `message`, `type`, `code`, `param`, and `request_id`. ### Unavailable Voice (`404 Not Found`) Supplying a voice ID that does not exist or is not accessible to your account returns `voice_unavailable`. Omitting `voice` returns `400 bad_voice`. ```json { "error": { "message": "Voice is unavailable", "type": "not_found_error", "code": "voice_unavailable", "param": null, "request_id": "" } } ``` ### Insufficient Scope (`403 Forbidden`) Using a restricted credential that lacks speech access returns a `permission_error`: ```json { "error": { "message": "Insufficient scope", "type": "permission_error", "code": "insufficient_scope", "param": null, "request_id": "" } } ``` ### Authentication Failure (`401 Unauthorized`) Omitting the authorization token returns `auth_missing`; an invalid key returns `auth_failed`. Both use `authentication_error`: ```json { "error": { "message": "Authentication failed", "type": "authentication_error", "code": "auth_failed", "param": null, "request_id": "" } } ``` ## Agent recipes > Deterministic workflows for Claude, Codex, and other coding agents. Canonical page: https://sawtakarabi.ai/docs/agent-recipes.md For a user asking an agent to create a voiceover, start with [Use with an AI agent](/docs/agent-skill) and the official skill instead of writing another synthesis client. These recipes are for direct API integrations. Use these recipes when an agent needs to integrate the API without guessing. Set the environment first, use the endpoint guides for supported fields, and treat the verification step as part of the task. ```bash export SAWTAK_API_BASE_URL="https://api.sawtakarabi.ai/v1" export SAWTAK_API_KEY="" ``` ## Generate a WAV file **Goal:** create a playable Arabic WAV file from text. Replace `` with a voice ID from `GET /v1/voices`. ```bash curl --fail --show-error "$SAWTAK_API_BASE_URL/audio/speech" \ -H "Authorization: Bearer $SAWTAK_API_KEY" \ -H "Content-Type: application/json" \ -d '{"model":"arabic-tts-1","voice":"","input":"مرحباً بك","response_format":"wav"}' \ --output speech.wav file speech.wav ``` **Done when:** the request exits successfully and `file speech.wav` identifies WAV audio. This is a format check, not a content or listening review; finalize streaming WAV lengths as described in [Streaming audio](/docs/streaming). Do not retry `400`, `401`, `402`, or `403` unchanged. For `429` and transient `5xx` responses, follow [Request limits](/docs/limits). ## Clone a voice **Goal:** obtain a ready private voice ID for later synthesis. 1. Send the multipart request in [Voice cloning](/docs/voice-cloning). It requires one reference file, a name, and JSON `labels` containing `dialect`. 2. Save the returned `id`. 3. Successful cloning returns HTTP 201 with `status: ready` after processing. Use `GET /v1/voices/{id}` if you need the full metadata. 4. Use that ID as `voice` in the text-to-speech request. **Done when:** the voice response reports `status: "ready"`. Never substitute a different voice when cloning is still processing or fails. ## Diagnose a failed request **Goal:** choose the next safe action from an API failure. 1. Read the HTTP status and `error.code`. 2. Save the server’s `X-Request-Id` response header as soon as headers arrive, and `error.request_id` when present. Keep any client-supplied `Idempotency-Key` separately; it is not the server request ID. 3. Correct client input for `400`, replace credentials for `401`, add balance for `402`, and use a key with the required scope for `403`. 4. For `429`, wait for `Retry-After` and use bounded exponential backoff with jitter. 5. Use [Operations and retries](/docs/operations) to inspect an accepted operation. Retry only idempotent work after a transient `5xx`. An interrupted speech stream has an uncertain outcome: preserve received audio as partial, keep the request identifiers, and get explicit approval before starting another paid generation. Do not automatically retry speech or duplicate a voice-creation request. Use [Text to speech](/docs/text-to-speech), [Voices](/docs/voices), and [Errors](/docs/errors) for supported request fields and responses. No public OpenAPI schema is currently served. API speech audio cannot be downloaded later through generation history; see [Generation records](/docs/generation-history). ## Authentication > Create a dashboard API key and connect your application or agent. Canonical page: https://sawtakarabi.ai/docs/authentication.md ## Create your API key 1. Open [API Keys in your dashboard](/dashboard/api-keys). Sign in or create an account if prompted. 2. Click **Create key**, enter a name such as `Claude voiceovers`, and click **Create**. 3. Copy the key from the result dialog. The full key is shown only once; if you lose it, create another and revoke the old one. 4. Configure it as `SAWTAK_API_KEY` in the environment that runs your application or agent. No scope selection is needed for dashboard-created keys. For Claude Code or Codex, see [Use with an AI agent](/docs/agent-skill). If an agent says the key is missing, it should give you these steps and wait for configuration, then continue your original task. ## Configure the key locally For Bash on macOS or Linux, enter the key at a hidden prompt, then launch your agent from that same terminal: ```bash read -r -s -p "Sawtak Arabi API key: " SAWTAK_API_KEY export SAWTAK_API_KEY printf '\n' ``` For Zsh (the default macOS shell): ```zsh read -r -s 'SAWTAK_API_KEY?Sawtak Arabi API key: ' export SAWTAK_API_KEY printf '\n' ``` An already-running agent does not inherit variables set in another terminal. For desktop or hosted execution, use its supported environment/secret settings if available. Do not paste the key into ordinary chat. Store persistent credentials in your application's secret manager; do not put them in source code, browser bundles, or public repositories. ## Send authenticated requests ```http Authorization: Bearer ``` The public API base URL is `https://api.sawtakarabi.ai/v1`. Website `/api/*` routes serve the dashboard and are not the integration API. For browser products, call the public API through your own server to keep the key private. ## Verify your key This read-only request checks the key without generating speech: ```bash curl --fail --show-error https://api.sawtakarabi.ai/v1/me \ -H "Authorization: Bearer $SAWTAK_API_KEY" ``` It returns account information, balance, permissions and limits. Do not publish its private account fields. Add credit in [Billing](/dashboard/billing) if your balance is insufficient for generation; a missing key does not establish that you need a top-up. ## Restricted credentials Dashboard-created keys work for the standard API without choosing permissions. The gateway also understands restricted credentials: `tts` for synthesis, `voices` for catalog/voice management, and `usage` for account/usage reporting. A restricted key missing a required permission receives `403 insufficient_scope`. A 403 from `/v1/me` alone does not establish that speech access is unavailable. Missing or invalid credentials return 401; repeated authentication failures can be throttled. See [Errors](/docs/errors) and [Request limits](/docs/limits). ## Text to speech > Stream Arabic speech from text with the audio endpoint. Canonical page: https://sawtakarabi.ai/docs/text-to-speech.md Send `POST /v1/audio/speech` with your dashboard API key. The endpoint follows the familiar OpenAI audio-speech request shape and streams the resulting audio in the response body. ```bash curl --fail --show-error https://api.sawtakarabi.ai/v1/audio/speech \ -H "Authorization: Bearer $SAWTAK_API_KEY" \ -H "Content-Type: application/json" \ -d '{ "model": "arabic-tts-1", "voice": "", "input": "مرحباً بك", "response_format": "wav" }' \ --output speech.wav ``` ## Request fields | Field | Required | Description | | --- | --- | --- | | `model` | No | Use `arabic-tts-1`. Omission uses this model; OpenAI-compatible model aliases are accepted and map to the same service. | | `input` | Yes | Nonblank text, at most 10,000 characters and 400 characters per word. | | `voice` | Yes | A voice ID returned by the [voice catalog](/docs/voices). Replace `` in the example with an accessible public or private voice. | | `response_format` | No | `pcm` (default) for raw audio, or `wav` for a WAV stream. Case-insensitive. | | `sample_rate` | No | Integer output rate from 8000 to 48000 Hz; default 24000. Read `X-Sample-Rate` in the response. | | `normalize_text` | No | Boolean, default `true`. Controls text normalization before synthesis. | | `normalization_locale` | No | Locale hint, default `auto`; only valid when normalization is enabled. | | `use_continuation` | No | Boolean, default `true`. Uses a saved voice’s continuation context when present. | | `enhance_pronunciation` | No | Optional boolean (`true`/`false`). Applies dialect-conditioned automatic diacritization (tashkeel) to Arabic text before synthesis to guide pronunciation. Preserves existing marks. Defaults to `false`. Use this field instead of the retired `apply_tashkeel` and `tashkeel` names. | Advanced tuning fields are optional: `cfg_value` accepts finite numbers clamped to 0–10 and `temperature` to 0–2. Omit them for normal voiceover generation. They are not speed, emotion, or duration controls. ## Text preparation The service validates the request before it is sent to a speech worker. Split unusually long passages at natural sentence boundaries, and keep words intact. Text normalization runs automatically by default. Pronunciation enhancement is a separate, opt-in step: `enhance_pronunciation: true` adds dialect-conditioned Arabic diacritics (tashkeel) after normalization, preserving existing marks. It does not select a different voice or dialect and does not guarantee better pronunciation. Set it explicitly when wanted; the official agent skill enables it unless the user asks to disable it. Pronunciation should be assessed by listening. The dialect comes from the selected saved voice. Do not send inline reference audio, `voice_dialect`, SSML, or invented speed/emotion controls. `max_generate_length` is not a client control; generation length is server-managed. The response is mono signed 16-bit PCM or streaming WAV. See [Streaming audio](/docs/streaming) for headers, finalizing WAV files, and partial-stream billing. Use [Operations and retries](/docs/operations) for status, cancellation, and duplicate protection. ## Use a custom voice Create a private voice through [Voice cloning](/docs/voice-cloning), wait until its status is ready, then place its ID in `voice`. A voice that is not available to the authenticated account returns an error instead of silently falling back to another voice. For response failures and retries, see [Errors](/docs/errors) and [Request limits](/docs/limits). ## Voices > Find the right voice and use its ID in a request. Canonical page: https://sawtakarabi.ai/docs/voices.md The Voices API lets you search, filter, and inspect voices from the catalog, audition preview audio, and retrieve the voice ID required for speech synthesis. ## Authentication and Scopes Listing voices and retrieving voice metadata require an API key, including for public voices. Pass the key in the standard HTTP `Authorization` header: ```http Authorization: Bearer ``` **Required Scope**: Dashboard-created keys have voice access without a scope-selection step. Restricted credentials must include `voices` (or have unrestricted access). If the scope is missing, the API returns HTTP `403 Forbidden` (`insufficient_scope`). Public voice **audio previews** can be played without authentication. Private previews require a key with the `voices` scope belonging to the voice owner. ## Find a dialect For an agent, use the [official skill](/docs/agent-skill): `dialects --search Najdi` finds the code and `voices --dialect saudi-najdi` resolves its alias. Direct API requests take exact codes such as `saudi-najdi`; they do not resolve English names. `eg` does not include `eg-cairene`. `search` searches voice names only. ## List Voices Retrieve a paginated list of available voices. ```http GET /v1/voices ``` ### Query Parameters All query parameters are optional: | Parameter | Type | Default | Description | | :--- | :--- | :--- | :--- | | `sort` | string | `newest` | Sort order for results. Supported values: `newest`, `oldest`, `most_liked`. Any other value returns HTTP `400` (`bad_request`). | | `limit` | integer | `25` | Number of voices to return per page. Must be an integer between `1` and `100` (inclusive). Values outside this range return HTTP `400` (`bad_request`). | | `after` | string | `null` | Opaque pagination cursor from a previous response (`next_cursor`). Used to fetch the next page of results. | | `sharing_status` | string | `null` | Filter by voice visibility: `public` or `private`. If omitted, returns the caller's own voices plus all public voices. Authentication is required for every listing request. | | `dialect` | string | `null` | Filter by dialect label. This parameter can be repeated multiple times (for example, `?dialect=saudi-najdi&dialect=eg-cairene`) to match any of the specified dialects. | | `gender` | string | `null` | Filter by gender label (for example, `male`, `female`, `neutral`). | | `age` | string | `null` | Filter by age label (for example, `young`, `middle`, `old`). | | `use_case` | string | `null` | Filter by use-case label (for example, `conversational`, `narration`, `advertisement`). | | `search` | string | `null` | Case-insensitive substring search matching against voice names. | ### Response Structure The endpoint returns a list envelope object: | Field | Type | Description | | :--- | :--- | :--- | | `object` | string | Always `"list"`. | | `data` | array | An array of voice objects matching the query. | | `first_id` | string or null | The `id` of the first voice in `data`, or `null` if the list is empty. | | `last_id` | string or null | The `id` of the last voice in `data`, or `null` if the list is empty. | | `has_more` | boolean | Indicates whether additional matching voices exist beyond the current page. | | `next_cursor` | string (optional) | An opaque cursor string included when `has_more` is `true`. Pass this value to the `after` query parameter to fetch the next page. | ## Voice Object Each item in `data`, as well as the response from `GET /v1/voices/{voice_id}`, contains the following fields: | Field | Type | Description | | :--- | :--- | :--- | | `object` | string | Always `"voice"`. | | `id` | string | Unique voice identifier. Use this value as the `voice` parameter in speech synthesis requests. | | `name` | string | Display name of the voice. | | `status` | string | Readiness status: `"ready"` or `"processing"`. Choose `"ready"` voices for generation. | | `labels` | object | Categorization metadata containing `dialect`, `gender`, `age`, and `use_case` string values. | | `preview_url` | string | URL pointing to the audio preview endpoint for this voice (`/v1/voices/{id}/preview`). | | `sharing` | object | Visibility details containing `status` (`"public"` or `"private"`) and `liked_by_count` (integer count of likes). | | `created_at` | string | Creation timestamp in UTC, formatted `YYYY-MM-DDTHH:mm:ss`. | | `description` | string (optional) | Text description of the voice, present if configured. | | `preview_text` | string (optional) | The spoken text used in the preview audio sample, present if configured. | | `continuation_status` | string (optional) | Status of continuation data (`"ready"` or `"processing"`), present only on voices created with a continuation sample. | ## Retrieve a Single Voice Fetch metadata and status for a specific voice by its identifier. ```http GET /v1/voices/{voice_id} ``` ### Access and Visibility - Returns the voice object if the voice is public or owned by the authenticated caller. - If the voice does not exist, or is private and belongs to another account, the endpoint returns HTTP `404 Not Found` (`voice_not_found`). Private voices belonging to other accounts are indistinguishable from nonexistent voices. ## Preview Voice Audio Audition a pre-rendered audio sample of a voice before using it. ```http GET /v1/voices/{voice_id}/preview ``` ### Audio Specifications and Headers - **Audio format**: Standard WAV (`audio/wav`). - **Cache-Control**: `no-store`. - **Public voices**: Unauthenticated access is allowed so that web players and client applications can stream the sample directly. - **Private voices**: Requires an authenticated API key with the `voices` scope belonging to the voice owner. - **Rate limit**: Preview requests are limited to `60` requests per minute per IP address. If this limit is exceeded, the endpoint returns HTTP `429 Too Many Requests` with a `Retry-After: 60` header. ## Choosing a Voice for Synthesis To synthesize speech using a selected voice: 1. **Find a voice**: Call `GET /v1/voices` with appropriate dialect, gender, or use-case filters. 2. **Verify readiness**: Ensure the voice object has `"status": "ready"`. Catalog entries without prepared conditioning may report `"processing"`; choose a ready voice for the straightforward path. 3. **Audition the sample**: Stream audio from `preview_url` or `GET /v1/voices/{voice_id}/preview` to evaluate the voice. 4. **Send the synthesis request**: Pass the voice `id` in the `voice` field of your speech synthesis request: ```bash curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "model": "arabic-tts-1", "voice": "", "input": "مرحبا بك", "response_format": "wav" }' \ --output speech.wav ``` The `voice` parameter is required. Omitting it returns HTTP `400` (`bad_voice`); there is no default-voice fallback. ## Error Reference The Voices API returns standard error responses: | HTTP Status | Error Code | Description | | :--- | :--- | :--- | | `400 Bad Request` | `bad_request` | Invalid parameter, such as an unrecognized `sort` value, `limit` outside `1`..`100`, an unparseable or mismatched pagination cursor, or an invalid `sharing_status`. | | `401 Unauthorized` | `auth_missing` / `auth_failed` | Missing or invalid API key when listing voices, retrieving metadata, or accessing private voice resources. | | `403 Forbidden` | `insufficient_scope` | The provided API key lacks the required `voices` scope. | | `404 Not Found` | `voice_not_found` | The requested voice ID does not exist, lacks preview audio, or is private and not owned by the caller. | | `429 Too Many Requests` | `rate_limited` | The client exceeded the preview rate limit of 60 requests per minute per IP address. Check the `Retry-After` header. | | `503 Service Unavailable` | `service_unavailable` | The voice catalog service is currently not configured or unavailable. | ## Voice cloning > Create and use an authorized cloned voice. Canonical page: https://sawtakarabi.ai/docs/voice-cloning.md The Voice Cloning API allows you to create custom voice clones from audio samples and use the resulting voice ID for speech synthesis. ## Authentication and Scope All voice cloning endpoints require authentication using a Bearer token in the HTTP `Authorization` header: ```http Authorization: Bearer ``` - **API key**: Use a dashboard-created key; no permission selection is needed. Restricted credentials require `voices`, otherwise the gateway returns `403 insufficient_scope`. - **Anonymous Access**: Keyless or anonymous requests are not permitted and return HTTP `401 Unauthorized` (`auth_missing`). ## Clone a Voice To create a new cloned voice, submit a `multipart/form-data` request to the voices endpoint: ```http POST /v1/voices Content-Type: multipart/form-data ``` ### Form Fields | Field | Type | Required | Description | | :--- | :--- | :--- | :--- | | `files` | binary | Yes | Exactly one nonempty reference audio file, at most 20 MiB. The whole multipart request must be at most 35 MiB. | | `name` | string | Yes | Nonblank display name, at most 80 characters. | | `labels` | string (JSON) | Yes | A JSON-encoded object containing categorization metadata. Must include `dialect`. | | `description` | string | No | Optional text description (maximum 180 characters). | | `continuation_file` | binary | No | Optional secondary audio clip, at most 20 MiB, providing delivery context. Must be paired with `continuation_text`. | | `continuation_text` | string | No | Spoken transcript matching `continuation_file`, at most 1000 characters. Must be paired with `continuation_file`. | #### Labels JSON Object The `labels` field must be passed as a serialized JSON object with the following properties: | Property | Type | Required | Description | | :--- | :--- | :--- | :--- | | `dialect` | string | Yes | Dialect identifier. Must match a supported dialect code. | | `gender` | string | No | Voice gender. Allowed values: `male`, `female`, `neutral` (defaults to `male`). | | `age` | string | No | Voice age group. Allowed values: `young`, `middle`, `old` (defaults to `middle`). | | `use_case` | string | No | Intended application. Allowed values: `conversational`, `narration`, `advertisement`, or empty string `""` (defaults to `""`). | ### Example Request ```bash curl -X POST "https://api.sawtakarabi.ai/v1/voices" \ -H "Authorization: Bearer " \ -F "name=Custom Voice" \ -F 'labels={"dialect":"saudi-najdi","gender":"male","age":"middle","use_case":"conversational"}' \ -F "description=Cloned voice for conversational interactions" \ -F "files=@reference_audio.wav" ``` ### Response A successful request returns HTTP `201 Created` with the assigned voice ID after processing has completed: ```json { "object": "voice", "id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4", "status": "ready" } ``` If a continuation sample was provided, the response also includes `continuation_status`: ```json { "object": "voice", "id": "7b8e62c1-9689-4a7b-a320-912b7a94f6f4", "status": "ready", "continuation_status": "ready" } ``` ## Readiness Cloning is synchronous: the request waits for audio processing and storage, then returns HTTP 201 with `status: ready`. When a continuation sample was supplied, `continuation_status` is also `ready`. Save the returned ID and use it for speech; ordinary successful clones do not require a polling loop. `GET /v1/voices/{voice_id}` retrieves current metadata. Catalog entries without prepared conditioning can report `processing`; this does not make a new clone request an asynchronous job. Do not submit a second clone merely because the first request timed out. Keep its request ID and inspect [operation status](/docs/operations) when available. ## Using the Resulting Voice ID Once `status` is `ready`, pass the voice `id` in the `voice` parameter of a speech synthesis request: ```bash curl -X POST "https://api.sawtakarabi.ai/v1/audio/speech" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "model": "arabic-tts-1", "voice": "", "input": "مرحبا بك، هذا اختبار للصوت المستنسخ.", "response_format": "wav" }' \ --output output.wav ``` ### Continuation Control For voices created with continuation audio, synthesis requests use the continuation context by default. You can disable this behavior by setting `"use_continuation": false` in the synthesis request payload. ## Recovery States and Resource Management ### Error Codes and Recovery | HTTP Status | Error Code | Cause and Recovery | | :--- | :--- | :--- | | `400 Bad Request` | `bad_request` | Invalid multipart body, missing or multiple audio files in `files`, empty audio, invalid JSON in `labels`, missing or unsupported `dialect`, invalid `gender`/`age`/`use_case`, description exceeding 180 characters, or unsupported audio data. Correct the input fields and resubmit. | | `401 Unauthorized` | `auth_missing` / `auth_failed` | Missing or invalid API key. | | `402 Payment Required` | `insufficient_balance` | Add credit before another clone attempt. | | `403 Forbidden` | `insufficient_scope` | Restricted credential lacks voice access. Dashboard-created keys need no scope selection. | | `404 Not Found` | `voice_not_found` | Voice is missing or not accessible to this account. | | `409 Conflict` | `duplicate_request` / `operation_changed` | Existing submission or a cancelled/closed operation. Inspect the operation rather than blindly uploading again. | | `413 Payload Too Large` | `body_too_large` | The multipart upload exceeds 35 MiB or did not arrive within the upload deadline. Individual files over 20 MiB return `400 bad_request`. | | `429 Too Many Requests` | `concurrency_limited` / `capacity_full` | Respect `Retry-After`. | | `503 Service Unavailable` | `no_backend` / `upload_failed` / `service_unavailable` | Processing capacity or infrastructure failed. Inspect an existing operation before retrying uncertain work. | Missing continuation pairs, overlong continuation text, and invalid metadata return `400 bad_request`. The API does not currently expose a per-account clone-count limit; upload and concurrency limits still apply. Successful cloning is charged when the voice is committed. See [Operations and retries](/docs/operations) for duplicate protection. ### Updating Voice Metadata You can update the name, description, or labels of an existing cloned voice: ```http PATCH /v1/voices/{voice_id} Authorization: Bearer Content-Type: application/json { "name": "Updated Voice Name", "description": "Updated voice description", "labels": { "use_case": "narration" } } ``` A successful update returns HTTP `200` with `{"ok":true}`. Only the owner can update or delete a voice. ### Deleting a Voice To delete a cloned voice you own: ```http DELETE /v1/voices/{voice_id} Authorization: Bearer ``` A successful deletion returns HTTP `204 No Content`. ## Streaming Audio > Receive, process, and safely play real-time streamed audio responses from POST /v1/audio/speech. Canonical page: https://sawtakarabi.ai/docs/streaming.md The Text-to-Speech API delivers audio in real time over HTTP. This guide explains the wire protocol, supported audio formats, how to read response headers, safe audio playback practices, and system behavior when a stream is interrupted. ## Wire Delivery Overview ```http POST /v1/audio/speech ``` When synthesis begins, the gateway responds with HTTP `200 OK` and streams audio data chunk-by-chunk: - **Transfer Mechanism**: Incremental response bytes. HTTP/1.1 may use chunked transfer encoding; HTTP/2 uses data frames. Do not depend on a `Transfer-Encoding` header. - **No Content-Length**: Because audio is synthesized dynamically in real time, the response omits the `Content-Length` header. - **Immediate Streaming**: Chunks are forwarded to the client as soon as they are produced by the inference engine, minimizing time-to-first-byte (TTFB). ## Supported Audio Formats Specify the target audio format in the JSON request body using the `response_format` parameter. | Format | `response_format` | Media Type (`Content-Type`) | Specifications | | :--- | :--- | :--- | :--- | | **Linear PCM** | `"pcm"` *(default)* | `audio/pcm` | Raw uncompressed 16-bit linear PCM, signed little-endian (`s16le`), mono (1 channel). Sample rate defaults to 24,000 Hz (`24000`), or matches the requested `sample_rate` (8,000–48,000 Hz). Contains no container header or metadata padding. | | **WAV** | `"wav"` | `audio/wav` | Standard 44-byte RIFF/WAVE header prepended to the 16-bit mono linear PCM stream. Uses streaming length conventions (`0xFFFFFFFF`). | ### Unsupported Formats and Codec Pitfalls - **No MP3 support**: Compressed audio formats such as `mp3` are not supported. Requesting an unsupported format returns HTTP `400 Bad Request` with error code `unsupported_response_format`. - **Client framework defaults**: Several client SDKs and agent frameworks (such as LiveKit) default to requesting `mp3`. You must explicitly set `response_format: "pcm"` or `"wav"`. - **Silent decoding failures**: Feeding raw PCM bytes directly into an MP3 or AAC decoder results in loud static distortion and noise without triggering an HTTP or decoder error. ## Actual Response Headers Inspect HTTP response headers to configure audio sinks, inspect sample rates, and correlate requests. | Header | Example Value | Description | | :--- | :--- | :--- | | `Content-Type` | `audio/pcm` or `audio/wav` | Media type corresponding to the requested `response_format`. Returns `application/json` on pre-stream errors. | | `X-Sample-Rate` | `24000` | The effective audio sample rate in Hz. Reflects either the default (`24000`) or the requested `sample_rate` (8000–48000). Always configure your audio output device using this value. | | `X-Audio-Sample-Rate` | `24000` | Mirror header for `X-Sample-Rate`. | | `X-Channels` | `1` | Number of audio channels (`1` for mono). | | `X-Audio-Channels` | `1` | Mirror header for `X-Channels`. | | `X-Request-Id` | `a1b2c3d4-...` | Unique request UUID generated by the gateway. Use this identifier for log correlation, debugging, and billing inquiries. | | `Cache-Control` | `no-store` | Instructs caches not to store the response. | ## Safe PCM and WAV Playback Guidance Because streamed audio is delivered in arbitrary chunk sizes over the network, client applications must handle buffering and audio sink initialization carefully. ### Playing Raw PCM (`audio/pcm`) Raw PCM contains purely audio sample values with no header bytes to identify format, channels, or sample rate. 1. **Pre-initialize the Audio Sink**: Configure the audio device or Web Audio context explicitly before feeding received chunks: - **Sample Format**: Signed 16-bit integer, little-endian (`s16le` / `Int16Array`). - **Channels**: 1 (mono). - **Sample Rate**: Read dynamically from the `X-Sample-Rate` response header (defaults to `24000` Hz). 2. **Preserve Sample Boundary Alignment**: - Each audio sample is 16 bits (2 bytes). - Network chunks may arrive with an odd number of bytes. Never pass an incomplete 2-byte sample to an integer conversion buffer. - If an incoming network chunk ends on an odd byte boundary, retain the dangling single byte and prepend it to the next chunk before converting to 16-bit PCM samples. 3. **Queue and Jitter Buffer**: Maintain a short jitter buffer (e.g., 50–100 ms of audio) before initiating playback to prevent audio dropouts or underruns caused by network latency variations. ### Playing Streaming WAV (`audio/wav`) The `"wav"` format wraps the 16-bit mono PCM stream with a 44-byte RIFF/WAVE header in the very first streamed chunk. 1. **Streaming Length Markers (`0xFFFFFFFF`)**: Because the total audio length is unknown while streaming, the gateway sets the 4-byte RIFF chunk size (offset 4) and the `data` subchunk size (offset 40) to `0xFFFFFFFF` (`4,294,967,295` bytes). This is the standard convention accepted by streaming-aware WAV players. 2. **Streaming vs. Non-Streaming Players**: - **Streaming decoders**: Tools and libraries such as FFmpeg or streaming browser audio decoders accept `0xFFFFFFFF` and play the stream continuously until the HTTP connection closes. - **Static file decoders**: Naive WAV decoders that expect a complete static file and seek to the file end based on the header length will report an invalid file. 3. **Persisting to Disk**: If saving streamed WAV audio to a file for later playback in standard desktop media players: - Accumulate total audio data bytes received (`data_bytes`). - Once the stream finishes, seek back to offset 4 and write the 32-bit little-endian value `data_bytes + 36` (the RIFF chunk size). - Seek to offset 40 and write the 32-bit little-endian value `data_bytes` (the `data` subchunk size). - Alternatively, strip the initial 44-byte header and store or process the payload as raw `s16le` PCM. ## Behavior on an Interrupted Response Understanding how the gateway meters requests, handles disconnections, and recovers from errors ensures resilient client integration. ### Metering and Stream Lifecycle The gateway checks that the balance covers the request before generation. It does not debit and then refund a prepayment. It settles once the operation ends: - If audio was delivered by the gateway, the request is charged for its input text, including partial streams and cancellations after audio delivery. - If no audio was delivered, the speech charge is zero. - Delivering a WAV header alone does not count as audio delivery. - Settlement can finish after the client disconnects. The gateway's delivery accounting cannot prove the client saved or listened to the bytes. Before streaming begins, failures return a JSON error and non-200 status. After response headers have been sent, a failure cannot change HTTP 200; the gateway terminates the stream. A successful header therefore does not prove complete audio. A client disconnect signals cancellation. Backend and concurrency resources are released as the operation finishes; cancellation is not an automatic refund. Network drops, client termination, timeouts, and backend failures can all interrupt a stream. ### What the client should do Save `X-Request-Id` immediately, preserve complete received PCM samples as partial audio, and inspect [operation status](/docs/operations) when possible. Drop any final unmatched byte rather than treating it as a complete 16-bit sample. Mark interrupted files incomplete even when playable. A completed operation reports `completion_bytes` for PCM audio only. Exclude the 44-byte WAV header when comparing it with received bytes when diagnosing a suspiciously short response. API audio is not archived for later recovery, and missing operation history does not authorize another paid generation. The [official skill](/docs/agent-skill) handles local metadata and partial WAV files for you. ## Operations and retries > Inspect generation status, cancel work, and avoid duplicate submissions. Canonical page: https://sawtakarabi.ai/docs/operations.md Speech and voice cloning create operations. Keep the server's `X-Request-Id` as soon as response headers arrive. For an accepted speech request or successful clone, this is the operation ID. Errors before an operation is created still have a tracing request ID, which may not identify an operation. ## Inspect an operation ```bash curl --fail --show-error "https://api.sawtakarabi.ai/v1/operations/$OPERATION_ID" \ -H "Authorization: Bearer $SAWTAK_API_KEY" ``` The owning account can inspect its operation. The response includes: | Field | Meaning | | --- | --- | | `id` | Server operation ID. | | `state` | `pending` or `terminal`. | | `outcome` | `null` while pending; otherwise `completed`, `stream_failed`, `cancelled`, `failed`, or `not_accepted`. | | `charged_micros` | Charge in millionths of a US dollar. A pending operation's value is not a final charge. | | `completion_bytes` | Present for completed speech; delivered PCM audio bytes, excluding the header for WAV output. | Use bounded polling when an operation is pending. Operation records are temporary: caller-supplied idempotency keys normally keep the record for 24 hours from creation; operations without a supplied key normally expire one hour after settlement. Cleanup can lag. This is not an audio archive or permanent billing history. A 404 can mean an expired, unknown, or other account's operation; it does not establish whether speech was generated or charged. ## Cancel an operation ```bash curl --fail --show-error -X POST \ "https://api.sawtakarabi.ai/v1/operations/$OPERATION_ID/cancel" \ -H "Authorization: Bearer $SAWTAK_API_KEY" ``` Returns `{"ok":true}` after signaling cancellation of running work. It is not confirmation of a refund or terminal state. Inspect the operation again for the settled result. Cancellation after audio delivery may still be charged. A terminal operation is not restarted or refunded by this call. ## Prevent duplicate submissions Send a client-generated `Idempotency-Key` header for a speech or cloning request and keep it separately from the server operation ID. It must contain 1–128 UTF-8 bytes. Without this header, the gateway generates a new key for each request, so separate submissions are not recognized as duplicates. For an existing pending, completed, or charged operation: - Reusing the key returns HTTP `409` with `error.code: duplicate_request`, the original `operation_id`, `state`, and `outcome`. It does not replay audio. - For speech, reusing the key with different request fields returns `409 idempotency_conflict`. - Cloning deduplicates by key without comparing the uploaded body; use a new key for a deliberately new clone. Keys are scoped to the account and operation type. They are not a permanent lock: a terminal operation that did not complete and was not charged releases its key for reuse, and expired records can be cleaned up. A retry can therefore start new work; do not treat submitting the same key as a read-only status check. ## Interrupted speech Save audio locally while streaming. Inspect status using the operation ID when available, preserve partial audio as incomplete, and ask before another paid generation. No endpoint retrieves missing audio from an interrupted API response. Neither a 200 response header nor a playable partial file proves completion. See [Streaming audio](/docs/streaming) and [Generation records](/docs/generation-history). Operation inspection and cancellation require the owning account's key. Restricted credentials need `tts` for speech operations or `voices` for clones. They share a separate allowance of 600 requests per minute, with burst protection, so normal generation limits do not prevent cancellation. ## Usage > Inspect the authenticated account and its API usage. Canonical page: https://sawtakarabi.ai/docs/usage.md Use your dashboard API key with `GET /v1/me` for account information and `GET /v1/usage` for daily usage. Restricted credentials require `usage` for both endpoints. ```bash curl --fail --show-error https://api.sawtakarabi.ai/v1/usage \ -H "Authorization: Bearer $SAWTAK_API_KEY" ``` Usage is an account-level record. Treat it as reporting data: display it in a dashboard, reconcile it with internal records, and avoid making availability decisions from a single response. Use the [Request limits](/docs/limits) guide to shape traffic instead of deriving a limit from usage data. If the key is scoped, it must include `usage`; otherwise the API returns an insufficient-scope error. ## History and date ranges `start_time` (inclusive) and `end_time` (exclusive) are Unix seconds. A query covers at most 366 days; `limit` is 1–366 daily results (default 31). Results are newest first; pass the returned `last_id` as `after` for the next page with the same date range. Usage is available for the last twenty-four calendar months as UTC daily totals. Both range boundaries and the `after` cursor must be UTC midnight (a Unix timestamp divisible by 86400); otherwise the API returns HTTP 400. The default range is the thirty UTC days ending at tomorrow's midnight, including today's exported activity. A range starting before the retention window also returns HTTP 400. Reporting is asynchronous: usage is exported hourly, and processing backlogs or storage outages may delay it further. The balance is updated when the request settles; it does not wait for the reporting dashboard. Administrators can inspect captured troubleshooting text for ninety days; it is not returned by this endpoint. See [Generation records](/docs/generation-history). Totals include completed requests and charged partial streams. Free failed work does not contribute to billable usage. Refunds and account credits remain in the separate financial ledger; usage totals describe the original generation charges. ## Response fields `GET /v1/me` returns `object: "me"`, `user_id`, `email`, `key` (`id`, `scopes`), `balance_usd` as a decimal string, and `limits` (`requests_per_minute`, `concurrent_requests`, `generation_suspended`, `voices_per_account`). A null request/concurrency limit means disabled; `voices_per_account` is currently null. Keep private account details out of public logs. `GET /v1/usage` returns `object: "list"`, `data`, `first_id`, `last_id`, and `has_more`. Each data item contains `object: "usage"`, `start_time`, `end_time`, `tts_characters`, and `cost_usd` as a decimal USD string. It includes days with recorded activity; do not assume an item for every calendar day. Unlike voice pagination, usage pagination uses `last_id` rather than `next_cursor`. ## Generation records > Understand saved generation history and troubleshooting retention. Canonical page: https://sawtakarabi.ai/docs/generation-history.md Use [Operations and retries](/docs/operations) for short-lived request status and billing inspection. This page describes longer-term troubleshooting records and dashboard playback history. Keep the `X-Request-Id` returned by `POST /v1/audio/speech` when contacting support. Authorised administrators can inspect recorded customer requests for troubleshooting. Request metadata, input text, processed text and diagnostic payloads are retained together for ninety days, using UTC-day boundaries, when capture succeeds. Capture is best-effort: missing diagnostics do not mean the speech request failed or was uncharged. History is browsed by customer and date. New records appear after hourly background export; processing backlogs or storage outages may delay it further. Daily account totals are retained for twenty-four months from the usage date. Read those totals through [Usage](/docs/usage). The dashboard's saved generation history is separate: it keeps the newest twenty entries per generation type, for up to thirty days, including saved text and audio. Deleting an entry schedules its stored files for deletion. The current gateway does not expose `/v1/audio/generations/:id` or a generation-audio retrieval endpoint. Ordinary API output is streamed, not automatically archived for later download. Save the audio in your application if you need a durable copy. If a stream is interrupted, preserve received audio as partial and keep the server `X-Request-Id` separately from any client-supplied `Idempotency-Key`. Neither identifier can restore missing audio. Missing diagnostics do not establish whether generation completed or was charged; do not automatically submit another paid generation. These limits do not control financial-record retention. Billing records remain separate from troubleshooting content. See the [Privacy Policy](/privacy). ## Request limits > Handle request, upload, and concurrency limits reliably. Canonical page: https://sawtakarabi.ai/docs/limits.md Standard accounts allow **60 requests per minute** and **5 concurrent requests**. These limits are shared across all API keys and dashboard tokens belonging to the account. Speech and voice cloning share the concurrent-request allowance. Requests per minute refill steadily with a small burst allowance; sending the entire minute's quota at once may return `429`. High-volume customers can receive custom limits, including no account-level rate or concurrency limit. GPU capacity and upload protections still apply. There is no characters-per-minute quota; request size limits and usage charges remain unchanged. `GET /v1/me` returns the effective `limits.requests_per_minute` and `limits.concurrent_requests`; `null` means that account limit is disabled. `limits.generation_suspended` indicates whether new speech and cloning are suspended. Administrative changes apply within 60 seconds. Running requests can finish. Operation status and cancellation have a separate allowance of 600 requests per minute (with burst protection), so exhausting the regular allowance does not prevent cancellation. ## When a request is limited The API returns `429` with an OpenAI-compatible error object. The response includes `Retry-After`. Wait at least that long before retrying. ```text 429 rate_limited Retry-After: 60 ``` Use exponential backoff with jitter, and bound the number of retries. Do not immediately retry a request that was rejected for concurrent work; wait for an in-flight request to finish first. ## Request sizes The free website demo allows 10 generation requests per minute and 3 simultaneous generations per visitor address, with up to 300 characters per request. Visitors sharing an address share these limits. Speech input is limited to 10,000 characters and 400 characters per word; JSON bodies are limited to 65,536 bytes. Clone requests accept at most 35 MiB of multipart data, with each audio file at most 20 MiB. Clone names allow 80 characters, descriptions 180, and continuation transcripts 1000. Split long text at sentence boundaries. ## Capacity Each account can have up to 20 active API keys and 100 voice collections. Revoked API keys do not count toward the active-key limit. Limits are enforced at the account level, including when several keys belong to the same account. If your production workload needs more capacity, contact support with the expected request volume, concurrent work, and audio duration. ## Errors > Read and handle API errors consistently. Canonical page: https://sawtakarabi.ai/docs/errors.md Errors use an OpenAI-compatible envelope. Check the HTTP status first, then use `error.code` for programmatic handling. Keep `request_id` when it is present so support can trace a request. ```json { "error": { "message": "Invalid API key", "type": "authentication_error", "code": "auth_failed", "param": null, "request_id": "…" } } ``` ## Common responses | Status | Meaning | Client action | | --- | --- | --- | | `400` | Invalid input or unsupported option | Correct the request; do not retry unchanged. | | `401` | Missing, invalid, or expired credential | Replace the credential. | | `402` | Insufficient balance | Add balance before retrying. | | `403` | The key lacks the needed scope | Use a key with the required permission. | | `404` | Resource is unavailable to this account | Check the identifier and access. | | `409` | Duplicate submission, conflicting idempotency key, or changed operation | Read `error.code`; inspect the existing operation. A duplicate does not replay audio. | | `413` | Request or upload is too large | Reduce the request or split the file. | | `429` | Rate or authentication throttle | Respect `Retry-After`, then retry with backoff. | | `5xx` | Temporary API or capacity failure | Retry idempotent work with bounded backoff. | Do not present raw API messages directly to end users. Map stable error codes to clear product copy and log the request ID separately. After speech headers have been sent, a stream failure cannot become a JSON error response. Preserve the server ID and partial audio, then use [Operations and retries](/docs/operations). A new synthesis request can incur another charge.