# OpenAI-compatible server

Source: https://docs.vieneu.io/docs/sdk/remote-mode

`apps/openai_speech.py` in the repo serves `POST /v1/audio/speech` exactly like OpenAI's TTS endpoint (`pcm`/`wav`, chunked body or SSE). The OpenAI SDK, Pipecat, LiveKit Agents and the Vercel AI SDK work by changing `base_url`.

## Start the server

Pick one; all listen on `http://localhost:8000`:

```bash
uv run python -m apps.openai_speech                                  # from a repo checkout, auto-detects GPU/CPU
docker compose -f docker/docker-compose.yml --profile api-gpu up     # Docker, GPU
docker compose -f docker/docker-compose.yml --profile api-cpu up     # Docker, CPU only (torch-free)
```

Measure time-to-first-audio and RTF on your own machine:

```bash
uv run python examples/openai_speech_client.py --bench 8
```

## Call it

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:8000/v1", api_key="x")
with client.audio.speech.with_streaming_response.create(
    model="vieneu-v3-turbo",
    voice="Mai Anh",
    input="Xin chào! Đây là chế độ streaming của VieNeu.",
    response_format="pcm",
) as r:
    for chunk in r.iter_bytes(4096):      # s16le 48 kHz mono, as it is generated
        play(chunk)
```

## Endpoints

| Method | Path | Purpose |
|---|---|---|
| `POST` | `/v1/audio/speech` | Synthesize; `response_format` `pcm` or `wav`, streamed |
| `GET` | `/v1/models` | Model list |
| `GET` | `/v1/voices` | Preset and enrolled voices |
| `POST` | `/v1/voices` | Clone from an uploaded clip |
| `GET` | `/health` | Liveness |

Concurrency is capped per backend with `VIENEU_MAX_STREAMS` (default 16 on GPU, 1 on CPU). Requests beyond the cap wait in a small queue, then get `429`. First audio arrives in ~115 ms with 16 streams on an RTX 3060; ~140–300 ms and 1–2 streams on CPU. Full numbers: [`docs/streaming.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/streaming.md).

## Web UI in Docker

```bash
docker compose -f docker/docker-compose.yml --profile gpu up   # or --profile cpu → http://localhost:7860
```

Production images and builds: the repo's [`docs/Deploy.md`](https://github.com/pnnbao97/VieNeu-TTS/blob/main/docs/Deploy.md) and our [Docker](/docs/deployment/docker) page.

## Hosted instead of self-hosted

If you do not want to run a GPU, the [VieNeu Cloud API](/docs/cloud-api/openai-compatible) exposes the same OpenAI-compatible shape at `api.vieneu.io`, including the v4 engine.

## Legacy: v2 `remote` mode (deprecated)

:::warning
The LMDeploy server on port 23333 and `Vieneu(mode="remote")` only work with **VieNeu-TTS v2**, which is no longer updated. They are kept for existing deployments. For v3 Turbo use the streaming server above.
:::

```bash
docker run --gpus all -p 23333:23333 -v huggingface_cache:/root/.cache/huggingface pnnbao/vieneu-tts:latest --tunnel
pip install "vieneu[legacy]"
```

```python
from vieneu import Vieneu

tts = Vieneu(mode="remote", api_base="http://your-server-ip:23333/v1",
             model_name="pnnbao-ump/VieNeu-TTS-v2", emotion="natural")
audio = tts.infer(text="Chào bạn!")
tts.save(audio, "remote_output.wav")
```

Fine-tuned v3 Turbo models are not served by that container; load them with the SDK (`Vieneu(mode="v3turbo", backbone_repo=...)`). See [Remote server](/docs/deployment/remote-server) for the old Docker flags.
