Skip to main content

MCP server — VieNeu in Claude, ChatGPT and more

VieNeu runs a hosted Model Context Protocol (MCP) server. Add one URL to an AI assistant, sign in with your VieNeu account, and ask for Vietnamese speech in plain language — the assistant finds a voice, synthesizes the text, and plays it back or hands you a download link.

https://api.vieneu.io/mcp

Nothing to install. The assistant signs in through your browser (OAuth 2.1) the first time you use it; tools and apps that cannot sign in can send an API key instead.

Before you start: you need a VieNeu account with an active token plan. Connecting is refused without one, because every synthesis would fail.

Pick your app​

AppHow you connectInline playerGuide
Claude (web, desktop, mobile)Add a custom connector, sign inYesClaude
ChatGPTDeveloper mode → add an app, sign inYesChatGPT
Claude CodeOne command, then /mcp—Claude Code
CursorOne click or mcp.json; sign in or API keyYesCursor
VS Code (GitHub Copilot)One click or mcp.json; sign in or API keyYes (experimental)VS Code
OpenAI API / Agents SDKmcp tool in your code—OpenAI API
Windsurf (Devin Desktop), Gemini CLI, Codex CLI, Zed, Goose, LM StudioConfig fileGoose onlyOther apps

Something not working? See Troubleshooting.

What you say to it​

"What can VieNeu do?"

The assistant calls list_capabilities and answers with what works here, one example request for each, and which ones spend tokens. In apps with the inline player it is a card: click an example and it is sent to the chat.

"Find me a northern female voice for news."

The assistant calls list_voices. Free.

"Read this paragraph with Thu Trang."

The assistant calls text_to_speech. In apps with the inline player, a small player appears right in the chat — press play. There is always a download link too (valid 24 hours; MP3 for texts up to 1,500 characters). The audio is saved in your Library on vieneu.io. This spends tokens from your plan, exactly like the same request made through the API.

"Add a laugh after the first sentence."

The assistant calls list_emotion_tags and writes a cue tag such as [cười] into the text before synthesizing.

"Make an audiobook of this story — a northern woman narrating, a voice for each character."

The assistant scripts your text, chooses the voices, tells you the cost, and after you agree VieNeu makes one MP3 per chapter in the background. See Audiobooks.

Tools​

ToolWhat it doesCost
list_capabilitiesWhat this connection can do, with an example request for eachFree
list_voicesSearch the voices your account can use, including your own cloned voicesFree
text_to_speechSynthesize Vietnamese text; returns a download link and the durationTokens
get_speech_statusCheck a long job that was still processing; returns the link when doneFree
list_emotion_tagsReading styles, and cue tags such as [cười] to write into the textFree
get_token_balanceTokens you can spend right now, your plan's caps and when they reset, and roughly how many characters that readsFree

Six more tools — create_audiobook, add_audiobook_chapter, start_audiobook, get_audiobook, list_audiobooks, cancel_audiobook — make audiobooks; they are described in Audiobooks.

list_voices​

ArgumentTypeNotes
searchstring, optionalA few words describing the voice (nữ miền Bắc trầm ấm, male south), or its name (Thu Trang). How the words are read is below.
limit1–100, optionalDefault 25. The reply says how many more matched.

Each line reads id (gender, region) "Name" - description. The id is what text_to_speech needs, exactly as shown — ids look like names (Thu Trang) and cannot be guessed.

How search reads the words:

  • Gender and region come from the voice's fields, never from its description: nữ, nam, miền Bắc, miền Nam, miền Trung, or female, male, north, south, central. nam on its own means male; write miền Nam for the South.
  • Other words match whole words of the voice's id, name or description, ignoring case. A word written with diacritics must match them (già, old, does not find a voice named Gia); written without, it matches either way, and from 4 letters also as the start of a word.
  • Words that describe no voice, such as giọng or đọc, are ignored.
  • A voice's name (Thùy An, giọng Thu Trang) puts that voice first.
  • When no voice has every word, the closest come back instead, those with the gender and region asked for first, and each line ends with the words that voice lacks: [thiếu: trầm, ấm].

text_to_speech​

ArgumentTypeNotes
textstringVietnamese text, up to 50,000 characters. Numbers, dates and English acronyms are handled.
voicestring, optionalA voice id from list_voices. Omitted: the default voice.
speed0.5–2.0, optional1.0 is natural pace.
stylestring, optionalA reading style from list_emotion_tags.

Up to 1,500 characters the tool makes one POST /v1/audio/speech call for MP3 and returns its link a moment later. Longer text becomes one POST /v1/tts job (WAV) that the tool waits on for about 50 seconds; a very long text may still be processing then — the reply gives a job id and tells the assistant to call get_speech_status later instead of synthesizing again (which would be charged again). Either way the reply carries the link, its format, the duration and the voice — and, once the audio is ready, how many tokens you have left.

get_speech_status​

ArgumentTypeNotes
job_idstringThe id text_to_speech returned.

get_token_balance​

No arguments. Ask "còn bao nhiêu token?" and the assistant answers with:

  • Usable now — the most one synthesis can spend at this moment: the plan's remaining tokens, or less when a daily or weekly cap is lower.
  • The plan's remaining / total, each cap and when it resets (Vietnam time), and when the plan expires.
  • The total across all your active plans, if you have more than one — the next one takes over when this one runs out.
  • About how many characters that reads with the default voice engine.

At zero it links to the pricing page. The same numbers are on GET /v1/balance for scripts.

The inline player​

In apps that support MCP Apps — Claude (web, desktop and mobile), ChatGPT, Cursor, VS Code (experimental) and Goose — text_to_speech and get_speech_status results show a player in the conversation: play/pause, seek, duration, and a download button. While a long job is still running it shows a small green light and keeps checking get_speech_status by itself (free), switching to the player when the audio is ready.

Audiobooks get their own player under start_audiobook and get_audiobook: the chapters with their progress, played one after another, refreshing itself while the book is being made.

Claude asks once before showing it — choose Allow (or Always allow). Apps without MCP Apps show the text reply with the link instead; nothing else changes.

Ready-made prompts​

The server also offers five prompts — ready-made requests with a form for their details. Apps that show MCP prompts list them in a menu: Claude Code as /mcp__vieneu__make_audiobook, VS Code as /mcp.vieneu.make_audiobook (the middle part is the name you gave the server). ChatGPT does not show prompts; ask "VieNeu làm được gì?" there instead.

PromptWhat it asks forDetails
what_can_vieneu_doThe list of what VieNeu can do—
read_textRead a text aloudtext, voice (optional, e.g. "nữ miền Bắc")
make_audiobookAn audiobook from your text, priced before it startstext (optional — paste it after), narrator (optional)
find_voiceVoices matching a descriptiondescription
check_balanceTokens left and the plan's limits—

Using an API key instead of signing in​

Scripts, the OpenAI API, and apps that cannot sign in can send one of your API keys on every request, in either header:

X-API-Key: vn_sk_...
Authorization: Bearer vn_sk_...

Create a key on the Developer page of vieneu.io. A config file holding a key is a password on disk — prefer signing in where the app supports it, and never commit a key to a repository. Each app's guide shows where the header goes.

Billing​

A tool call costs exactly what the equivalent API call costs: one synthesis per text_to_speech call, charged per character (minimum 50) at the engine's rate, from your plan's tokens. A failed synthesis is refunded automatically. An audiobook is billed the same way, chapter by chapter as each one is made, after you agree to start_audiobook. Every other tool is free.

MCP usage appears under Developer → Usage like any API traffic, attributed to the connection (for example "Claude (MCP)").

See and remove connections​

Developer → Dùng VieNeu trong Claude, ChatGPT on vieneu.io shows the URL to copy and lists every app you have connected, with the date and the last use. Gỡ kết nối (Disconnect) signs that app out immediately; it has to ask for permission again to come back.

Errors​

The tools translate API errors into sentences the assistant can act on.

What you seeCauseFix
Asked to sign in againThe connection was removed on vieneu.io, or its sign-in expiredReconnect from the app
"Tài khoản không đủ quyền hoặc hết token…"The plan is out of tokens, expired, or does not include the engine (HTTP 403)Top up or renew on vieneu.io — retrying does not help
"Đang gọi quá nhanh…"Rate limited (HTTP 429)Wait the number of seconds quoted
"Máy chủ tạo giọng đang bận"No synthesis worker free (HTTP 503)Retry in a few seconds
"Voice … is not available"The voice id is not in your cataloguePick an id from list_voices

More in Troubleshooting.

What is deliberately not exposed​

Voice cloning, dubbing and SRT dubbing are not tools here. They take file uploads and spend noticeably more per call, and an assistant invoking one by accident is a bad first experience. Use the web app or the Cloud API for those.

For client developers​

  • Transport: Streamable HTTP, stateless. Protocol revision 2026-07-28, with 2025-era clients served too.
  • Authorization: OAuth 2.1 per the MCP authorization spec. An unauthenticated request gets 401 with WWW-Authenticate: Bearer resource_metadata="https://api.vieneu.io/.well-known/oauth-protected-resource".
  • Metadata: /.well-known/oauth-protected-resource (RFC 9728) and /.well-known/oauth-authorization-server (RFC 8414).
  • Clients: Client ID Metadata Documents and Dynamic Client Registration (/oauth/register). PKCE S256 is always required. A metadata document must allow none (preferred or among token_endpoint_auth_methods_supported). Registration accepts none, client_secret_post and client_secret_basic and always returns a client_secret; it is required only for the two secret methods. Loopback redirects (http://127.0.0.1 / localhost) match on any port.
  • Scope: tts (plus offline_access for a refresh token). Access tokens last one hour; refresh tokens rotate on every use.
  • MCP Apps: ui://vieneu/player.html, ui://vieneu/audiobook.html and ui://vieneu/capabilities.html (text/html;profile=mcp-app); media allowed from the storage origin only. The capabilities card sends an example to the chat with ui/message when the host offers it.
  • Prompts: what_can_vieneu_do, read_text, make_audiobook, find_voice, check_balance.