Skip to main content

Voice cloning

Clone any voice from a 3–8 second clip. No transcript, no fine-tuning. The clip is denoised and trimmed to ≤ 8 s automatically before cloning.

from vieneu import Vieneu

tts = Vieneu()
audio = tts.infer(
"Đây là giọng được nhân bản tức thì.",
ref_audio="my_voice.wav", # 3–8 s reference clip
denoise=True, # default; False if the clip is already clean
)
tts.save(audio, "cloned.wav")
No ref_text on v3

v1/v2 needed the exact transcript of the reference clip (ref_text). v3 Turbo extracts a speaker embedding and reference codes instead, so a transcript is not required.

Enrol once, reuse by name​

tts.add_voice("Giọng của tôi", "my_voice.wav")           # denoise + extract the profile once
audio = tts.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")

tts.save_voices() # persist to the default voices file
# tts.remove_voice("Giọng của tôi")

tts.add_voice("Giọng sạch", "already_clean.wav", denoise=False)

Enrolled voices work everywhere a preset does, including infer_batch, infer_stream and the conversation mode.

Denoise on its own​

wav, sr = tts.denoise("noisy.wav", out_path="clean.wav")   # 44.1 kHz mono

Reading style follows the reference​

The style argument (tin_tuc, doc_truyen, …) is deprecated and ignored on v3 Turbo. The reading character is baked into the reference: pick a preset voice or a clip that already reads the way you want. Passing style still runs for backwards compatibility.

Emotion cues (experimental)​

Inline tags work with cloned voices too: [cười] (chuckle), [thở dài] (sigh), [hắng giọng] (clear throat).

audio = tts.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Giọng của tôi")

Backends​

denoise, add_voice and cloning work on every backend, including the torch-free CPU/ONNX install (the pipeline runs on onnxruntime + soxr + kaldi-native-fbank). v3 Nano clones the same way; its three cloning graphs (~110 MB) download on first use.

Tips for a good reference​

  • 3–8 s of one speaker, no music, no second voice.
  • Natural, continuous speech beats isolated words.
  • Keep denoise=True unless you cleaned the clip yourself.
  • Want a tighter match or a specific reading style? Fine-tune with LoRA on 10–30 minutes of audio.

Higher fidelity: v4 on the Cloud API​

The proprietary VieNeu-TTS v4 reproduces a reference with near-original speaker similarity. It is not open source and is available only through the Cloud API.