Voice cloning
Clone any voice from a 3–8 second clip. No transcript, no fine-tuning. The clip is denoised and trimmed to ≤ 8 s automatically before cloning.
from vieneu import Vieneu
tts = Vieneu()
audio = tts.infer(
"Đây là giọng được nhân bản tức thì.",
ref_audio="my_voice.wav", # 3–8 s reference clip
denoise=True, # default; False if the clip is already clean
)
tts.save(audio, "cloned.wav")
ref_text on v3v1/v2 needed the exact transcript of the reference clip (ref_text). v3 Turbo extracts a speaker embedding and reference codes instead, so a transcript is not required.
Enrol once, reuse by name
tts.add_voice("Giọng của tôi", "my_voice.wav") # denoise + extract the profile once
audio = tts.infer("Câu này dùng giọng đã lưu.", voice="Giọng của tôi")
tts.save_voices() # persist to the default voices file
# tts.remove_voice("Giọng của tôi")
tts.add_voice("Giọng sạch", "already_clean.wav", denoise=False)
Enrolled voices work everywhere a preset does, including infer_batch, infer_stream and the conversation mode.
Denoise on its own
wav, sr = tts.denoise("noisy.wav", out_path="clean.wav") # 44.1 kHz mono
Reading style follows the reference
The style argument (tin_tuc, doc_truyen, …) is deprecated and ignored on v3 Turbo. The reading character is baked into the reference: pick a preset voice or a clip that already reads the way you want. Passing style still runs for backwards compatibility.
Emotion cues (experimental)
Inline tags work with cloned voices too: [cười] (chuckle), [thở dài] (sigh), [hắng giọng] (clear throat).
audio = tts.infer("Nghe hay quá đi [cười]. Để mình nói tiếp [hắng giọng].", voice="Giọng của tôi")
Backends
denoise, add_voice and cloning work on every backend, including the torch-free CPU/ONNX install (the pipeline runs on onnxruntime + soxr + kaldi-native-fbank). v3 Nano clones the same way; its three cloning graphs (~110 MB) download on first use.
Tips for a good reference
- 3–8 s of one speaker, no music, no second voice.
- Natural, continuous speech beats isolated words.
- Keep
denoise=Trueunless you cleaned the clip yourself. - Want a tighter match or a specific reading style? Fine-tune with LoRA on 10–30 minutes of audio.
Higher fidelity: v4 on the Cloud API
The proprietary VieNeu-TTS v4 reproduces a reference with near-original speaker similarity. It is not open source and is available only through the Cloud API.