TTS: Síntesi de veu
Contingut
1a prova: F5-TTS, no funciona
(no acaba de funcionar, resultats inconsistents)
(m'he fet un lio amb el XTTS-v2 i la instal·lació de Coqui TTS)
prompt: "vull fer clonació i síntesi de veu en català amb F5-TTS. Tinc una RTX5060 Ti 16GB."
Per català, però, hi ha un punt important: el model base de F5-TTS no és específicament un model català, així que la qualitat de pronunciació catalana dependrà molt del text, del model/checkpoint i de l'àudio de referència. Si vols qualitat alta i consistent en català, podem plantejar també fine-tuning.
mkdir -p F5-TTS cd F5-TTS python3 -m venv .venv source .venv/bin/activate python -m pip install --upgrade pip setuptools wheel
3. Instal·la PyTorch El repositori actual de F5-TTS documenta PyTorch amb CUDA 12.8, i PyTorch publica wheels específiques cu128.
Jo tinc CUDA 13.2, però no passa res
pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 \
--extra-index-url https://download.pytorch.org/whl/cu128
$ python3
import torch
print("PyTorch:", torch.__version__)
print("CUDA:", torch.version.cuda)
print("CUDA disponible:", torch.cuda.is_available())
print("GPU:", torch.cuda.get_device_name(0))
print("Compute capability:", torch.cuda.get_device_capability(0))
PyTorch: 2.8.0+cu128
CUDA: 12.8
CUDA disponible: True
GPU: NVIDIA GeForce RTX 5060 Ti
Compute capability: (12, 0)
4. Instal·la F5-TTS
git clone https://github.com/SWivid/F5-TTS.git cd F5-TTS pip install -e .
ara el meu directori de treball és ~/F5-TTS/F5-TTS. No confondre el primer amb el segon.
$ f5-tts_infer-cli --help $ f5-tts_infer-gradio --help -> es descarrega el model
5. Primera prova: interfície web
$ f5-tts_infer-gradio
Normalment et donarà una adreça local del tipus:
Obre-la al navegador.
La interfície de F5-TTS permet fer inferència a partir d'una veu de referència i generar text nou. El projecte actual també disposa de funcionalitats d'inferència per chunks i altres modes.
però jo no estic davant de l'ordinador. I no puc fer:
perquè el port 7860 no està obert (m'he quedat aquí) (...)
Gravo una veu que representa que és de referència, i a partir d'aquesta veu podré fer síntesi.
$ arecord -f cd gravacio.wav scp gravacio.* joan@192.168.1.184:/home/joan/Baixades $ f5-tts_infer-cli --model F5TTS_Base --ref_audio /home/joan/Baixades/gravacio.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood. The streets were almost empty, and a thick fog covered the old houses. As she walked towards the house she had recently inherited from her grandmother, she noticed a strange light shining from one of the upstairs windows. Emily stopped and stared at it for a moment. She was certain that the house had been empty for years, so she could not understand who could be inside." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova.wav --speed 0.06
es genera el fitxer /home/joan/Baixades/prova.wav
NOTA: recordem que en el portàtil tinc un punt de muntatge a /home/joan/projectes en el servidor.
$ find ~/.cache/gface -type d -iname "*F5*" 2>/dev/null /home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS /home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS/snapshots/84e5a410d9cead4de2f847e7c9369a6440bdfaca/F5TTS_v1_Base /home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS/snapshots/84e5a410d9cead4de2f847e7c9369a6440bdfaca/F5TTS_Base /home/joan/.cache/huggingface/hub/.locks/models--SWivid--F5-TTS $ f5-tts_infer-cli --model F5TTS_v1_Base --ref_audio /home/joan/Baixades/gravacio.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood. The streets were almost empty, and a thick fog covered the old houses. As she walked towards the house she had recently inherited from her grandmother, she noticed a strange light shining from one of the upstairs windows. Emily stopped and stared at it for a moment. She was certain that the house had been empty for years, so she could not understand who could be inside." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova2.wav --speed 0.06 $ f5-tts_infer-cli --model F5TTS_Base --ref_audio /home/joan/projectes/gravacio2.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova.wav --speed 0.06
No m'acaba de sortir. Obtinc uns resultats en anglès bastant inconsistents, i en català no ho aconseguiré.
creat per Joan Quintana Compte, setembre 2026