TTS: Síntesi de veu

De wikijoan
La revisió el 21:09, 10 oct 2026 per Joan (discussió | contribucions) (Es crea la pàgina amb «__TOC__ =1a prova: F5-TTS, no funciona= (no acaba de funcionar, resultats inconsistents) *https://chatgpt.com/c/6ab042c5-4344-83eb-90bb-83cad23ab322 (m'he fet un l...».)
(dif) ← Versió més antiga | Versió actual (dif) | Versió més nova → (dif)
Salta a la navegació Salta a la cerca

1a prova: F5-TTS, no funciona

(no acaba de funcionar, resultats inconsistents)

(m'he fet un lio amb el XTTS-v2 i la instal·lació de Coqui TTS)

prompt: "vull fer clonació i síntesi de veu en català amb F5-TTS. Tinc una RTX5060 Ti 16GB."

Per català, però, hi ha un punt important: el model base de F5-TTS no és específicament un model català, així que la qualitat de pronunciació catalana dependrà molt del text, del model/checkpoint i de l'àudio de referència. Si vols qualitat alta i consistent en català, podem plantejar també fine-tuning.

mkdir -p F5-TTS
cd F5-TTS
python3 -m venv .venv
source .venv/bin/activate
python -m pip install --upgrade pip setuptools wheel

3. Instal·la PyTorch El repositori actual de F5-TTS documenta PyTorch amb CUDA 12.8, i PyTorch publica wheels específiques cu128.

Jo tinc CUDA 13.2, però no passa res

pip install torch==2.8.0+cu128 torchaudio==2.8.0+cu128 \
    --extra-index-url https://download.pytorch.org/whl/cu128

$ python3
import torch

print("PyTorch:", torch.__version__)
print("CUDA:", torch.version.cuda)
print("CUDA disponible:", torch.cuda.is_available())
print("GPU:", torch.cuda.get_device_name(0))
print("Compute capability:", torch.cuda.get_device_capability(0))

PyTorch: 2.8.0+cu128
CUDA: 12.8
CUDA disponible: True
GPU: NVIDIA GeForce RTX 5060 Ti
Compute capability: (12, 0)

4. Instal·la F5-TTS

git clone https://github.com/SWivid/F5-TTS.git

cd F5-TTS
pip install -e .

ara el meu directori de treball és ~/F5-TTS/F5-TTS. No confondre el primer amb el segon.

$ f5-tts_infer-cli --help
$ f5-tts_infer-gradio --help -> es descarrega el model

5. Primera prova: interfície web

$ f5-tts_infer-gradio

Normalment et donarà una adreça local del tipus:

Obre-la al navegador.

La interfície de F5-TTS permet fer inferència a partir d'una veu de referència i generar text nou. El projecte actual també disposa de funcionalitats d'inferència per chunks i altres modes.

però jo no estic davant de l'ordinador. I no puc fer:

perquè el port 7860 no està obert (m'he quedat aquí) (...)

Gravo una veu que representa que és de referència, i a partir d'aquesta veu podré fer síntesi.

$ arecord -f cd  gravacio.wav

scp gravacio.* joan@192.168.1.184:/home/joan/Baixades


$ f5-tts_infer-cli --model F5TTS_Base --ref_audio /home/joan/Baixades/gravacio.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood. The streets were almost empty, and a thick fog covered the old houses. As she walked towards the house she had recently inherited from her grandmother, she noticed a strange light shining from one of the upstairs windows. Emily stopped and stared at it for a moment. She was certain that the house had been empty for years, so she could not understand who could be inside." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova.wav --speed 0.06

es genera el fitxer /home/joan/Baixades/prova.wav

NOTA: recordem que en el portàtil tinc un punt de muntatge a /home/joan/projectes en el servidor.

$ find ~/.cache/gface -type d -iname "*F5*" 2>/dev/null
/home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS
/home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS/snapshots/84e5a410d9cead4de2f847e7c9369a6440bdfaca/F5TTS_v1_Base
/home/joan/.cache/huggingface/hub/models--SWivid--F5-TTS/snapshots/84e5a410d9cead4de2f847e7c9369a6440bdfaca/F5TTS_Base
/home/joan/.cache/huggingface/hub/.locks/models--SWivid--F5-TTS

$ f5-tts_infer-cli --model F5TTS_v1_Base --ref_audio /home/joan/Baixades/gravacio.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood. The streets were almost empty, and a thick fog covered the old houses. As she walked towards the house she had recently inherited from her grandmother, she noticed a strange light shining from one of the upstairs windows. Emily stopped and stared at it for a moment. She was certain that the house had been empty for years, so she could not understand who could be inside." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova2.wav --speed 0.06

$ f5-tts_infer-cli --model F5TTS_Base --ref_audio /home/joan/projectes/gravacio2.wav --ref_text "It was a cold and quiet evening when Emily arrived in the small village of Blackwood." --gen_text "Good morning. This is a simple test of generating an English speech from a text" --output_dir /home/joan/projectes --output_file prova.wav --speed 0.06

No m'acaba de sortir. Obtinc uns resultats en anglès bastant inconsistents, i en català no ho aconseguiré.


creat per Joan Quintana Compte, setembre 2026