2 Star 6 Fork 107

FORK/PaddleSpeech

forked from PaddlePaddle/PaddleSpeech 
加入 Gitee
与超过 1200万 开发者一起发现、参与优秀开源项目,私有仓库也完全免费 :)
免费加入
文件
克隆/下载
贡献代码
同步代码
取消
提示: 由于 Git 不支持空文件夾,创建文件夹后会生成空的 .keep 文件
Loading...
README

WaveFlow with LJSpeech

Dataset

Download and Extract

Download LJSpeech-1.1 from it's Official Website and extract it to ~/datasets. Then the dataset is in the directory ~/datasets/LJSpeech-1.1.

Get Started

Assume the path to the dataset is ~/datasets/LJSpeech-1.1. Assume the path to the Tacotron2 generated mels is ../tts0/output/test. Run the command below to

  1. source path.
  2. preprocess the dataset.
  3. train the model.
  4. synthesize wavs from mels.
./run.sh

You can choose a range of stages you want to run, or set stage equal to stop-stage to use only one stage, for example, running the following command will only preprocess the dataset.

./run.sh --stage 0 --stop-stage 0

Data Preprocessing

./local/preprocess.sh ${preprocess_path}

Model Training

./local/train.sh calls ${BIN_DIR}/train.py.

CUDA_VISIBLE_DEVICES=${gpus} ./local/train.sh ${preprocess_path} ${train_output_path}

The training script requires 4 command line arguments.

  1. --data is the path of the training dataset.
  2. --output is the path of the output directory.
  3. --ngpu is the number of gpus to use, if ngpu == 0, use cpu.

If you want distributed training, set a larger --ngpu (e.g. 4). Note that distributed training with cpu is not supported yet.

Synthesizing

./local/synthesize.sh calls ${BIN_DIR}/synthesize.py, which can synthesize waveform from mels.

CUDA_VISIBLE_DEVICES=${gpus} ./local/synthesize.sh ${input_mel_path} ${train_output_path} ${ckpt_name}

Synthesize waveform.

  1. We assume the --input is a directory containing several mel spectrograms(log magnitude) in .npy format.
  2. The output would be saved in the --output directory, containing several .wav files, each with the same name as the mel spectrogram does.
  3. --checkpoint_path should be the path of the parameter file (.pdparams) to load. Note that the extention name .pdparmas is not included here.
  4. --ngpu is the number of gpus to use, if ngpu == 0, use cpu.

Pretrained Model

Pretrained Model with residual channel equals 128 can be downloaded here:

马建仓 AI 助手
尝试更多
代码解读
代码找茬
代码优化
1
https://gitee.com/diven-fork/PaddleSpeech.git
git@gitee.com:diven-fork/PaddleSpeech.git
diven-fork
PaddleSpeech
PaddleSpeech
develop

搜索帮助