Model | GitHub
Flow-matching based Japanese TTS model (approximately 766M parameters). Generates speech from text, optional reference audio, and optional style caption.
- Reference audio: Optional. One or more clips can be concatenated in the displayed order, up to the model's 120-second limit.
- Caption: Optional style prompt for emotion, tone, speaking style, or acoustic scene.
- Duration: By default, v4-Small predicts the output duration automatically. Use Duration Scale for small adjustments or Seconds for exact manual control.
For longer references, multiple clean, shorter clips from the same speaker are recommended. This matches v4-Small training. A single uninterrupted long recording is accepted but has not been evaluated.