Update org links to studio-dots-ai / dots-studio
Browse files
README.md
CHANGED
|
@@ -11,15 +11,15 @@ tags:
|
|
| 11 |
- flow-matching
|
| 12 |
- post-trained
|
| 13 |
library_name: dots_tts
|
| 14 |
-
base_model:
|
| 15 |
---
|
| 16 |
|
| 17 |
# dots.tts-soar
|
| 18 |
|
| 19 |
<p align="left">
|
| 20 |
-
<a href="https://github.com/
|
| 21 |
-
<a href="https://huggingface.co/spaces/
|
| 22 |
-
<a href="https://
|
| 23 |
</p>
|
| 24 |
|
| 25 |
**dots.tts** is a **2B-parameter fully continuous, end-to-end autoregressive (AR) text-to-speech system**. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE β no discrete codec tokens anywhere in the pipeline.
|
|
@@ -28,15 +28,15 @@ This repository hosts **`dots.tts-soar`** β the pretrained backbone further re
|
|
| 28 |
|
| 29 |
<table>
|
| 30 |
<tr>
|
| 31 |
-
<td align="left" valign="middle"><a href="https://huggingface.co/
|
| 32 |
<td>Pretrain (~1.5M h). Fine-tuning, full CFG / NFE control.</td>
|
| 33 |
</tr>
|
| 34 |
<tr>
|
| 35 |
-
<td align="left" valign="middle"><a href="https://huggingface.co/
|
| 36 |
<td>β <em>you are here</em> β + Self-corrective Alignment. <strong>Highest zero-shot fidelity and speaker similarity</strong>; also recommended for fine-tuning.</td>
|
| 37 |
</tr>
|
| 38 |
<tr>
|
| 39 |
-
<td align="left" valign="middle"><a href="https://huggingface.co/
|
| 40 |
<td>+ MeanFlow distillation. Few-step inference (NFE = 4), low latency.</td>
|
| 41 |
</tr>
|
| 42 |
</table>
|
|
@@ -52,8 +52,8 @@ conda create -n dots_tts python=3.10 -y
|
|
| 52 |
conda activate dots_tts
|
| 53 |
|
| 54 |
python -m pip install --upgrade pip
|
| 55 |
-
python -m pip install "git+https://github.com/
|
| 56 |
-
-c "https://raw.githubusercontent.com/
|
| 57 |
```
|
| 58 |
|
| 59 |
### CLI
|
|
@@ -61,7 +61,7 @@ python -m pip install "git+https://github.com/rednote-hilab/dots.tts.git" \
|
|
| 61 |
```bash
|
| 62 |
# Continuation voice cloning (reference audio + transcript) β recommended
|
| 63 |
dots.tts \
|
| 64 |
-
--model-name-or-path
|
| 65 |
--text "Hello, this is a zero-shot voice cloning demonstration." \
|
| 66 |
--prompt-audio /path/to/reference.wav \
|
| 67 |
--prompt-text "The exact transcript of the reference audio." \
|
|
@@ -75,7 +75,7 @@ from dots_tts.runtime import DotsTtsRuntime
|
|
| 75 |
import soundfile as sf
|
| 76 |
|
| 77 |
runtime = DotsTtsRuntime.from_pretrained(
|
| 78 |
-
"
|
| 79 |
precision="bfloat16",
|
| 80 |
)
|
| 81 |
|
|
@@ -99,7 +99,7 @@ sf.write("output.wav", result["audio"].float().cpu().squeeze().numpy(), result["
|
|
| 99 |
|
| 100 |
### Fine-tuning
|
| 101 |
|
| 102 |
-
Both `dots.tts-base` and `dots.tts-soar` are valid fine-tuning starting points. Pick `dots.tts-soar` when you want to inherit its tightened text/timbre alignment on top of the pretrained backbone. See the [training script](https://github.com/
|
| 103 |
|
| 104 |
```bash
|
| 105 |
accelerate launch scripts/train_dots_tts.py --config configs/dots_tts.yaml
|
|
@@ -153,7 +153,7 @@ A frozen **AudioVAE** encodes 48 kHz mono waveform into a continuous latent and
|
|
| 153 |
|
| 154 |
On head-to-head judging vs. `gpt-4o-mini-tts`, `dots.tts-soar` posts **65.7% on Syntactic Complexity β above every closed-source system listed**, while keeping competitive Emotions / Questions scores.
|
| 155 |
|
| 156 |
-
See the [project README](https://github.com/
|
| 157 |
|
| 158 |
---
|
| 159 |
|
|
|
|
| 11 |
- flow-matching
|
| 12 |
- post-trained
|
| 13 |
library_name: dots_tts
|
| 14 |
+
base_model: dots-studio/dots.tts-base
|
| 15 |
---
|
| 16 |
|
| 17 |
# dots.tts-soar
|
| 18 |
|
| 19 |
<p align="left">
|
| 20 |
+
<a href="https://github.com/studio-dots-ai/dots.tts"><img src="https://img.shields.io/badge/GitHub-studio--dots--ai%2Fdots.tts-blue?logo=github" alt="GitHub"></a>
|
| 21 |
+
<a href="https://huggingface.co/spaces/dots-studio/dots.tts"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Spaces-Playground-orange" alt="Playground"></a>
|
| 22 |
+
<a href="https://studio-dots-ai.github.io/dots.tts-demo/"><img src="https://img.shields.io/badge/Demo%20Page-Live-red" alt="Demo Page"></a>
|
| 23 |
</p>
|
| 24 |
|
| 25 |
**dots.tts** is a **2B-parameter fully continuous, end-to-end autoregressive (AR) text-to-speech system**. The backbone pairs a semantic encoder, an LLM, and an autoregressive flow-matching acoustic head over a 48 kHz AudioVAE β no discrete codec tokens anywhere in the pipeline.
|
|
|
|
| 28 |
|
| 29 |
<table>
|
| 30 |
<tr>
|
| 31 |
+
<td align="left" valign="middle"><a href="https://huggingface.co/dots-studio/dots.tts-base"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-dots.tts--base-yellow" alt="dots.tts-base"></a></td>
|
| 32 |
<td>Pretrain (~1.5M h). Fine-tuning, full CFG / NFE control.</td>
|
| 33 |
</tr>
|
| 34 |
<tr>
|
| 35 |
+
<td align="left" valign="middle"><a href="https://huggingface.co/dots-studio/dots.tts-soar"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-dots.tts--soar-yellow" alt="dots.tts-soar"></a></td>
|
| 36 |
<td>β <em>you are here</em> β + Self-corrective Alignment. <strong>Highest zero-shot fidelity and speaker similarity</strong>; also recommended for fine-tuning.</td>
|
| 37 |
</tr>
|
| 38 |
<tr>
|
| 39 |
+
<td align="left" valign="middle"><a href="https://huggingface.co/dots-studio/dots.tts-mf"><img src="https://img.shields.io/badge/%F0%9F%A4%97%20Hugging%20Face-dots.tts--mf-yellow" alt="dots.tts-mf"></a></td>
|
| 40 |
<td>+ MeanFlow distillation. Few-step inference (NFE = 4), low latency.</td>
|
| 41 |
</tr>
|
| 42 |
</table>
|
|
|
|
| 52 |
conda activate dots_tts
|
| 53 |
|
| 54 |
python -m pip install --upgrade pip
|
| 55 |
+
python -m pip install "git+https://github.com/studio-dots-ai/dots.tts.git" \
|
| 56 |
+
-c "https://raw.githubusercontent.com/studio-dots-ai/dots.tts/main/constraints/recommended.txt"
|
| 57 |
```
|
| 58 |
|
| 59 |
### CLI
|
|
|
|
| 61 |
```bash
|
| 62 |
# Continuation voice cloning (reference audio + transcript) β recommended
|
| 63 |
dots.tts \
|
| 64 |
+
--model-name-or-path dots-studio/dots.tts-soar \
|
| 65 |
--text "Hello, this is a zero-shot voice cloning demonstration." \
|
| 66 |
--prompt-audio /path/to/reference.wav \
|
| 67 |
--prompt-text "The exact transcript of the reference audio." \
|
|
|
|
| 75 |
import soundfile as sf
|
| 76 |
|
| 77 |
runtime = DotsTtsRuntime.from_pretrained(
|
| 78 |
+
"dots-studio/dots.tts-soar",
|
| 79 |
precision="bfloat16",
|
| 80 |
)
|
| 81 |
|
|
|
|
| 99 |
|
| 100 |
### Fine-tuning
|
| 101 |
|
| 102 |
+
Both `dots.tts-base` and `dots.tts-soar` are valid fine-tuning starting points. Pick `dots.tts-soar` when you want to inherit its tightened text/timbre alignment on top of the pretrained backbone. See the [training script](https://github.com/studio-dots-ai/dots.tts/blob/main/scripts/train_dots_tts.py) and [smoke config](https://github.com/studio-dots-ai/dots.tts/blob/main/configs/dots_tts.yaml) in the source repository:
|
| 103 |
|
| 104 |
```bash
|
| 105 |
accelerate launch scripts/train_dots_tts.py --config configs/dots_tts.yaml
|
|
|
|
| 153 |
|
| 154 |
On head-to-head judging vs. `gpt-4o-mini-tts`, `dots.tts-soar` posts **65.7% on Syntactic Complexity β above every closed-source system listed**, while keeping competitive Emotions / Questions scores.
|
| 155 |
|
| 156 |
+
See the [project README](https://github.com/studio-dots-ai/dots.tts#-performance) for full benchmark tables.
|
| 157 |
|
| 158 |
---
|
| 159 |
|