97 7 84

Yanis L

Pendrokar

AI & ML interests

STT/STS/TTS you know, something that is solveable

Recent Activity

updated a dataset 40 minutes ago

Pendrokar/TTS_Arena

liked a model about 3 hours ago

hexgrad/Kokoro-82M

new activity 3 days ago

Pendrokar/open_tts_tracker:merging with ttsds/datasets

View all activity

Organizations

Posts 4

Post

386

TTS: Sorry, I just cannot get the hype behind F5 TTS. It has now gathered a thousand votes in the TTS Arena fork and **has remained in #8 spot** against the _mostly_ Open TTS adversaries.

The voice sample used is the same as XTTS. F5 has so far been unstable, being unemotional/monotone/depressed and mispronouncing words (_awestruck_).

If you have suggestions please give feedback in the following thread:
mrfakename/E2-F5-TTS#32

Post

924

Added @amphion MaskGCT & @hexgrad StyleTTS fine tuned model by the name of kokoro to the forked TTS Arena Space. If things keep up from what is seen in the preliminary results, then these two may end up in the TOP 5 of all TTS models. 🤞️🍀️

Pendrokar/TTS-Spaces-Arena
Svngoku/maskgct-audio-lab
hexgrad/Kokoro-TTS

I chose @Svngoku 's forked HF space over amphion's due to the overly high ZeroGPU duration demand on the latter. 300s!

amphion/maskgct

Had to remove @mrfakename 's MetaVoice-1B Space from the available models as that space has been down for quite some time. 🤕️

mrfakename/MetaVoice-1B-v0.1

I'm close to syncing the code to the original Arena's code structure. Then I'd like to use ASR in order to validate and create synthetic public datasets from the generated samples. And then make the Arena multilingual, which will surely attract quite the crowd!

View all posts