AI & ML interests
voice-conversion speech-separation speech-enhancement speech-translation speech-synthesis speech-recognition spoken-language-understanding
Recent Activity
Articles
TheESPnetLeaderBoard
ESPnet Leaderboard
Forced alignment
When each line was said, and how sure the model is
POWSM-CTC
What you said in phones, IPA, from a phonetic model
MELD speech emotion recognition demo
Run speech model demos with a web interface
Speaker verification
Are these two recordings the same speaker?
OWSM-CTC v4
151 languages in, 25 translation targets out, fast
Universal speech enhancement
One model for noise, reverb, any mics, any rate
LJSpeech VITS
English text to speech with a VITS model
OWSM v4
151 languages in, 25 translation targets out, promptable
MELD speech emotion recognition demo
Launch a web demo for speech models with Gradio
OWSM V4 Demo
This is a demo for OWSM-V4 CTC and medium model.
Voice Assistant Demo
SingingSDS
Generate text with a customizable interface
Svs
Generate singing voice from lyrics, duration, and pitch
TTS
Greet someone by name
S2st
Translate Spanish speech to English