Instructions to use caesar-abrham/Chaka-ASR with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use caesar-abrham/Chaka-ASR with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="caesar-abrham/Chaka-ASR")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForCTC processor = AutoProcessor.from_pretrained("caesar-abrham/Chaka-ASR") model = AutoModelForCTC.from_pretrained("caesar-abrham/Chaka-ASR", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Chaka-ASR
Amharic automatic speech recognition
Chaka-ASR converts Amharic speech into text. The model is fine-tuned from badrex/Ethio-ASR-amharic using Leyu Amharic speech data, with WAXAL replay during training.
It is intended for transcription, voice interfaces, and speech applications in Ethiopian contexts, including environments where background noise may be present.
Model details
| Property | Details |
|---|---|
| Model ID | caesar-abrham/Chaka-ASR |
| Maintainer | caesar-abrham |
| Language | Amharic |
| Task | Automatic speech recognition |
| Architecture | Wav2Vec2-BERT with a CTC output head |
| Parameters | Approximately 606 million |
| Starting checkpoint | badrex/Ethio-ASR-amharic |
| Input | Mono audio at 16,000 Hz |
| Output | Amharic text |
| Weights | Safetensors, float32 |
Training overview
Chaka-ASR was adapted using Leyu Amharic speech datasets. Training also included sampled WAXAL examples as replay data. The training summary records 8,900 training steps and speaker-separated Leyu training, validation, and test splits.
The starting checkpoint is an Amharic model from the Ethio-ASR project, based on facebook/w2v-bert-2.0. Chaka-ASR builds on that work; it is not trained from scratch.
Evaluation
| Dataset | WER (%) | CER (%) |
|---|---|---|
| Leyu | 20.48 | 6.02 |
The reported Leyu result was measured on its Addis Ababa dialect test subset. It is a subset result rather than an aggregate score for the complete Leyu dataset. WER is word error rate and CER is character error rate; lower values indicate fewer errors.
These values come from the recorded evaluation of the released checkpoint. They are not a direct comparison against the starting model. The original evaluation's text normalization and decoding settings should be consulted before reproducing or comparing these scores.
Quick start
Install the example dependencies in a separate Python environment:
python -m pip install -r requirements-example.txt
Run the included example on a WAV or FLAC recording:
python examples/transcribe.py audio.wav
The example converts stereo audio to mono, resamples it to 16 kHz, and uses greedy CTC decoding. It uses an NVIDIA GPU when CUDA is available and otherwise runs on CPU. Audio is processed as a complete clip; this example does not implement live streaming or long-recording segmentation.
The checkpoint configuration records Transformers 5.16.1. The dependency file uses that version as its minimum; the example must still be checked in your target runtime before production use. Install a PyTorch build appropriate for your hardware if GPU inference is required.
Intended uses
- Amharic audio transcription.
- Voice assistants and conversational applications.
- Speech interfaces for customer service and accessibility.
- Research and evaluation of Amharic speech recognition.
Background noise
Chaka-ASR is intended for applications that may encounter background noise. Actual performance depends on microphone placement, speech clarity, reverberation, and the level and type of noise. The published Leyu result does not establish a separate noise robustness benchmark. Evaluate the model on representative recordings from your deployment environment.
Limitations
Recognition quality may vary with dialect, recording conditions, speaking style, and overlapping speech. Names, numbers, abbreviations, and mixed-language speech may require additional review. The model produces transcriptions; it does not verify identity, amounts, or factual content.
The reported error rates describe the evaluated subset and should not be treated as guaranteed performance for all Amharic recordings.
Attribution
Chaka-ASR acknowledges:
- Ethio-ASR, including the Amharic starting checkpoint.
- facebook/w2v-bert-2.0, the upstream architecture and pretrained model.
- Leyu Amharic speech datasets, used for adaptation.
- WAXAL, used as replay data during adaptation.
License
Chaka-ASR is released under the Creative Commons Attribution 4.0 International license (CC BY 4.0). You may share and adapt the model, including for commercial use, provided you give appropriate credit, link to the license, and indicate any changes.
Feedback
Report transcription issues and technical questions in the Community tab. Include the model revision and runtime versions. Share recordings only when you have permission and remove personal or confidential information.
Citation
@misc{chaka_asr_2026,
author = {caesar-abrham},
title = {Chaka-ASR: Amharic Automatic Speech Recognition},
year = {2026},
url = {https://huggingface.co/caesar-abrham/Chaka-ASR}
}
- Downloads last month
- 19