Chaka-ASR

Amharic automatic speech recognition

Chaka-ASR converts Amharic speech into text. The model is fine-tuned from badrex/Ethio-ASR-amharic using Leyu Amharic speech data, with WAXAL replay during training.

It is intended for transcription, voice interfaces, and speech applications in Ethiopian contexts, including environments where background noise may be present.

Model details

Property Details
Model ID caesar-abrham/Chaka-ASR
Maintainer caesar-abrham
Language Amharic
Task Automatic speech recognition
Architecture Wav2Vec2-BERT with a CTC output head
Parameters Approximately 606 million
Starting checkpoint badrex/Ethio-ASR-amharic
Input Mono audio at 16,000 Hz
Output Amharic text
Weights Safetensors, float32

Training overview

Chaka-ASR was adapted using Leyu Amharic speech datasets. Training also included sampled WAXAL examples as replay data. The training summary records 8,900 training steps and speaker-separated Leyu training, validation, and test splits.

The starting checkpoint is an Amharic model from the Ethio-ASR project, based on facebook/w2v-bert-2.0. Chaka-ASR builds on that work; it is not trained from scratch.

Evaluation

Dataset WER (%) CER (%)
Leyu 20.48 6.02

The reported Leyu result was measured on its Addis Ababa dialect test subset. It is a subset result rather than an aggregate score for the complete Leyu dataset. WER is word error rate and CER is character error rate; lower values indicate fewer errors.

These values come from the recorded evaluation of the released checkpoint. They are not a direct comparison against the starting model. The original evaluation's text normalization and decoding settings should be consulted before reproducing or comparing these scores.

Quick start

Install the example dependencies in a separate Python environment:

python -m pip install -r requirements-example.txt

Run the included example on a WAV or FLAC recording:

python examples/transcribe.py audio.wav

The example converts stereo audio to mono, resamples it to 16 kHz, and uses greedy CTC decoding. It uses an NVIDIA GPU when CUDA is available and otherwise runs on CPU. Audio is processed as a complete clip; this example does not implement live streaming or long-recording segmentation.

The checkpoint configuration records Transformers 5.16.1. The dependency file uses that version as its minimum; the example must still be checked in your target runtime before production use. Install a PyTorch build appropriate for your hardware if GPU inference is required.

Intended uses

  • Amharic audio transcription.
  • Voice assistants and conversational applications.
  • Speech interfaces for customer service and accessibility.
  • Research and evaluation of Amharic speech recognition.

Background noise

Chaka-ASR is intended for applications that may encounter background noise. Actual performance depends on microphone placement, speech clarity, reverberation, and the level and type of noise. The published Leyu result does not establish a separate noise robustness benchmark. Evaluate the model on representative recordings from your deployment environment.

Limitations

Recognition quality may vary with dialect, recording conditions, speaking style, and overlapping speech. Names, numbers, abbreviations, and mixed-language speech may require additional review. The model produces transcriptions; it does not verify identity, amounts, or factual content.

The reported error rates describe the evaluated subset and should not be treated as guaranteed performance for all Amharic recordings.

Attribution

Chaka-ASR acknowledges:

License

Chaka-ASR is released under the Creative Commons Attribution 4.0 International license (CC BY 4.0). You may share and adapt the model, including for commercial use, provided you give appropriate credit, link to the license, and indicate any changes.

Feedback

Report transcription issues and technical questions in the Community tab. Include the model revision and runtime versions. Share recordings only when you have permission and remove personal or confidential information.

Citation

@misc{chaka_asr_2026,
  author = {caesar-abrham},
  title = {Chaka-ASR: Amharic Automatic Speech Recognition},
  year = {2026},
  url = {https://huggingface.co/caesar-abrham/Chaka-ASR}
}
Downloads last month
19
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for caesar-abrham/Chaka-ASR

Finetuned
(1)
this model