- Problem
- Translation and speech recognition models for Ugandan languages were
strong in evaluation but had no path to real users — no serving layer,
no pipeline, no way to know when they degraded.
- What I did
- Designed and deployed the production systems serving real-time translation
and speech inference, built end-to-end data pipelines covering collection,
preprocessing, validation, training, evaluation and deployment, and stood up
monitoring for system performance, reliability and usage.
- Outcome
- Both models run in production with high availability, with
evaluation frameworks using BLEU, chrF and COMET so performance is
benchmarked measurably rather than asserted.