When I first had the opportunity to join Leyu, I spent time thinking about whether it was the right next step for me. I wanted to continue growing in my career, take on bigger responsibilities, challenge myself, and reach my full potential as a project manager by leading complex projects from planning to completion.
What made the opportunity particularly interesting was the problem Leyu was working on. At the intersection of language, technology, and AI, language data is foundational to technologies such as Automatic Speech Recognition, text-to-speech, translation, and conversational systems. Yet many African languages remain significantly underrepresented in the digital resources needed to develop these technologies.
Leyu works to address this gap by building high-quality language datasets for low-resource languages. That meant the work was not simply about collecting recordings. It was about building language resources that could better represent how people actually speak and creating the processes and infrastructure needed to collect that data at scale.
One of the experiences that shaped that growth was an Amharic language data-collection initiative that involved contributors across five cities, multiple rounds of recruitment, training, constant coordination across teams, and an initial target of 500 hours of quality dataset. At some point, my team and I started checking that number constantly. We would look at the dashboard, share updates, compare progress, discuss what needed to happen next, and then somehow come back to the same question: How do we get to 500 hours? We even talked about how we would celebrate once we reached it. It became a regular part of our weekly updates as we watched the number grow and tried to imagine what the finish line would look like. And when we finally exceeded the number, we celebrated—not once, not twice, but three times. Somehow, 550.22 hours deserved three celebrations.
On paper, the target looked straightforward. In practice, it became a lesson in what it actually takes to implement a large-scale high quality language data set collection.
The initiative covered five dialect areas; Addis Ababa, Wollo, Shewa, Gonder, and Gojam and involved contributors, focal persons, linguistic reviewers, call-center agents, and different teams working together to keep the collection moving. My days involved coordinating across these groups, following up on progress, reviewing dashboards, leading recruitment, training, meetings and responding to whatever unexpected issue came up. With contributors spread across areas, keeping everyone aligned became a project of its own, requiring constant communication, monitoring, and adjustment.
One of the earliest lessons was that literacy did not automatically translate into readiness to collect language data. We saw lower than expected conversion because of differences in interest, digital skills, financial constraints, and linguistic fit. Training brought similar challenges: fewer participants than expected progressed from training session and assessment completion. These experiences changed how I approached planning. It is easy to build a process around the numbers you expect but implementation requires you to build around what people actually need, can do, and are willing to participate in.
Recruitment was only one part of the process. Once contributors joined, we had to manage documentation, training, task assignment, data monitoring, and ongoing communication. Telegram groups became an important part of that work, giving us a place to share updates, answer questions, address concerns, and keep contributors informed.
I remember one day, a contributor asked a question in the group, and because the team was dealing with several things at once, the question was not answered quickly. The contributor responded with an Amharic saying: "ቄሱም ዝም መጽሃፉም ዝም"—essentially, the priest is silent and the book is silent too. We laughed when we saw it, but the message was clear. Being busy was not a good enough reason for someone waiting for an answer. After that, we made our communication more structured, sharing key updates and ensuring questions were addressed within 24 hours. It was a small change, but it reflected something we kept learning: sometimes improving implementation does not require a major system. It requires noticing where the process is breaking down and fixing it.
As the data collection grew, another question became increasingly important: how do we scale this? Initially, we were trying to expand within the operating model we already had, but it became clear that simply asking the existing system to do more would not be enough. We began bringing university students into the contributor pool, which became one of the more important lessons for us about scalability. Scaling is not always about doing more of the same thing. Sometimes, it requires changing the strategy to support implementation.
Much of this thinking happened through the team's brainstorming sessions. At Leyu, we often filled the conference glass walls and doors of our conference rooms with ideas, breaking down what was slowing the process, exploring possible solutions, and thinking through where they might fail before putting them into practice. That approach taught me to look beyond immediate problems and think about how to anticipate challenges and build mitigation into the process. I also realized that I enjoy taking something complex and breaking it into smaller, workable parts.
For a while, 500 hours felt far away. Then the numbers started moving. We tracked progress across cities, followed up with contributors, adjusted our approach, and kept finding ways to move the collection forward. Eventually, we reached the target and went beyond it, collecting 550.22 hours of data with 619 contributors across the five dialect areas. For the team, that number represented months of recruitment, training, coordination, collaboration, problem-solving, and constant adaptation coming together. It also showed me how large-scale implementation feels when you are responsible for making it happen.
The impact of the initiative became even more visible once the dataset was shared. People started commenting about the language and the different ways Amharic is spoken. One person wrote, "I appreciate y'all for making it free." Another asked, "Please do Tigrigna as well." A developer commented "Now it is my turn". There were also people expressing how long they had been thinking about something like this but had not been able to make it happen themselves. Some even expressed surprise that the dataset was openly available. One comment in particular that gave me an aha moment: "ኦራ እዋ እና ክላም በAI ሊታወቁ ነው፡፡ እንዳሻኝ ማውራቴ ነው እግዲህ"—a playful realization that local nuances and dialects could actually be recognized by AI. It made me pause and think about what we were really building: not just a dataset, but a way for the way people naturally speak to be represented in AI.
It made me see that we were building more than a technical dataset; we were creating a way for people to see their language, dialect, and everyday speech represented in AI. The response also opened conversations that led to additional Afaan Oromo data-collection work, highlighting the broader gap in African language data and the opportunity to keep building where those gaps exist.
This initiative was Leyu's first experience collecting language data, and as the work progressed, we improved our own platform to support the process. We continuously improved the workflows, processes, and technical infrastructure needed to collect, manage, and validate language data at scale. Those improvements have since grown into the Leyu Data Collection Platform, including an open-source version that others can use for their own language-data collection efforts. Future initiatives can build on what we developed rather than starting from scratch. In that sense, the initiative produced more than a dataset; it also helped create infrastructure for future language-data work. You can explore the Leyu Data Collection Platform on GitHub.
Working on Leyu strengthened my understanding of what it takes to lead a complex, multi-stakeholder initiative. The 550.22 hour target was an important milestone, but the experience showed me that successful delivery depends on everything happening behind the number keeping people, processes, quality, technology, and timelines moving together while adapting to what the work demands.
That work has now resulted in an open Amharic dataset, available on Hugging Face for researchers, developers, and others to use and build with. It also demonstrates what the Leyu platform can make possible, enabling communities to contribute their languages and dialects, creating more opportunities for those voices to be represented in AI, and making the resulting data available for others to build on. The lessons from this work will continue to shape how the platform evolves and how future datasets can be collected across more languages and communities. And this is not where the story ends. There is more to come, and your feedback whether you are contributing data, using the platform, exploring the dataset, or building with it can help shape what comes next and contribute to a stronger ecosystem around African language data.

