Research datasets

Open data from our research

Publicly available datasets from Marconi Lab research programs — ready for use in your own work.

CC01.0

Dataset of Crops part one

This dataset consists of five classes of data in the training set: cassava, sugarcane, maize, cashew, and coffee images. This is the first part, which consists of five out of seven classes used in crop classification projects; train data, weeds, …

Jan 2024
CC0 1.0

Data of Crops Part Two

This is the second part of the data "https://doi.org/10.7910/DVN/J0OS9R". It consists of two classes that remained from the training set's data: weeds and unknown. Plus, the validation and test data with all classes. Please, to use it, combine the first …

Jan 2024
CC BY 4.0

Multilingual Parallel Text Corpora for East African Languages

This is a partial multilingual parallel corpora of 5 East African languages. The dataset contains an English text corpus that has been translated into five East African languages: Acholi, Runyankore, Luganda, Lumasaba, and Swahili. (2023-12-05)

Dec 2023
CC BY 4.0

Coffee and Cashew Nut Dataset

The datasets presented in this work consist of high-resolution images of coffee and cashew plants acquired using Unmanned Aerial Equipment (UAV) equipment from small and large-scale farms across Uganda. Images range approximately between 10 MB and 12 MB in size, …

Nov 2023
CC0 1.0

Makerere Luganda Agricultural Text Data

The dataset consists of sentences in the Luganda language that solely pertain to the agricultural domain. These sentences cover a wide range of topics within agriculture, such as farming, animal breeding, crop cultivation, crop storage and yield, marketing of produce, …

May 2023
CC BY 4.0

Sentiment Tagged Parallel Corpus for Luganda and Swahili

This dataset contains 10,000 parallel sentiment-tagged sentences. English sentences were translated to both Luganda and Swahili. The translations were done by language experts and professional translators in collaboration with researchers at Makerere University. All sentences were tagged with a sentiment …

Mar 2023
CC BY 4.0

Kiswahili Monolingual Corpus

This dataset contains 100,000 Kiswahili sentences. We want to thank the team at the Makerere AI and Marconi Labs at Makerere University, TAVODET Youth Development (TYD) Innovation Incubator, Ai Kenya, Maseno University, United States International University-Africa (USIU-Africa), and Kabarak University …

Mar 2023
CC BY 4.0

Acoli Monolingual Corpus

Acoli is a very low-resourced language spoken in parts of Northern Uganda. This dataset contains 40,037 Acoli sentences. The sentences were collected and evaluated by Acoli linguists with the collaboration of teams at Marconi Research and Innovation Lab and Makerere …

Mar 2023
CC BY 4.0

Lumasaba Monolingual Corpus

Lumasaba sometimes known as Lugisu is a Bantu language spoken in the Eastern part of Uganda. This dataset contains a total of 39,999 sentences. The sentences are split into two separate files. One file contains 20,764 sentences from the Northern …

Mar 2023
CC BY 4.0

Luganda Monolingual Corpus

This dataset contains 100,000 Luganda sentences. Luganda is a Bantu language and is one of the major languages spoken in Uganda. This dataset was compiled by researchers at the Makerere AI and Data Science Research Lab and Marconi Research and …

Mar 2023
CC0 1.0

Makerere University Cassava Image Dataset

The dataset was created to provide an open-source and well-curated image dataset showing diseased and healthy cassava leaf images from Uganda. This will be used by data scientists, researchers, the wider machine learning community, and experts from other domains to …

Aug 2022
CC0 1.0

Makerere University Beans Image Dataset

This beans dataset was created to provide an open and accessible, well-labeled, sufficiently curated image dataset. This is to enable researchers to build various machine learning experiments to aid innovations that may include; bean crop disease diagnosis and spatial analysis. …

Jul 2022
Creative Commons Attribution 4.0 International

The Makerere Gendered Corpus: A Gendered English to Luganda Parallel Corpus

This English-Luganda parallel sentence corpus consists of gendered examples created by a team of researchers from Makerere AI Lab at Makerere University with a team of Luganda teachers, students and freelancers. The collaborative work which involves generating English sentences under …

Jan 2022