Fine-tuned whisper-large-v3 model for speech recognition in Macedonian

Authors:

Dejan Porjazovski
Ilina Jakimovska
Ordan Chukaliev
Nikola Stikov

This collaboration is part of the activities of the Center for Advanced Interdisciplinary Research (CAIR) at UKIM.

Data used for training

In training of the model, we used the following data sources:

Digital Archive for Ethnological and Anthropological Resources (DAEAR) at the Institutе of Ethnology and Anthropology, PMF, UKIM.
Audio version of the international journal "EthnoAnthropoZoom" at the Institutе of Ethnology and Anthropology, PMF, UKIM.
The podcast "Обични луѓе" by Ilina Jakimovska.
The scientific videos from the series "Наука за деца", foundation KANTAROT.
Macedonian version of the Mozilla Common Voice (version 18).

Usage

from speechbrain.inference.interfaces import foreign_class
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
asr_classifier = foreign_class(source="Macedonian-ASR/whisper-large-v3-macedonian-asr", pymodule_file="custom_interface.py", classname="ASR")
asr_classifier = asr_classifier.to(device)
predictions = asr_classifier.classify_file("audio_file.wav", device)
print(predictions)

Macedonian-ASR
/

whisper-large-v3-macedonian-asr

Fine-tuned whisper-large-v3 model for speech recognition in Macedonian

Data used for training

Usage

Model tree for Macedonian-ASR/whisper-large-v3-macedonian-asr

Spaces using Macedonian-ASR/whisper-large-v3-macedonian-asr 2