Data Augmentation For Sorani Kurdish News Headline Classification Using Back-Translation And Deep Learning Model

With the increase in the volume of news articles and headlines being generated, it is becoming more difficult for individuals to keep up with the latest developments and find relevant news articles in the Kurdish language. To address this issue, this paper proposes a novel data augmentation approach...

Full description

Saved in:
Bibliographic Details
Main Author: Soran Badawi
Format: Article
Language:English
Published: Sulaimani Polytechnic University 2023-06-01
Series:Kurdistan Journal of Applied Research
Subjects:
Online Access:https://kjar.spu.edu.iq/index.php/kjar/article/view/852
Tags: Add Tag
No Tags, Be the first to tag this record!
Description
Summary:With the increase in the volume of news articles and headlines being generated, it is becoming more difficult for individuals to keep up with the latest developments and find relevant news articles in the Kurdish language. To address this issue, this paper proposes a novel data augmentation approach for improving the performance of Kurdish news headline classification using back-translation and a proposed deep learning Bidirectional Long Short-Term Memory (BiLSTM) model. The approach involves generating synthetic training data by translating Kurdish headlines into a target language in this context English language and back-translating them to the Kurdish language, resulting in an augmented dataset. The proposed BiLSTM model is trained on the augmented data and compared with baseline models SVM (Support-Vector-Machines) and Naïve Bayes an trained on the original data. The experimental results demonstrate that the proposed BiLSTM model outperforms the baseline model and other existing models, achieving state-of-the-art performance on the Kurdish news headline classification task. The findings suggest that the combination of back-translation and a proposed BiLSTM model is a promising approach for data augmentation in low-resource languages, contributing to the advancement of natural language processing in under-resourced languages. Moreover, having a Kurdish news headline classification model can improve access to news and information for Kurdish speakers. With the classification model, they can easily and quickly search for news articles that interest them based on their preferred categories, such as politics, sports, or entertainment.
ISSN:2411-7684
2411-7706