Image Document Classification Prediction based on SVM and gradient-boosting Algorithms

DOI:

https://doi.org/10.36371/port.2023.4.5

Authors

  • Ahmed Hussein Salman Iraqi Commission for Computers and Informatics, Informatics Institute of Postgraduate Studies, Baghdad, Iraq.
  • Waleed A, Mahmoud Al-Jawher College of Engineering, Uruk University, Baghdad, Iraq.

Image document classification is crucial in various domains, including healthcare, finance, and security. Automatically categorizing images into predefined classes can significantly improve data management and decision-making processes. For this research, we investigate the effectiveness of two machine learning algorithms, Support Vector Machines (SVM) and Gradient Boosting, for image document classification. First, we preprocess the image data by extracting relevant features, such as Image Embedding, to create a feature vector for each image. These features are essential for representing the content of the images accurately. Next, we apply SVM, a robust supervised learning algorithm, to train a classification model. SVM aims to Determine the optimal hyperplane for effectively distinguishing the images into different classes while maximizing the margin. Furthermore, we explore the Gradient Boosting algorithm, an ensemble learning method combining multiple weak learners to create a robust classifier. We experimented with different classification results with ten classes. We employ Multiple measures, including accuracy, precision, recall, F1-score, and ROC-AUC, are used to assess the performance of the SVM and Gradient Boosting models. The higher result of 0.964 for SVM compared with Adaboost is achieved. 0.853. 

Keywords:

Document Classification, Feature Extraction, Gradient Boosting algorithm, Support Vector Machines (SVM)

[1] Image document classification is crucial in various domains, including healthcare, finance, and security. Automatically categorizing images into predefined classes can significantly improve data management and decision-making processes. For this research, we investigate the effectiveness of two machine learning algorithms, Support Vector Machines (SVM) and Gradient Boosting, for image document classification. First, we preprocess the image data by extracting relevant features, such as Image Embedding, to create a feature vector for each image. These features are essential for representing the content of the images accurately. Next, we apply SVM, a robust supervised learning algorithm, to train a classification model. SVM aims to Determine the optimal hyperplane for effectively distinguishing the images into different classes while maximizing the margin. Furthermore, we explore the Gradient Boosting algorithm, an ensemble learning method combining multiple weak learners to create a robust classifier. We experimented with different classification results with ten classes. We employ Multiple measures, including accuracy, precision, recall, F1-score, and ROC-AUC, are used to assess the performance of the SVM and Gradient Boosting models. The higher result of 0.964 for SVM compared with Adaboost is achieved. 0.853

Salman, A. H. ., & Al-Jawher, W. A. M. . (2023). Image Document Classification Prediction based on SVM and gradient-boosting Algorithms. Journal Port Science Research, 6(4), 348–356. https://doi.org/10.36371/port.2023.4.5

Downloads

Download data is not yet available.