Open AccessJournal ArticleDOI

An up-to-date comparison of state-of-the-art classification algorithms

- 01 Oct 2017 -

- Vol. 82, Iss: 82, pp 128-150

Chats0

TLDR

It is found that Stochastic Gradient Boosting Trees (GBDT) matches or exceeds the prediction performance of Support Vector Machines and Random Forests, while being the fastest algorithm in terms of prediction efficiency.

Abstract:

Up-to-date report on the accuracy and efficiency of state-of-the-art classifiers.We compare the accuracy of 11 classification algorithms pairwise and groupwise.We examine separately the training, parameter-tuning, and testing time.GBDT and Random Forests yield highest accuracy, outperforming SVM.GBDT is the fastest in testing, Naive Bayes the fastest in training. Current benchmark reports of classification algorithms generally concern common classifiers and their variants but do not include many algorithms that have been introduced in recent years. Moreover, important properties such as the dependency on number of classes and features and CPU running time are typically not examined. In this paper, we carry out a comparative empirical study on both established classifiers and more recently proposed ones on 71 data sets originating from different domains, publicly available at UCI and KEEL repositories. The list of 11 algorithms studied includes Extreme Learning Machine (ELM), Sparse Representation based Classification (SRC), and Deep Learning (DL), which have not been thoroughly investigated in existing comparative studies. It is found that Stochastic Gradient Boosting Trees (GBDT) matches or exceeds the prediction performance of Support Vector Machines (SVM) and Random Forests (RF), while being the fastest algorithm in terms of prediction efficiency. ELM also yields good accuracy results, ranking in the top-5, alongside GBDT, RF, SVM, and C4.5 but this performance varies widely across all data sets. Unsurprisingly, top accuracy performers have average or slow training time efficiency. DL is the worst performer in terms of accuracy but second fastest in prediction efficiency. SRC shows good accuracy performance but it is the slowest classifier in both training and testing.

Fig. 20. Box plots for AUC (raw values).

Table 15 Training time (in seconds) efficiency results for different classification algorithms.

Fig. 19. The number of data sets on which each classifier achieves the best accuracy, grouped by the number of features.

Fig. 18. The number of data sets on which each classifier achieves the best accuracy, grouped by the number of classes.

Fig. 22. Box plots for AUC (mean ranks).

Table 6 Accuracy results for different classification algorithms on 71 data sets.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

Impact-slip experiments and systematic study of coal gangue “category” recognition technology Part I: Impact-slip experiments between coal gangue mixture and top coal caving hydraulic support and the study of coal gangue “category” recognition technology

Yang Yang, +2 more

- 01 Nov 2021 -

Powder Technology

TL;DR: Research show that coal gangue “category” random recognition accuracy can reach to 0.935, which proves the effectiveness of “ category” by IoT system and impact-slip contact characteristics, which provides the theoretical basis for further classification recognition effect improving research.

...read moreread less

Proceedings ArticleDOI

Feature selection and resampling in class imbalance learning: Which comes first? An empirical study in the biological domain

Chongsheng Zhang, +2 more

TL;DR: There is no constant winner between the two pipelines, practitioners should test both pipelines in order to derive the best classification model for imbalance learning, in particular, the resampling before feature selection pipeline should not be neglected; but it is shown that, the feature selection before resamplings pipeline outperforms the other in more cases than not.

...read moreread less

Journal ArticleDOI

User preference modeling based on meta paths and diversity regularization in heterogeneous information networks

Hongzhi Liu, +4 more

- 01 Oct 2019 -

Knowledge Based Systems

TL;DR: A new diversity measure is presented and used to encourage diversity among meta paths to improve recommendation performance and compare with traditional collaborative filtering and state-of-the-art HIN-based recommendation methods.

...read moreread less

Journal ArticleDOI

A machine learning model to assess the ecosystem response to water policy measures in the Tagus River Basin (Spain).

Carlotta Valerio, +3 more

- 01 Jan 2021 -

Science of The Total Environment

TL;DR: Coupling more restrictive nutrient thresholds with measures that improve the riparian habitat yields up to 85% of water bodies with biological indices in good status, thus proving to be a key approach to restore the status of the ecosystem.

...read moreread less

Journal ArticleDOI

Fast kernel extreme learning machine for ordinal regression

Yong Shi, +4 more

- 01 Aug 2019 -

Knowledge Based Systems

TL;DR: A new KELM model for ordinal regression is proposed by exploiting a quadratic cost-sensitive encoding scheme and a fast algorithm is designed based on the low rank approximation to make the training process more efficient in the big data scenario.

...read moreread less

Collapse

References

PDF

Open Access

More filters

Journal ArticleDOI

Random Forests

Leo Breiman

TL;DR: Internal estimates monitor error, strength, and correlation and these are used to show the response to increasing the number of features used in the forest, and are also applicable to regression.

...read moreread less

Proceedings Article

ImageNet Classification with Deep Convolutional Neural Networks

Alex Krizhevsky, +2 more

TL;DR: The state-of-the-art performance of CNNs was achieved by Deep Convolutional Neural Networks (DCNNs) as discussed by the authors, which consists of five convolutional layers, some of which are followed by max-pooling layers, and three fully-connected layers with a final 1000-way softmax.

...read moreread less

Journal ArticleDOI

LIBSVM: A library for support vector machines

Chih-Chung Chang, +1 more

- 06 May 2011 -

ACM Transactions on Intelligent Systems ...

TL;DR: Issues such as solving SVM optimization problems theoretical convergence multiclass classification probability estimates and parameter selection are discussed in detail.

...read moreread less

Journal ArticleDOI

Support-Vector Networks

Corinna Cortes, +1 more

- 15 Sep 1995 -

Machine Learning

TL;DR: High generalization ability of support-vector networks utilizing polynomial input transformations is demonstrated and the performance of the support- vector network is compared to various classical learning algorithms that all took part in a benchmark study of Optical Character Recognition.

...read moreread less

Book

C4.5: Programs for Machine Learning

J. Ross Quinlan

TL;DR: A complete guide to the C4.5 system as implemented in C for the UNIX environment, which starts from simple core learning methods and shows how they can be elaborated and extended to deal with typical problems such as missing data and over hitting.

...read moreread less

Collapse

Random Forests

Leo Breiman

Greedy function approximation: A gradient boosting machine.

Jerome H. Friedman

- 01 Oct 2001 -

Annals of Statistics

UCI Machine Learning Repository

A. Asuncion

A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting

Yoav Freund, +1 more

Frequently Asked Questions (2)

Q1. What have the authors contributed in "An up-to-date comparison of state- of-the-art classification algorithms" ?

Moreover, important properties such as the dependency on number of classes and features and CPU running time are typically not examined. In this paper, the authors carry out a comparative empirical study on both established classifiers and more recently proposed ones on 71 data sets originating from different domains, publicly available at UCI and KEEL repositories. The list of 11 algorithms studied includes Extreme Learning Machine ( ELM ), Sparse Representation based Classification ( SRC ), and Deep Learning ( DL ), which have not been thoroughly investigated in existing comparative studies.

Q2. What have the authors stated for future works in "An up-to-date comparison of state- of-the-art classification algorithms" ?

In the future work, the authors will further investigate the performance of the 11 classifiers in specific application domains and with different feature selection methods.

An up-to-date comparison of state-of-the-art classification algorithms

Figures

Citations

Impact-slip experiments and systematic study of coal gangue “category” recognition technology Part I: Impact-slip experiments between coal gangue mixture and top coal caving hydraulic support and the study of coal gangue “category” recognition technology

Feature selection and resampling in class imbalance learning: Which comes first? An empirical study in the biological domain

User preference modeling based on meta paths and diversity regularization in heterogeneous information networks

A machine learning model to assess the ecosystem response to water policy measures in the Tagus River Basin (Spain).

Fast kernel extreme learning machine for ordinal regression

References

Random Forests

ImageNet Classification with Deep Convolutional Neural Networks

LIBSVM: A library for support vector machines

Support-Vector Networks

C4.5: Programs for Machine Learning

Related Papers (5)

Random Forests

Scikit-learn: Machine Learning in Python

Greedy function approximation: A gradient boosting machine.

UCI Machine Learning Repository

A Decision-Theoretic Generalization of On-Line Learning and an Application to Boosting

Frequently Asked Questions (2)

Q1. What have the authors contributed in "An up-to-date comparison of state- of-the-art classification algorithms" ?

Q2. What have the authors stated for future works in "An up-to-date comparison of state- of-the-art classification algorithms" ?

Trending Questions (1)