Nearest Neighbor Imputation for Categorical Data by Weighting of Attributes

Open AccessPosted Content

Nearest Neighbor Imputation for Categorical Data by Weighting of Attributes

- 03 Oct 2017 -

TLDR

The weighted nearest neighbors approach is extended to impute missing values in categorical variables and shows that the weighting of attributes yields smaller imputation errors than existing approaches.

Abstract:

Missing values are a common phenomenon in all areas of applied research. While various imputation methods are available for metrically scaled variables, methods for categorical data are scarce. An imputation method that has been shown to work well for high dimensional metrically scaled variables is the imputation by nearest neighbor methods. In this paper, we extend the weighted nearest neighbors approach to impute missing values in categorical variables. The proposed method, called $\mathtt{wNNSel_{cat}}$, explicitly uses the information on association among attributes. The performance of different imputation methods is compared in terms of the proportion of falsely imputed values. Simulation results show that the weighting of attributes yields smaller imputation errors than existing approaches. A variety of real data sets is used to support the results obtained by simulations.

Citations

PDF

Open Access

More filters

Journal ArticleDOI

An integrated BIM-LEED application to automate sustainable design assessment framework at the conceptual stage of building projects

Farzad Jalaei, +2 more

- 01 Feb 2020 -

Sustainable Cities and Society

TL;DR: A plug-in is developed to calculate and predict the potential accumulated LEED credits with access to the Application Program Interface of the BIM tool, energy analysis and lighting simulation tool, Google Map and their associated library, and to propose the whole scale innovative green building evaluation interface for building projects.

...read moreread less

Journal ArticleDOI

Sparse Convolutional Denoising Autoencoders for Genotype Imputation.

Junjie Chen, +1 more

- 28 Aug 2019 -

Genes

TL;DR: This study proposes a deep model called a sparse convolutional denoising autoencoder (SCDA) to impute missing genotypes and shows that SCDA has strong robustness and significantly outperforms popular reference-free imputation methods.

...read moreread less

A stimulation study to evaluate the performance of model-based multiple imputations in HCHS health examination surveys

T. M. Ezatti-Rice, +5 more

Journal ArticleDOI

Multi-objective ensemble deep learning using electronic health records to predict outcomes after lung cancer radiotherapy

Rongfang Wang, +6 more

- 13 Dec 2019 -

Physics in Medicine and Biology

TL;DR: A reliable multi-objective ensemble deep learning method that uses features extracted from EHRs to predict high risk of treatment failure after radiotherapy in patients with lung cancer and the experimental results demonstrate that MoEDL can perform better than other conventional methods.

...read moreread less

Proceedings ArticleDOI

Prediction of mortality in patients with cardiovascular disease using data mining methods

Damir Imamovic, +2 more

TL;DR: Based on data on patients with cardiovascular disease, collected from 2011 to 2017 at Mostar Hospital, models for mortality prediction using techniques for data tree mining, neural network and logistic regression are presented.

...read moreread less

References

PDF

Open Access

More filters

Journal ArticleDOI

Random Forests

Leo Breiman

TL;DR: Internal estimates monitor error, strength, and correlation and these are used to show the response to increasing the number of features used in the forest, and are also applicable to regression.

...read moreread less

Journal ArticleDOI

A Coefficient of agreement for nominal Scales

Jacob Cohen

- 01 Apr 1960 -

Educational and Psychological Measuremen...

TL;DR: In this article, the authors present a procedure for having two or more judges independently categorize a sample of units and determine the degree, significance, and significance of the units. But they do not discuss the extent to which these judgments are reproducible, i.e., reliable.

...read moreread less

Book

Statistical Analysis with Missing Data

Roderick J. A. Little, +1 more

TL;DR: This work states that maximum Likelihood for General Patterns of Missing Data: Introduction and Theory with Ignorable Nonresponse and large-Sample Inference Based on Maximum Likelihood Estimates is likely to be high.

...read moreread less

Book

Multiple imputation for nonresponse in surveys

Donald B. Rubin

TL;DR: In this article, a survey of drinking behavior among men of retirement age was conducted and the results showed that the majority of the participants reported that they did not receive any benefits from the Social Security Administration.

...read moreread less

Journal ArticleDOI

Missing data: Our view of the state of the art.

Joseph L. Schafer, +1 more

- 01 Jun 2002 -

Psychological Methods

TL;DR: 2 general approaches that come highly recommended: maximum likelihood (ML) and Bayesian multiple imputation (MI) are presented and may eventually extend the ML and MI methods that currently represent the state of the art.

...read moreread less