A Study of EM Algorithm as an Imputation Method: A Model-Based Simulation Study with Application to a Synthetic Compositional Data - Open Journal of Modelling and Simulation

OJMSi > Vol.12 No.2, April 2024

Open Journal of Modelling and Simulation

Volume 12, Issue 2 (April 2024)

ISSN Print: 2327-4018 ISSN Online: 2327-4026

Google-based Impact Factor: 2.79 Citations

A Study of EM Algorithm as an Imputation Method: A Model-Based Simulation Study with Application to a Synthetic Compositional Data ()

HTML XML

Download as PDF (Size: 447KB) PP. 33-42

DOI: 10.4236/ojmsi.2024.122002 219 Downloads 806 Views

Author(s)

Yisa Adeniyi Abolade, Yichuan Zhao

Affiliation(s)

Department of Mathematics and Statistics, Georgia State University, Atlanta, Georgia, USA.

ABSTRACT

Compositional data, such as relative information, is a crucial aspect of machine learning and other related fields. It is typically recorded as closed data or sums to a constant, like 100%. The statistical linear model is the most used technique for identifying hidden relationships between underlying random variables of interest. However, data quality is a significant challenge in machine learning, especially when missing data is present. The linear regression model is a commonly used statistical modeling technique used in various applications to find relationships between variables of interest. When estimating linear regression parameters which are useful for things like future prediction and partial effects analysis of independent variables, maximum likelihood estimation (MLE) is the method of choice. However, many datasets contain missing observations, which can lead to costly and time-consuming data recovery. To address this issue, the expectation-maximization (EM) algorithm has been suggested as a solution for situations including missing data. The EM algorithm repeatedly finds the best estimates of parameters in statistical models that depend on variables or data that have not been observed. This is called maximum likelihood or maximum a posteriori (MAP). Using the present estimate as input, the expectation (E) step constructs a log-likelihood function. Finding the parameters that maximize the anticipated log-likelihood, as determined in the E step, is the job of the maximization (M) phase. This study looked at how well the EM algorithm worked on a made-up compositional dataset with missing observations. It used both the robust least square version and ordinary least square regression techniques. The efficacy of the EM algorithm was compared with two alternative imputation techniques, k-Nearest Neighbor (k-NN) and mean imputation (

), in terms of Aitchison distances and covariance.

KEYWORDS

Compositional Data, Linear Regression Model, Least Square Method, Robust Least Square Method, Synthetic Data, Aitchison Distance, Maximum Likelihood Estimation, Expectation-Maximization Algorithm, k-Nearest Neighbor, and Mean imputation

Share and Cite:

Abolade, Y. and Zhao, Y. (2024) A Study of EM Algorithm as an Imputation Method: A Model-Based Simulation Study with Application to a Synthetic Compositional Data. Open Journal of Modelling and Simulation, 12, 33-42. doi: 10.4236/ojmsi.2024.122002.

Cited by

No relevant information.

Journals Menu

Follow SCIRP

	customer@scirp.org
	+86 18163351462(WhatsApp)
	1655362766

	Paper Publishing WeChat

Journals Menu

Home

About SCIRP

Service

Policies