Uncategorized
-
Exploring the Science of Variable Relationships: A Guide to Structural Equation Modeling
PROJECT OVERVIEW This project provides a detailed exploration of structural equation modeling, from theoretical foundations and assumptions through model specification, evaluation, comparison, and implementation in R. The work examines how SEM can represent relationships among observed and latent variables, how models are identified and estimated, and how researchers evaluate fit and refine models. It also… Continue reading
-
Predicting Employee Engagement and Satisfaction with Transformational Leadership
PROJECT OVERVIEW This original research project examines whether transformational leadership measures can predict employee engagement and employee satisfaction. The analysis develops a research objective and hypotheses, prepares a modified survey dataset, and compares multiple predictive approaches. It evaluates decision trees, random forests, logistic regression, and support vector machines alongside Pearson correlations to examine both predictive… Continue reading
-
Behavioral Insights from Website Analytics: An Ecommerce Case Study
Today I will explore several insights derived from my analysis. The focal points of discussion will encompass a range of subjects, including the assessment of browser and device compatibility with the webpage, identification of challenges encountered by users employing mobile devices, the concerning decline in add-to-cart rates, as well as the overall stagnant nature of… Continue reading
-
Bayesian Network: Infant Clinical Presentations
PROJECT OVERVIEW This project builds a Bayesian Network in R to reason about the likelihood of possible diseases from clinical presentations in infants. The analysis starts by defining the network structure from an established model, parameterizing it with medical data, visualizing the resulting network, and performing evidence-based probability queries. The goal is to show how… Continue reading
-
Exploring Neonatal and Maternal Predictors of Germinal Matrix Hemorrhage in Low Birth Weight Infants with Logistic Regression and Odds Ratios
A low birth weight dataset records 100 births and whether or not these babies experiences a germinal matrix hemorrhage (grmhem). The babies’ 5 minute apgar score is recorded (apgar5) and also whether or not the mother had toxemia (tox) during her pregnancy. The odds ratio of whether a baby experienced a germinal matrix hemorrhage associated… Continue reading
-
Supply Chain Analytics: Mapping Coffeehouse Locations, Understanding Customers with Cohort Analysis, RFM Analysis, and K-Means Clustering, & Exploring Top-Quality Coffee Producers for Supplier Recommendations
SELECTED PROJECT Supply Chain Analytics Mapping coffeehouse locations, understanding customers with cohort analysis, RFM analysis, and K-means clustering, and exploring top-quality coffee producers for supplier recommendations. Overview This project applies multiple analytical approaches to a supply chain and customer analytics problem. It combines geographic visualization, cohort analysis, RFM analysis, and K-means clustering to understand customer… Continue reading
-
Bladder and Brain Cancer Survival Analysis
The Kaplan-Meier survival function gives the probability of surviving past time, t: 86 bladder cancer patients had tumors removed. After removal, these patients were separated into two groups, a placebo group (group 0), and a drug Thiopeta treatment group (group 1). The variable time in this dataset represents how many months until a tumor reoccurred… Continue reading
-
Comparing Self-Organizing Map, Complete Linkage Hierarchical Clustering, and Principal Component Analysis on the NCI Microarray Data
PROJECT OVERVIEW This project examines high-dimensional gene-expression data from cancer cell lines using unsupervised learning and dimensionality-reduction methods. The analysis compares self-organizing maps, hierarchical clustering, and principal component analysis to identify structure in the NCI microarray data. It evaluates how different approaches group cancer cell lines and considers the tradeoffs between complex clustering workflows and… Continue reading
-
Comparing Machine Learning Models to Predict Death Due to Heart Failure and Diabetes
PROJECT OVERVIEW This project compares machine-learning approaches for predicting death events associated with heart failure and diabetes. The analysis uses clinical datasets to build, evaluate, and compare several predictive models. It examines decision trees, random forests, bagging, boosting, k-nearest neighbors, and logistic regression, with attention to test error, variable importance, model performance, and interpretability. What… Continue reading
-
Predicting Chronic Kidney Disease with Machine Learning in R
My goal with this analysis was to predict chronic kidney disease based on 24 attributes. The dimensions of the data are 400 by 25, including the dependent variables. There are 150 observations of patients who do not have chronic kidney disease and 250 observations of patients with chronic kidney disease. I hoped to create a… Continue reading
-
Exploring the State and Arrests Data with Hierarchical Clustering, Stars Plots, and a Self-Organizing Map
In the first part of this analysis I am going to show an example of hierarchical clustering and how correlation can help aid in the understanding of dendrogram results. I will then explore some star plots. State data released from the US department of Commerce, Bureau of the Census is available in R. I will… Continue reading
-
Exploring the State and Arrests Data with Hierarchical Clustering, Stars Plots, and a Self-Organizing Map
In the first part of this analysis I am going to show an example of hierarchical clustering and how correlation can help aid in the understanding of dendrogram results. I will then explore some star plots. State data released from the US department of Commerce, Bureau of the Census is available in R. I will… Continue reading
-
Using the Apriori Algorithm to Discover Association Rules
For this tutorial, I am going to use the transaction Income dataset from the R package arules. This dataset comes from the website for the book The Elements of Statistical Learning. Chapter 14 has information about association rules. Here is a link and a citation for that book: Hastie, T., Tibshirani, R. & Friedman, J. (2001) The Elements… Continue reading
-
Derive Generalized Association Rules by Disguising an Unsupervised Learning Problem as a Supervised Learning Problem with CART
The income data used for this analysis can be found under the marketing database provided by The Elements of Statistical Learning by Hastie, T., Tibshirani, R., & Friedman, J. H.. Here is the link: https://hastie.su.domains/ElemStatLearn/data.html Attribute information includes household income, sex, marital status, age, education, occupation, how long the person has lived in the San Francisco/Oakland/San Jose area… Continue reading
-
Hypothesis Testing: ANOVA, Chi-Square Goodness of Fit, One Sample Z-Test, One Sample T-Test, One Sample Variance Test, Two Sample Z-Test, Two Sample T-Test, Paired T-Test, Two Sample Variance Test
Hypothesis Testing: ANOVA The following is a compilation of a series of ANOVA tests conducted over the course of a few years. Analysis of Variance (ANOVA) is a statistical technique used to compare the means of groups to determine if there are any statistically significant differences between them. It helps in understanding whether there are variations… Continue reading
