The project was built in MATLAB, and structured into the following pipeline:

Data Cleaning Removed non-numeric and high-missing columns Eliminated outliers (1st–99th percentile) Dropped rows with more than 10 missing values Filled remaining missing values with zeros

Feature Selection Applied sequentialfs using a decision tree classifier to select the most relevant features based on a 3-fold cross-validation.

Model Training Used fitcensemble with bagging to build a robust classifier Split the dataset randomly into 80% training and 20% testing using cvpartition

Evaluation Assessed performance with a confusion matrix and ROC curve Measured accuracy and AUC (area under the curve) Prediction on Unlabeled Data Preprocessed the unlabeled dataset using the same feature and cleaning logic Predicted diabetes outcomes and saved results to a CSV file

πŸ“š What We Learned How to structure a complete ML pipeline in MATLAB The importance of consistent preprocessing when applying a model to new data Practical use of feature selection (sequentialfs) and ensemble methods How small issues in missing data handling or column mismatch can significantly affect model performance

⚠️ Challenges I Faced The dataset was messy, with many missing values and mixed data types Ensuring training and testing consistency was tricky β€” especially when removing or imputing missing features Understanding the impact of data leakage and preventing it by proper train/test separation MATLAB is powerful but has a steeper learning curve for certain ML workflows compared to Python

βœ… Outcome Train an accurate diabetes prediction model Evaluate it confidently Deploy it on unseen data with a clean and reproducible pipeline

Built With

Share this project:

Updates

Submission history