From 1fcce03e9f86b9edda3f738821fb9878101dc3bf Mon Sep 17 00:00:00 2001 From: mudabbir-ahmad Date: Sun, 26 Apr 2026 16:00:29 +0100 Subject: [PATCH] Refined Answers! --- Q1.ipynb | 22 +++++++++++++++------- Q2.ipynb | 8 ++++++++ 2 files changed, 23 insertions(+), 7 deletions(-) diff --git a/Q1.ipynb b/Q1.ipynb index f9127f1..ab7238e 100644 --- a/Q1.ipynb +++ b/Q1.ipynb @@ -4863,28 +4863,36 @@ "source": [ "### Part G Answer (with Cross-Validation)\n", "\n", - "For Part G, the goal is to predict **`Size in bytes`** using the other inputs:\n", + "For Part G, the goal was to predict **`Size in bytes`** using the other inputs:\n", "- `Category` \n", "- `Reviews` \n", "- `Content Rating` \n", "- `Rating` \n", "- `Numeric Installs` \n", "\n", - "I compared the allowed regression models using a standard train/test split, but this time I also added **5-fold cross-validation** on the training set.\n", + "I compared the regression models using a standard train/test split, but this time I also added 5-fold cross-validation on the training set.\n", "\n", "**How I decided which model is best:**\n", - "- First, I looked for the model with the highest **Cross-Validation mean R²** since that tells me how well the model generalizes across different splits of the data.\n", - "- Then, I used **Test RMSE** as a secondary check (lower is better).\n", + "First, I looked for the model with the highest Cross-Validation mean R2 since that tells me how well the model generalizes across different splits of the data.\n", + "Then, I used the test RMSE as a secondary check (lower is better).\n", "\n", "**What the results show:**\n", - "- **Linear Regression** is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n", - "- **Polynomial Regression** actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n", - "- **Ridge Regression** is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n", + "- Linear Regression is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n", + "- Polynomial Regression actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n", + "- Ridge Regression is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n", "\n", "So, based on the cross-validation and test metrics, **Linear Regression is the best model for predicting `Size in bytes`**.\n", "\n", "I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).\n" ] + }, + { + "metadata": {}, + "cell_type": "code", + "outputs": [], + "execution_count": null, + "source": "", + "id": "bb62142045c39ce8" } ], "metadata": { diff --git a/Q2.ipynb b/Q2.ipynb index fea82af..3d94c0b 100644 --- a/Q2.ipynb +++ b/Q2.ipynb @@ -5322,6 +5322,14 @@ "The KNN Model with the K value being set to 41, provides the closest predicted data distribution to the actual distribution of the test set, and also has a higher weighted F1 score compared to Logistic Regression. Therefore, KNN is the best classification model for predicting \"Installs_Num\" in Part C." ], "id": "76125c5a3ab3075a" + }, + { + "metadata": {}, + "cell_type": "code", + "outputs": [], + "execution_count": null, + "source": "", + "id": "157be5d5faea96fe" } ], "metadata": {