mirror of
https://github.com/mudabbir-ahmad/UNI-PROG3-CW2-MLWP.git
synced 2026-10-07 20:10:20 +00:00
Refined Answers!
This commit is contained in:
1 parent
c0a3394f80
commit
1fcce03e9f
2 files changed
+23
-7
No files matched your search
@@ -4863,28 +4863,36 @@
|
||||
"source": [
|
||||
"### Part G Answer (with Cross-Validation)\n",
|
||||
"\n",
|
||||
"For Part G, the goal is to predict **`Size in bytes`** using the other inputs:\n",
|
||||
"For Part G, the goal was to predict **`Size in bytes`** using the other inputs:\n",
|
||||
"- `Category` \n",
|
||||
"- `Reviews` \n",
|
||||
"- `Content Rating` \n",
|
||||
"- `Rating` \n",
|
||||
"- `Numeric Installs` \n",
|
||||
"\n",
|
||||
"I compared the allowed regression models using a standard train/test split, but this time I also added **5-fold cross-validation** on the training set.\n",
|
||||
"I compared the regression models using a standard train/test split, but this time I also added 5-fold cross-validation on the training set.\n",
|
||||
"\n",
|
||||
"**How I decided which model is best:**\n",
|
||||
"- First, I looked for the model with the highest **Cross-Validation mean R²** since that tells me how well the model generalizes across different splits of the data.\n",
|
||||
"- Then, I used **Test RMSE** as a secondary check (lower is better).\n",
|
||||
"First, I looked for the model with the highest Cross-Validation mean R2 since that tells me how well the model generalizes across different splits of the data.\n",
|
||||
"Then, I used the test RMSE as a secondary check (lower is better).\n",
|
||||
"\n",
|
||||
"**What the results show:**\n",
|
||||
"- **Linear Regression** is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n",
|
||||
"- **Polynomial Regression** actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n",
|
||||
"- **Ridge Regression** is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n",
|
||||
"- Linear Regression is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n",
|
||||
"- Polynomial Regression actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n",
|
||||
"- Ridge Regression is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n",
|
||||
"\n",
|
||||
"So, based on the cross-validation and test metrics, **Linear Regression is the best model for predicting `Size in bytes`**.\n",
|
||||
"\n",
|
||||
"I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).\n"
|
||||
]
|
||||
},
|
||||
{
|
||||
"metadata": {},
|
||||
"cell_type": "code",
|
||||
"outputs": [],
|
||||
"execution_count": null,
|
||||
"source": "",
|
||||
"id": "bb62142045c39ce8"
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
|
||||
@@ -5322,6 +5322,14 @@
|
||||
"The KNN Model with the K value being set to 41, provides the closest predicted data distribution to the actual distribution of the test set, and also has a higher weighted F1 score compared to Logistic Regression. Therefore, KNN is the best classification model for predicting \"Installs_Num\" in Part C."
|
||||
],
|
||||
"id": "76125c5a3ab3075a"
|
||||
},
|
||||
{
|
||||
"metadata": {},
|
||||
"cell_type": "code",
|
||||
"outputs": [],
|
||||
"execution_count": null,
|
||||
"source": "",
|
||||
"id": "157be5d5faea96fe"
|
||||
}
|
||||
],
|
||||
"metadata": {
|
||||
|
||||
Reference in new issue
Block a user