Refined Answers!

This commit is contained in:
bobbert committed 2026-04-26 16:00:29 +01:00
1 parent c0a3394f80
commit 1fcce03e9f
2 files changed
+23 -7

No files matched your search

+15 -7
View File
@@ -4863,28 +4863,36 @@
"source": [ "source": [
"### Part G Answer (with Cross-Validation)\n", "### Part G Answer (with Cross-Validation)\n",
"\n", "\n",
"For Part G, the goal is to predict **`Size in bytes`** using the other inputs:\n", "For Part G, the goal was to predict **`Size in bytes`** using the other inputs:\n",
"- `Category` \n", "- `Category` \n",
"- `Reviews` \n", "- `Reviews` \n",
"- `Content Rating` \n", "- `Content Rating` \n",
"- `Rating` \n", "- `Rating` \n",
"- `Numeric Installs` \n", "- `Numeric Installs` \n",
"\n", "\n",
"I compared the allowed regression models using a standard train/test split, but this time I also added **5-fold cross-validation** on the training set.\n", "I compared the regression models using a standard train/test split, but this time I also added 5-fold cross-validation on the training set.\n",
"\n", "\n",
"**How I decided which model is best:**\n", "**How I decided which model is best:**\n",
"- First, I looked for the model with the highest **Cross-Validation mean R²** since that tells me how well the model generalizes across different splits of the data.\n", "First, I looked for the model with the highest Cross-Validation mean R2 since that tells me how well the model generalizes across different splits of the data.\n",
"- Then, I used **Test RMSE** as a secondary check (lower is better).\n", "Then, I used the test RMSE as a secondary check (lower is better).\n",
"\n", "\n",
"**What the results show:**\n", "**What the results show:**\n",
"- **Linear Regression** is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n", "- Linear Regression is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.\n",
"- **Polynomial Regression** actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n", "- Polynomial Regression actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.\n",
"- **Ridge Regression** is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n", "- Ridge Regression is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.\n",
"\n", "\n",
"So, based on the cross-validation and test metrics, **Linear Regression is the best model for predicting `Size in bytes`**.\n", "So, based on the cross-validation and test metrics, **Linear Regression is the best model for predicting `Size in bytes`**.\n",
"\n", "\n",
"I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).\n" "I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).\n"
] ]
},
{
"metadata": {},
"cell_type": "code",
"outputs": [],
"execution_count": null,
"source": "",
"id": "bb62142045c39ce8"
} }
], ],
"metadata": { "metadata": {
+8
View File
@@ -5322,6 +5322,14 @@
"The KNN Model with the K value being set to 41, provides the closest predicted data distribution to the actual distribution of the test set, and also has a higher weighted F1 score compared to Logistic Regression. Therefore, KNN is the best classification model for predicting \"Installs_Num\" in Part C." "The KNN Model with the K value being set to 41, provides the closest predicted data distribution to the actual distribution of the test set, and also has a higher weighted F1 score compared to Logistic Regression. Therefore, KNN is the best classification model for predicting \"Installs_Num\" in Part C."
], ],
"id": "76125c5a3ab3075a" "id": "76125c5a3ab3075a"
},
{
"metadata": {},
"cell_type": "code",
"outputs": [],
"execution_count": null,
"source": "",
"id": "157be5d5faea96fe"
} }
], ],
"metadata": { "metadata": {