Part G done

This commit is contained in:
bobbert committed 2026-04-25 20:41:17 +01:00
1 parent 9d8ae2eea4
commit 8368c4417a
1 file changed
+49 -10
+49 -10
View File
@@ -55,7 +55,14 @@
{
"metadata": {},
"cell_type": "markdown",
"source": "## Part A\n",
"source": [
"## Part A\n",
"\n",
"\n",
"Q: Check for any missing values. If any, delete all rows that contain null values Then,\n",
"check for duplicates (similar contents for every feature). If any, delete the\n",
"redundant rows. **[1 mark]**"
],
"id": "751a6161e8e12bdd"
},
{
@@ -305,7 +312,12 @@
}
},
"cell_type": "markdown",
"source": "## Part B",
"source": [
"## Part B\n",
"\n",
"Q: Create a new column called 'Size in bytes' (numeric) and convert the entries\n",
"from 'Size' column (M means megabyte and k means kilobytes). **[1 mark]**\n"
],
"id": "5e0e0e76635904b6"
},
{
@@ -772,21 +784,27 @@
}
},
"cell_type": "markdown",
"source": "# Part C",
"source": [
"## Part C\n",
"\n",
"Q: Create a new column called ‘Numeric_installs' (numeric) and convert the entries\n",
"from ‘Installs’ column (remove ‘+’ and ‘,’; for example, “5,000,000+” becomes\n",
"“5000000” as an Integer). **[1 mark]**"
],
"id": "9883e69704d0a69e"
},
{
"metadata": {
"ExecuteTime": {
"end_time": "2026-04-25T19:24:54.541583100Z",
"start_time": "2026-04-25T19:24:54.526548900Z"
"end_time": "2026-04-25T19:39:31.412418700Z",
"start_time": "2026-04-25T19:39:31.404386600Z"
}
},
"cell_type": "code",
"source": "df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)",
"id": "c8d4f46526918c20",
"outputs": [],
"execution_count": 1298
"execution_count": 1351
},
{
"metadata": {
@@ -1048,7 +1066,11 @@
{
"metadata": {},
"cell_type": "markdown",
"source": "## Part D",
"source": [
"## Part D\n",
"\n",
"Q: Save the updated dataset as “googleplaystore_new_new.csv” **[1 mark]**\n"
],
"id": "1818d3183dae5398"
},
{
@@ -1067,7 +1089,13 @@
{
"metadata": {},
"cell_type": "markdown",
"source": "# Part E",
"source": [
"# Part E\n",
"\n",
"Q: Using (Category + Reviews + Content Rating + Size in Bytes +\n",
"Installs_Num), find and discuss the best regression model to predict “rating”\n",
"(use the standard training/test partition without cross-validation). **[7 marks]**\n"
],
"id": "5372c799c0b2662"
},
{
@@ -2327,7 +2355,12 @@
{
"metadata": {},
"cell_type": "markdown",
"source": "## Part F",
"source": [
"## Part F\n",
"\n",
"Q: Train and test all necessary model(s) to show and discuss the most predictive\n",
"feature in the previous question **[8 marks]**\n"
],
"id": "3c89e2f8ae31de3b"
},
{
@@ -2632,7 +2665,13 @@
}
},
"cell_type": "markdown",
"source": "## Part G",
"source": [
"## Part G\n",
"\n",
"Q: Using (Category + Reviews + Content Rating + Rating + Installs_Num),\n",
"find and discuss the best regression model to predict “Size in Bytes” (use the\n",
"standard training/test partition with cross-validation). **[8 marks]**\n"
],
"id": "de14342fb4966baf"
},
{