mirror of
https://github.com/mudabbir-ahmad/UNI-PROG3-CW2-MLWP.git
synced 2026-10-08 04:10:20 +00:00
Part G done
This commit is contained in:
1 parent
9d8ae2eea4
commit
8368c4417a
1 file changed
+49
-10
@@ -55,7 +55,14 @@
|
|||||||
{
|
{
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "## Part A\n",
|
"source": [
|
||||||
|
"## Part A\n",
|
||||||
|
"\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Check for any missing values. If any, delete all rows that contain null values Then,\n",
|
||||||
|
"check for duplicates (similar contents for every feature). If any, delete the\n",
|
||||||
|
"redundant rows. **[1 mark]**"
|
||||||
|
],
|
||||||
"id": "751a6161e8e12bdd"
|
"id": "751a6161e8e12bdd"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -305,7 +312,12 @@
|
|||||||
}
|
}
|
||||||
},
|
},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "## Part B",
|
"source": [
|
||||||
|
"## Part B\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Create a new column called 'Size in bytes' (numeric) and convert the entries\n",
|
||||||
|
"from 'Size' column (M means megabyte and k means kilobytes). **[1 mark]**\n"
|
||||||
|
],
|
||||||
"id": "5e0e0e76635904b6"
|
"id": "5e0e0e76635904b6"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -772,21 +784,27 @@
|
|||||||
}
|
}
|
||||||
},
|
},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "# Part C",
|
"source": [
|
||||||
|
"## Part C\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Create a new column called ‘Numeric_installs' (numeric) and convert the entries\n",
|
||||||
|
"from ‘Installs’ column (remove ‘+’ and ‘,’; for example, “5,000,000+” becomes\n",
|
||||||
|
"“5000000” as an Integer). **[1 mark]**"
|
||||||
|
],
|
||||||
"id": "9883e69704d0a69e"
|
"id": "9883e69704d0a69e"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"metadata": {
|
"metadata": {
|
||||||
"ExecuteTime": {
|
"ExecuteTime": {
|
||||||
"end_time": "2026-04-25T19:24:54.541583100Z",
|
"end_time": "2026-04-25T19:39:31.412418700Z",
|
||||||
"start_time": "2026-04-25T19:24:54.526548900Z"
|
"start_time": "2026-04-25T19:39:31.404386600Z"
|
||||||
}
|
}
|
||||||
},
|
},
|
||||||
"cell_type": "code",
|
"cell_type": "code",
|
||||||
"source": "df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)",
|
"source": "df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)",
|
||||||
"id": "c8d4f46526918c20",
|
"id": "c8d4f46526918c20",
|
||||||
"outputs": [],
|
"outputs": [],
|
||||||
"execution_count": 1298
|
"execution_count": 1351
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
"metadata": {
|
"metadata": {
|
||||||
@@ -1048,7 +1066,11 @@
|
|||||||
{
|
{
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "## Part D",
|
"source": [
|
||||||
|
"## Part D\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Save the updated dataset as “googleplaystore_new_new.csv” **[1 mark]**\n"
|
||||||
|
],
|
||||||
"id": "1818d3183dae5398"
|
"id": "1818d3183dae5398"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -1067,7 +1089,13 @@
|
|||||||
{
|
{
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "# Part E",
|
"source": [
|
||||||
|
"# Part E\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Using (Category + Reviews + Content Rating + Size in Bytes +\n",
|
||||||
|
"Installs_Num), find and discuss the best regression model to predict “rating”\n",
|
||||||
|
"(use the standard training/test partition without cross-validation). **[7 marks]**\n"
|
||||||
|
],
|
||||||
"id": "5372c799c0b2662"
|
"id": "5372c799c0b2662"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -2327,7 +2355,12 @@
|
|||||||
{
|
{
|
||||||
"metadata": {},
|
"metadata": {},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "## Part F",
|
"source": [
|
||||||
|
"## Part F\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Train and test all necessary model(s) to show and discuss the most predictive\n",
|
||||||
|
"feature in the previous question **[8 marks]**\n"
|
||||||
|
],
|
||||||
"id": "3c89e2f8ae31de3b"
|
"id": "3c89e2f8ae31de3b"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
@@ -2632,7 +2665,13 @@
|
|||||||
}
|
}
|
||||||
},
|
},
|
||||||
"cell_type": "markdown",
|
"cell_type": "markdown",
|
||||||
"source": "## Part G",
|
"source": [
|
||||||
|
"## Part G\n",
|
||||||
|
"\n",
|
||||||
|
"Q: Using (Category + Reviews + Content Rating + Rating + Installs_Num),\n",
|
||||||
|
"find and discuss the best regression model to predict “Size in Bytes” (use the\n",
|
||||||
|
"standard training/test partition with cross-validation). **[8 marks]**\n"
|
||||||
|
],
|
||||||
"id": "de14342fb4966baf"
|
"id": "de14342fb4966baf"
|
||||||
},
|
},
|
||||||
{
|
{
|
||||||
|
|||||||
Reference in new issue
Block a user