From 8368c4417ab6cdc3f8949383762efd3e1533bd62 Mon Sep 17 00:00:00 2001 From: mudabbir-ahmad Date: Sat, 25 Apr 2026 20:41:17 +0100 Subject: [PATCH] Part G done --- Q1.ipynb | 59 ++++++++++++++++++++++++++++++++++++++++++++++---------- 1 file changed, 49 insertions(+), 10 deletions(-) diff --git a/Q1.ipynb b/Q1.ipynb index 4cd26cc..8196db6 100644 --- a/Q1.ipynb +++ b/Q1.ipynb @@ -55,7 +55,14 @@ { "metadata": {}, "cell_type": "markdown", - "source": "## Part A\n", + "source": [ + "## Part A\n", + "\n", + "\n", + "Q: Check for any missing values. If any, delete all rows that contain null values Then,\n", + "check for duplicates (similar contents for every feature). If any, delete the\n", + "redundant rows. **[1 mark]**" + ], "id": "751a6161e8e12bdd" }, { @@ -305,7 +312,12 @@ } }, "cell_type": "markdown", - "source": "## Part B", + "source": [ + "## Part B\n", + "\n", + "Q: Create a new column called 'Size in bytes' (numeric) and convert the entries\n", + "from 'Size' column (M means megabyte and k means kilobytes). **[1 mark]**\n" + ], "id": "5e0e0e76635904b6" }, { @@ -772,21 +784,27 @@ } }, "cell_type": "markdown", - "source": "# Part C", + "source": [ + "## Part C\n", + "\n", + "Q: Create a new column called ‘Numeric_installs' (numeric) and convert the entries\n", + "from ‘Installs’ column (remove ‘+’ and ‘,’; for example, “5,000,000+” becomes\n", + "“5000000” as an Integer). **[1 mark]**" + ], "id": "9883e69704d0a69e" }, { "metadata": { "ExecuteTime": { - "end_time": "2026-04-25T19:24:54.541583100Z", - "start_time": "2026-04-25T19:24:54.526548900Z" + "end_time": "2026-04-25T19:39:31.412418700Z", + "start_time": "2026-04-25T19:39:31.404386600Z" } }, "cell_type": "code", "source": "df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)", "id": "c8d4f46526918c20", "outputs": [], - "execution_count": 1298 + "execution_count": 1351 }, { "metadata": { @@ -1048,7 +1066,11 @@ { "metadata": {}, "cell_type": "markdown", - "source": "## Part D", + "source": [ + "## Part D\n", + "\n", + "Q: Save the updated dataset as “googleplaystore_new_new.csv” **[1 mark]**\n" + ], "id": "1818d3183dae5398" }, { @@ -1067,7 +1089,13 @@ { "metadata": {}, "cell_type": "markdown", - "source": "# Part E", + "source": [ + "# Part E\n", + "\n", + "Q: Using (Category + Reviews + Content Rating + Size in Bytes +\n", + "Installs_Num), find and discuss the best regression model to predict “rating”\n", + "(use the standard training/test partition without cross-validation). **[7 marks]**\n" + ], "id": "5372c799c0b2662" }, { @@ -2327,7 +2355,12 @@ { "metadata": {}, "cell_type": "markdown", - "source": "## Part F", + "source": [ + "## Part F\n", + "\n", + "Q: Train and test all necessary model(s) to show and discuss the most predictive\n", + "feature in the previous question **[8 marks]**\n" + ], "id": "3c89e2f8ae31de3b" }, { @@ -2632,7 +2665,13 @@ } }, "cell_type": "markdown", - "source": "## Part G", + "source": [ + "## Part G\n", + "\n", + "Q: Using (Category + Reviews + Content Rating + Rating + Installs_Num),\n", + "find and discuss the best regression model to predict “Size in Bytes” (use the\n", + "standard training/test partition with cross-validation). **[8 marks]**\n" + ], "id": "de14342fb4966baf" }, {