{ "cells": [ { "metadata": { "collapsed": true }, "cell_type": "markdown", "source": [ "## This is the Q2 Notebook!\n", "\n", "It's tracked via GitHub! hence the need for this line for the init commit" ], "id": "bb3519b1aa083259" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.826174300Z", "start_time": "2026-04-25T21:15:54.817214200Z" } }, "cell_type": "code", "source": [ "import pandas as pd\n", "import numpy as np\n", "from sklearn.model_selection import train_test_split\n", "from sklearn.preprocessing import StandardScaler\n", "from sklearn.linear_model import LogisticRegression\n", "from sklearn.neighbors import KNeighborsClassifier\n", "from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, ConfusionMatrixDisplay\n", "import matplotlib.pyplot as plt" ], "id": "76e70b2ed9af0b56", "outputs": [], "execution_count": 19 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.842931600Z", "start_time": "2026-04-25T21:15:54.827174Z" } }, "cell_type": "code", "source": "df = pd.read_csv('data/googleplaystore_new_new.csv')", "id": "f44380615d3aba25", "outputs": [], "execution_count": 20 }, { "metadata": {}, "cell_type": "markdown", "source": [ "## Part A\n", "\n", "Q: A. Using (Rating + Reviews + Content Rating + Size in Bytes +\n", "Installs_Num), using Logistic regression and KNN, find and discuss the best\n", "classification model to predict “Category” (use the training/validation/test\n", "partition without cross-validation). **[8 marks]**\n" ], "id": "1baa7daa49445720" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.857539500Z", "start_time": "2026-04-25T21:15:54.844931800Z" } }, "cell_type": "code", "source": [ "columns_to_keep = ['Rating', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Category']\n", "df_part_a = df[columns_to_keep]" ], "id": "f6fd7137bf91f31e", "outputs": [], "execution_count": 21 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.877630200Z", "start_time": "2026-04-25T21:15:54.857539500Z" } }, "cell_type": "code", "source": [ "X_raw = df_part_a.drop('Category', axis=1)\n", "y = df_part_a['Category']" ], "id": "b5daf475a5d5ca15", "outputs": [], "execution_count": 22 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.890619Z", "start_time": "2026-04-25T21:15:54.878630300Z" } }, "cell_type": "code", "source": "X = pd.get_dummies(X_raw, columns=['Content Rating'])", "id": "99a441c665dcc19b", "outputs": [], "execution_count": 23 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.896884800Z", "start_time": "2026-04-25T21:15:54.891625500Z" } }, "cell_type": "code", "source": "X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=101)", "id": "c6eb622c0f7ec63c", "outputs": [], "execution_count": 24 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.904746500Z", "start_time": "2026-04-25T21:15:54.897883400Z" } }, "cell_type": "code", "source": "X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=101)", "id": "89e8f117f057e0e2", "outputs": [], "execution_count": 25 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.914005600Z", "start_time": "2026-04-25T21:15:54.905746300Z" } }, "cell_type": "code", "source": [ "scaler = StandardScaler()\n", "scaled_X_train = scaler.fit_transform(X_train)\n", "scaled_X_val = scaler.transform(X_val)\n", "scaled_X_test = scaler.transform(X_test)" ], "id": "d529991171fafb3e", "outputs": [], "execution_count": 26 }, { "metadata": {}, "cell_type": "markdown", "source": "## Logistic Regression", "id": "66dc17cdf1a00a06" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.935689400Z", "start_time": "2026-04-25T21:15:54.915006600Z" } }, "cell_type": "code", "source": [ "log_model = LogisticRegression(max_iter=1000)\n", "log_model.fit(scaled_X_train, y_train)" ], "id": "84128fa4e7823565", "outputs": [ { "data": { "text/plain": [ "LogisticRegression(max_iter=1000)" ], "text/html": [ "
LogisticRegression(max_iter=1000)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
" ] }, "execution_count": 27, "metadata": {}, "output_type": "execute_result" } ], "execution_count": 27 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.948876Z", "start_time": "2026-04-25T21:15:54.937689300Z" } }, "cell_type": "code", "source": "y_val_pred_log = log_model.predict(scaled_X_val)", "id": "965e16ce48443c97", "outputs": [], "execution_count": 28 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.976908500Z", "start_time": "2026-04-25T21:15:54.948876Z" } }, "cell_type": "code", "source": [ "log_val_f1 = f1_score(y_val, y_val_pred_log, average='weighted')\n", "print(f\"Validation Accuracy: {accuracy_score(y_val, y_val_pred_log):.4f}\")\n", "print(f\"Validation F1-Score: {log_val_f1:.4f}\")" ], "id": "9be803af2f34cd6", "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "Validation Accuracy: 0.3333\n", "Validation F1-Score: 0.3069\n" ] } ], "execution_count": 29 }, { "metadata": {}, "cell_type": "markdown", "source": "## KNN", "id": "2b7b85284cddecb2" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.988744700Z", "start_time": "2026-04-25T21:15:54.977907500Z" } }, "cell_type": "code", "source": [ "best_k = 1\n", "best_f1 = 0" ], "id": "8c4835d7a413a6e6", "outputs": [], "execution_count": 30 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:55.029426700Z", "start_time": "2026-04-25T21:15:54.989744500Z" } }, "cell_type": "code", "source": [ "for k in range(1, 15, 2):\n", " knn_temp = KNeighborsClassifier(n_neighbors=k)\n", " knn_temp.fit(scaled_X_train, y_train)\n", " y_val_pred_knn = knn_temp.predict(scaled_X_val)\n", "\n", " current_f1 = f1_score(y_val, y_val_pred_knn, average='weighted')\n", " print(f\"KNN (K={k}) Validation F1-Score: {current_f1:.4f}\")\n", "\n", " if current_f1 > best_f1:\n", " best_f1 = current_f1\n", " best_k = k" ], "id": "a8d69c56d1a21825", "outputs": [ { "name": "stdout", "output_type": "stream", "text": [ "KNN (K=1) Validation F1-Score: 0.3201\n", "KNN (K=3) Validation F1-Score: 0.3245\n", "KNN (K=5) Validation F1-Score: 0.2939\n", "KNN (K=7) Validation F1-Score: 0.2916\n", "KNN (K=9) Validation F1-Score: 0.3120\n", "KNN (K=11) Validation F1-Score: 0.3097\n", "KNN (K=13) Validation F1-Score: 0.3181\n" ] } ], "execution_count": 31 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:55.045639100Z", "start_time": "2026-04-25T21:15:55.030432300Z" } }, "cell_type": "code", "source": [ "knn_final = KNeighborsClassifier(n_neighbors=best_k)\n", "knn_final.fit(scaled_X_train, y_train)" ], "id": "9644f7a78b4ff687", "outputs": [ { "data": { "text/plain": [ "KNeighborsClassifier(n_neighbors=3)" ], "text/html": [ "
KNeighborsClassifier(n_neighbors=3)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
" ] }, "execution_count": 32, "metadata": {}, "output_type": "execute_result" } ], "execution_count": 32 }, { "metadata": {}, "cell_type": "markdown", "source": [ "## Part B\n", "\n", "Q: Using (Rating + Reviews + Category + Size in Bytes + Installs_Num),\n", "using Logistic regression and KNN, find and discuss the best classification model\n", "to predict “Content Rating” (use the training/validation/test partition without\n", "cross-validation). **[7 marks]**" ], "id": "f15aa6cca58826fc" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:55.065809200Z", "start_time": "2026-04-25T21:15:55.059661900Z" } }, "cell_type": "code", "source": "", "id": "65e55b70ad44b2dc", "outputs": [], "execution_count": 32 }, { "metadata": {}, "cell_type": "markdown", "source": [ "## Part C\n", "\n", "Q: By considering Installs_Num as categorical feature and using (Rating +\n", "Reviews + Category + Size in Bytes + Content Rating), find and discuss the\n", "best classification model to predict “Installs_Num” (using Logistic regression and\n", "KNN) (use the training/test partition without cross-validation) **[8 marks]**" ], "id": "86c040f5e043a39c" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:55.070341Z", "start_time": "2026-04-25T21:15:55.066313200Z" } }, "cell_type": "code", "source": "", "id": "f3535461938f94f5", "outputs": [], "execution_count": 32 } ], "metadata": { "kernelspec": { "display_name": "Python 3", "language": "python", "name": "python3" }, "language_info": { "codemirror_mode": { "name": "ipython", "version": 2 }, "file_extension": ".py", "mimetype": "text/x-python", "name": "python", "nbconvert_exporter": "python", "pygments_lexer": "ipython2", "version": "2.7.6" } }, "nbformat": 4, "nbformat_minor": 5 }