{ "cells": [ { "metadata": { "collapsed": true }, "cell_type": "markdown", "source": [ "## This is the Q2 Notebook!\n", "\n", "It's tracked via GitHub! hence the need for this line for the init commit" ], "id": "bb3519b1aa083259" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.826174300Z", "start_time": "2026-04-25T21:15:54.817214200Z" } }, "cell_type": "code", "source": [ "import pandas as pd\n", "import numpy as np\n", "from sklearn.model_selection import train_test_split\n", "from sklearn.preprocessing import StandardScaler\n", "from sklearn.linear_model import LogisticRegression\n", "from sklearn.neighbors import KNeighborsClassifier\n", "from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, ConfusionMatrixDisplay\n", "import matplotlib.pyplot as plt" ], "id": "76e70b2ed9af0b56", "outputs": [], "execution_count": 19 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.842931600Z", "start_time": "2026-04-25T21:15:54.827174Z" } }, "cell_type": "code", "source": "df = pd.read_csv('data/googleplaystore_new_new.csv')", "id": "f44380615d3aba25", "outputs": [], "execution_count": 20 }, { "metadata": {}, "cell_type": "markdown", "source": [ "## Part A\n", "\n", "Q: A. Using (Rating + Reviews + Content Rating + Size in Bytes +\n", "Installs_Num), using Logistic regression and KNN, find and discuss the best\n", "classification model to predict “Category” (use the training/validation/test\n", "partition without cross-validation). **[8 marks]**\n" ], "id": "1baa7daa49445720" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.857539500Z", "start_time": "2026-04-25T21:15:54.844931800Z" } }, "cell_type": "code", "source": [ "columns_to_keep = ['Rating', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Category']\n", "df_part_a = df[columns_to_keep]" ], "id": "f6fd7137bf91f31e", "outputs": [], "execution_count": 21 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.877630200Z", "start_time": "2026-04-25T21:15:54.857539500Z" } }, "cell_type": "code", "source": [ "X_raw = df_part_a.drop('Category', axis=1)\n", "y = df_part_a['Category']" ], "id": "b5daf475a5d5ca15", "outputs": [], "execution_count": 22 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.890619Z", "start_time": "2026-04-25T21:15:54.878630300Z" } }, "cell_type": "code", "source": "X = pd.get_dummies(X_raw, columns=['Content Rating'])", "id": "99a441c665dcc19b", "outputs": [], "execution_count": 23 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.896884800Z", "start_time": "2026-04-25T21:15:54.891625500Z" } }, "cell_type": "code", "source": "X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=101)", "id": "c6eb622c0f7ec63c", "outputs": [], "execution_count": 24 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.904746500Z", "start_time": "2026-04-25T21:15:54.897883400Z" } }, "cell_type": "code", "source": "X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=101)", "id": "89e8f117f057e0e2", "outputs": [], "execution_count": 25 }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.914005600Z", "start_time": "2026-04-25T21:15:54.905746300Z" } }, "cell_type": "code", "source": [ "scaler = StandardScaler()\n", "scaled_X_train = scaler.fit_transform(X_train)\n", "scaled_X_val = scaler.transform(X_val)\n", "scaled_X_test = scaler.transform(X_test)" ], "id": "d529991171fafb3e", "outputs": [], "execution_count": 26 }, { "metadata": {}, "cell_type": "markdown", "source": "## Logistic Regression", "id": "66dc17cdf1a00a06" }, { "metadata": { "ExecuteTime": { "end_time": "2026-04-25T21:15:54.935689400Z", "start_time": "2026-04-25T21:15:54.915006600Z" } }, "cell_type": "code", "source": [ "log_model = LogisticRegression(max_iter=1000)\n", "log_model.fit(scaled_X_train, y_train)" ], "id": "84128fa4e7823565", "outputs": [ { "data": { "text/plain": [ "LogisticRegression(max_iter=1000)" ], "text/html": [ "
LogisticRegression(max_iter=1000)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
KNeighborsClassifier(n_neighbors=3)In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.