Files
UNI-PROG3-CW2-MLWP/Q2.ipynb
T
2026-04-25 23:15:31 +01:00

327 KiB

This is the Q2 Notebook!

It's tracked via GitHub! hence the need for this line for the init commit

In [148]:
import pandas as pd
import numpy as np
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score, f1_score, classification_report, confusion_matrix, ConfusionMatrixDisplay
import matplotlib.pyplot as plt
In [149]:
df = pd.read_csv('data/googleplaystore_new_new.csv')

Part A

Q: A. Using (Rating + Reviews + Content Rating + Size in Bytes + Installs_Num), using Logistic regression and KNN, find and discuss the best classification model to predict “Category” (use the training/validation/test partition without cross-validation). [8 marks]

In [150]:
columns_to_keep = ['Rating', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Category']
df_part_a = df[columns_to_keep]
In [151]:
X_raw = df_part_a.drop('Category', axis=1)
y = df_part_a['Category']
In [152]:
X = pd.get_dummies(X_raw, columns=['Content Rating'])
In [153]:
X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=101)
In [154]:
X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=101)
In [155]:
scaler = StandardScaler()
scaled_X_train = scaler.fit_transform(X_train)
scaled_X_val = scaler.transform(X_val)
scaled_X_test = scaler.transform(X_test)
In [156]:
def eval_metrics(y_true, y_pred, label="Model"):
    acc = accuracy_score(y_true, y_pred)
    f1_w = f1_score(y_true, y_pred, average='weighted')
    print(f"{label} -> Accuracy: {acc:.4f} | Weighted F1: {f1_w:.4f}")
    return acc, f1_w
In [156]:

Logistic Regression

In [157]:
log_model = LogisticRegression(max_iter=1000, random_state=101)
log_model.fit(scaled_X_train, y_train)
Out [157]:
LogisticRegression(max_iter=1000, random_state=101)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
penalty penalty: {'l1', 'l2', 'elasticnet', None}, default='l2'

Specify the norm of the penalty:

- `None`: no penalty is added;
- `'l2'`: add a L2 penalty term and it is the default choice;
- `'l1'`: add a L1 penalty term;
- `'elasticnet'`: both L1 and L2 penalty terms are added.

.. warning::
Some penalties may not work with some solvers. See the parameter
`solver` below, to know the compatibility between the penalty and
solver.

.. versionadded:: 0.19
l1 penalty with SAGA solver (allowing 'multinomial' + L1)

.. deprecated:: 1.8
`penalty` was deprecated in version 1.8 and will be removed in 1.10.
Use `l1_ratio` instead. `l1_ratio=0` for `penalty='l2'`, `l1_ratio=1` for
`penalty='l1'` and `l1_ratio` set to any float between 0 and 1 for
`'penalty='elasticnet'`.
'deprecated'
C C: float, default=1.0

Inverse of regularization strength; must be a positive float.
Like in support vector machines, smaller values specify stronger
regularization. `C=np.inf` results in unpenalized logistic regression.
For a visual example on the effect of tuning the `C` parameter
with an L1 penalty, see:
:ref:`sphx_glr_auto_examples_linear_model_plot_logistic_path.py`.
1.0
l1_ratio l1_ratio: float, default=0.0

The Elastic-Net mixing parameter, with `0 <= l1_ratio <= 1`. Setting
`l1_ratio=1` gives a pure L1-penalty, setting `l1_ratio=0` a pure L2-penalty.
Any value between 0 and 1 gives an Elastic-Net penalty of the form
`l1_ratio * L1 + (1 - l1_ratio) * L2`.

.. warning::
Certain values of `l1_ratio`, i.e. some penalties, may not work with some
solvers. See the parameter `solver` below, to know the compatibility between
the penalty and solver.

.. versionchanged:: 1.8
Default value changed from None to 0.0.

.. deprecated:: 1.8
`None` is deprecated and will be removed in version 1.10. Always use
`l1_ratio` to specify the penalty type.
0.0
dual dual: bool, default=False

Dual (constrained) or primal (regularized, see also
:ref:`this equation `) formulation. Dual formulation
is only implemented for l2 penalty with liblinear solver. Prefer `dual=False`
when n_samples > n_features.
False
tol tol: float, default=1e-4

Tolerance for stopping criteria.
0.0001
fit_intercept fit_intercept: bool, default=True

Specifies if a constant (a.k.a. bias or intercept) should be
added to the decision function.
True
intercept_scaling intercept_scaling: float, default=1

Useful only when the solver `liblinear` is used
and `self.fit_intercept` is set to `True`. In this case, `x` becomes
`[x, self.intercept_scaling]`,
i.e. a "synthetic" feature with constant value equal to
`intercept_scaling` is appended to the instance vector.
The intercept becomes
``intercept_scaling * synthetic_feature_weight``.

.. note::
The synthetic feature weight is subject to L1 or L2
regularization as all other features.
To lessen the effect of regularization on synthetic feature weight
(and therefore on the intercept) `intercept_scaling` has to be increased.
1
class_weight class_weight: dict or 'balanced', default=None

Weights associated with classes in the form ``{class_label: weight}``.
If not given, all classes are supposed to have weight one.

The "balanced" mode uses the values of y to automatically adjust
weights inversely proportional to class frequencies in the input data
as ``n_samples / (n_classes * np.bincount(y))``.

Note that these weights will be multiplied with sample_weight (passed
through the fit method) if sample_weight is specified.

.. versionadded:: 0.17
*class_weight='balanced'*
None
random_state random_state: int, RandomState instance, default=None

Used when ``solver`` == 'sag', 'saga' or 'liblinear' to shuffle the
data. See :term:`Glossary ` for details.
101
solver solver: {'lbfgs', 'liblinear', 'newton-cg', 'newton-cholesky', 'sag', 'saga'}, default='lbfgs'

Algorithm to use in the optimization problem. Default is 'lbfgs'.
To choose a solver, you might want to consider the following aspects:

- 'lbfgs' is a good default solver because it works reasonably well for a wide
class of problems.
- For :term:`multiclass` problems (`n_classes >= 3`), all solvers except
'liblinear' minimize the full multinomial loss, 'liblinear' will raise an
error.
- 'newton-cholesky' is a good choice for
`n_samples` >> `n_features * n_classes`, especially with one-hot encoded
categorical features with rare categories. Be aware that the memory usage
of this solver has a quadratic dependency on `n_features * n_classes`
because it explicitly computes the full Hessian matrix.
- For small datasets, 'liblinear' is a good choice, whereas 'sag'
and 'saga' are faster for large ones;
- 'liblinear' can only handle binary classification by default. To apply a
one-versus-rest scheme for the multiclass setting one can wrap it with the
:class:`~sklearn.multiclass.OneVsRestClassifier`.

.. warning::
The choice of the algorithm depends on the penalty chosen (`l1_ratio=0`
for L2-penalty, `l1_ratio=1` for L1-penalty and `0 < l1_ratio < 1` for
Elastic-Net) and on (multinomial) multiclass support:

================= ======================== ======================
solver l1_ratio multinomial multiclass
================= ======================== ======================
'lbfgs' l1_ratio=0 yes
'liblinear' l1_ratio=1 or l1_ratio=0 no
'newton-cg' l1_ratio=0 yes
'newton-cholesky' l1_ratio=0 yes
'sag' l1_ratio=0 yes
'saga' 0<=l1_ratio<=1 yes
================= ======================== ======================

.. note::
'sag' and 'saga' fast convergence is only guaranteed on features
with approximately the same scale. You can preprocess the data with
a scaler from :mod:`sklearn.preprocessing`.

.. seealso::
Refer to the :ref:`User Guide ` for more
information regarding :class:`LogisticRegression` and more specifically the
:ref:`Table `
summarizing solver/penalty supports.

.. versionadded:: 0.17
Stochastic Average Gradient (SAG) descent solver. Multinomial support in
version 0.18.
.. versionadded:: 0.19
SAGA solver.
.. versionchanged:: 0.22
The default solver changed from 'liblinear' to 'lbfgs' in 0.22.
.. versionadded:: 1.2
newton-cholesky solver. Multinomial support in version 1.6.
'lbfgs'
max_iter max_iter: int, default=100

Maximum number of iterations taken for the solvers to converge.
1000
verbose verbose: int, default=0

For the liblinear and lbfgs solvers set verbose to any positive
number for verbosity.
0
warm_start warm_start: bool, default=False

When set to True, reuse the solution of the previous call to fit as
initialization, otherwise, just erase the previous solution.
Useless for liblinear solver. See :term:`the Glossary `.

.. versionadded:: 0.17
*warm_start* to support *lbfgs*, *newton-cg*, *sag*, *saga* solvers.
False
n_jobs n_jobs: int, default=None

Does not have any effect.

.. deprecated:: 1.8
`n_jobs` is deprecated in version 1.8 and will be removed in 1.10.
None
In [158]:
y_val_pred_log = log_model.predict(scaled_X_val)
log_val_acc, log_val_f1 = eval_metrics(y_val, y_val_pred_log, "LogReg (Validation)")
LogReg (Validation) -> Accuracy: 0.3333 | Weighted F1: 0.3069
In [159]:
y_test_pred_log = log_model.predict(scaled_X_test)
log_test_acc, log_test_f1 = eval_metrics(y_test, y_test_pred_log, "LogReg (Test)")
LogReg (Test) -> Accuracy: 0.3108 | Weighted F1: 0.2646
In [160]:
log_probs = log_model.predict_proba(scaled_X_test)
log_pred_idx = np.argmax(log_probs, axis=1)
log_pred_conf = log_probs[np.arange(len(log_probs)), log_pred_idx]

log_pred_df = pd.DataFrame({
    "True_Label": y_test.reset_index(drop=True),
    "Predicted_Label": pd.Series(y_test_pred_log).reset_index(drop=True),
    "Predicted_Confidence": log_pred_conf
})
log_pred_df["Correct"] = log_pred_df["True_Label"] == log_pred_df["Predicted_Label"]

display(log_pred_df.head(15))
True_Label Predicted_Label Predicted_Confidence Correct
0 FINANCE FINANCE 0.201206 True
1 COMMUNICATION FINANCE 0.127063 False
2 LIBRARIES_AND_DEMO EDUCATION 0.134790 False
3 LIFESTYLE EDUCATION 0.146064 False
4 EDUCATION EDUCATION 0.136238 True
5 ART_AND_DESIGN LIFESTYLE 0.143387 False
6 BEAUTY EDUCATION 0.125436 False
7 DATING DATING 0.865158 True
8 HOUSE_AND_HOME LIFESTYLE 0.120546 False
9 FINANCE HEALTH_AND_FITNESS 0.274936 False
10 HEALTH_AND_FITNESS GAME 0.407868 False
11 EDUCATION EDUCATION 0.143047 True
12 HEALTH_AND_FITNESS HEALTH_AND_FITNESS 0.198159 True
13 FINANCE EDUCATION 0.136472 False
14 FINANCE FINANCE 0.180731 True

KNN

In [161]:
best_k = 1
best_f1 = 0.0
In [162]:
test_error_rates = []
k_candidates = list(range(1, 30))  # 1..29

for k in k_candidates:
    knn_model = KNeighborsClassifier(n_neighbors=k)
    knn_model.fit(scaled_X_train, y_train)
    y_pred_test_k = knn_model.predict(scaled_X_test)

    # Elbow metric based on test error (as requested)
    test_error = 1 - accuracy_score(y_test, y_pred_test_k)
    test_error_rates.append(test_error)

    # Keep track of best k by validation weighted F1 (good model selection practice)
    y_pred_val_k = knn_model.predict(scaled_X_val)
    current_f1 = f1_score(y_val, y_pred_val_k, average='weighted')
    if current_f1 > best_f1:
        best_f1 = current_f1
        best_k = k
In [163]:
knn_probs = knn_final.predict_proba(scaled_X_test)
knn_pred_idx = np.argmax(knn_probs, axis=1)
knn_pred_conf = knn_probs[np.arange(len(knn_probs)), knn_pred_idx]

knn_pred_df = pd.DataFrame({
    "True_Label": y_test.reset_index(drop=True),
    "Predicted_Label": pd.Series(y_test_pred_knn).reset_index(drop=True),
    "Predicted_Confidence": knn_pred_conf
})
knn_pred_df["Correct"] = knn_pred_df["True_Label"] == knn_pred_df["Predicted_Label"]

display(knn_pred_df.head(15))
True_Label Predicted_Label Predicted_Confidence Correct
0 FINANCE BEAUTY 0.333333 False
1 COMMUNICATION ART_AND_DESIGN 0.333333 False
2 LIBRARIES_AND_DEMO COMMUNICATION 0.333333 False
3 LIFESTYLE BUSINESS 0.333333 False
4 EDUCATION GAME 0.666667 False
5 ART_AND_DESIGN AUTO_AND_VEHICLES 0.333333 False
6 BEAUTY BEAUTY 0.333333 True
7 DATING DATING 1.000000 True
8 HOUSE_AND_HOME AUTO_AND_VEHICLES 0.333333 False
9 FINANCE HEALTH_AND_FITNESS 0.666667 False
10 HEALTH_AND_FITNESS FINANCE 0.333333 False
11 EDUCATION HEALTH_AND_FITNESS 0.666667 False
12 HEALTH_AND_FITNESS EDUCATION 0.666667 False
13 FINANCE ART_AND_DESIGN 0.333333 False
14 FINANCE FINANCE 1.000000 True
In [164]:
best_acc_idx = int(np.argmax(knn_val_acc_scores))
best_f1_idx = int(np.argmax(knn_val_f1_scores))

plt.figure(figsize=(9, 5))
plt.plot(k_values, knn_val_acc_scores, marker='o', linewidth=2, label='Validation Accuracy')
plt.plot(k_values, knn_val_f1_scores, marker='s', linewidth=2, label='Validation Weighted F1')

# Highlight best metric points
plt.scatter(
    k_values[best_acc_idx],
    knn_val_acc_scores[best_acc_idx],
    s=120,
    color='green',
    zorder=5,
    label=f'Best Accuracy (k={k_values[best_acc_idx]})'
)
plt.scatter(
    k_values[best_f1_idx],
    knn_val_f1_scores[best_f1_idx],
    s=120,
    color='red',
    zorder=5,
    label=f'Best Weighted F1 (k={k_values[best_f1_idx]})'
)

plt.axvline(best_k, color='red', linestyle='--', alpha=0.5, label=f'Selected k = {best_k}')

plt.title('KNN Validation Performance vs K (Dynamic)')
plt.xlabel('Number of Neighbors (k)')
plt.ylabel('Score')
plt.xticks(k_values)
plt.grid(alpha=0.3)
plt.legend()
plt.show()
In [165]:
plt.figure(figsize=(10, 6), dpi=120)
plt.plot(k_candidates, test_error_rates, label='test error')
plt.xlabel('k value')
plt.title('Elbow method for KNN')
plt.legend()
plt.grid(alpha=0.3)
plt.show()
In [166]:
knn_final = KNeighborsClassifier(n_neighbors=best_k)
knn_final.fit(scaled_X_train, y_train)
Out [166]:
KNeighborsClassifier(n_neighbors=3)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
n_neighbors n_neighbors: int, default=5

Number of neighbors to use by default for :meth:`kneighbors` queries.
3
weights weights: {'uniform', 'distance'}, callable or None, default='uniform'

Weight function used in prediction. Possible values:

- 'uniform' : uniform weights. All points in each neighborhood
are weighted equally.
- 'distance' : weight points by the inverse of their distance.
in this case, closer neighbors of a query point will have a
greater influence than neighbors which are further away.
- [callable] : a user-defined function which accepts an
array of distances, and returns an array of the same shape
containing the weights.

Refer to the example entitled
:ref:`sphx_glr_auto_examples_neighbors_plot_classification.py`
showing the impact of the `weights` parameter on the decision
boundary.
'uniform'
algorithm algorithm: {'auto', 'ball_tree', 'kd_tree', 'brute'}, default='auto'

Algorithm used to compute the nearest neighbors:

- 'ball_tree' will use :class:`BallTree`
- 'kd_tree' will use :class:`KDTree`
- 'brute' will use a brute-force search.
- 'auto' will attempt to decide the most appropriate algorithm
based on the values passed to :meth:`fit` method.

Note: fitting on sparse input will override the setting of
this parameter, using brute force.
'auto'
leaf_size leaf_size: int, default=30

Leaf size passed to BallTree or KDTree. This can affect the
speed of the construction and query, as well as the memory
required to store the tree. The optimal value depends on the
nature of the problem.
30
p p: float, default=2

Power parameter for the Minkowski metric. When p = 1, this is equivalent
to using manhattan_distance (l1), and euclidean_distance (l2) for p = 2.
For arbitrary p, minkowski_distance (l_p) is used. This parameter is expected
to be positive.
2
metric metric: str or callable, default='minkowski'

Metric to use for distance computation. Default is "minkowski", which
results in the standard Euclidean distance when p = 2. See the
documentation of `scipy.spatial.distance
`_ and
the metrics listed in
:class:`~sklearn.metrics.pairwise.distance_metrics` for valid metric
values.

If metric is "precomputed", X is assumed to be a distance matrix and
must be square during fit. X may be a :term:`sparse graph`, in which
case only "nonzero" elements may be considered neighbors.

If metric is a callable function, it takes two arrays representing 1D
vectors as inputs and must return one value indicating the distance
between those vectors. This works for Scipy's metrics, but is less
efficient than passing the metric name as a string.
'minkowski'
metric_params metric_params: dict, default=None

Additional keyword arguments for the metric function.
None
n_jobs n_jobs: int, default=None

The number of parallel jobs to run for neighbors search.
``None`` means 1 unless in a :obj:`joblib.parallel_backend` context.
``-1`` means using all processors. See :term:`Glossary `
for more details.
Doesn't affect :meth:`fit` method.
None
In [167]:
y_val_pred_knn = knn_final.predict(scaled_X_val)
knn_val_acc, knn_val_f1 = eval_metrics(y_val, y_val_pred_knn, f"KNN k={best_k} (Validation)")
KNN k=3 (Validation) -> Accuracy: 0.3288 | Weighted F1: 0.3245
In [168]:
y_test_pred_knn = knn_final.predict(scaled_X_test)
knn_test_acc, knn_test_f1 = eval_metrics(y_test, y_test_pred_knn, f"KNN k={best_k} (Test)")
KNN k=3 (Test) -> Accuracy: 0.2523 | Weighted F1: 0.2288
In [169]:
results_df = pd.DataFrame([
    {"Model": "Logistic Regression", "Split": "Validation", "Accuracy": log_val_acc, "Weighted F1": log_val_f1},
    {"Model": "Logistic Regression", "Split": "Test",       "Accuracy": log_test_acc, "Weighted F1": log_test_f1},
    {"Model": f"KNN (k={best_k})",   "Split": "Validation", "Accuracy": knn_val_acc, "Weighted F1": knn_val_f1},
    {"Model": f"KNN (k={best_k})",   "Split": "Test",       "Accuracy": knn_test_acc, "Weighted F1": knn_test_f1},
])

results_df
Out [169]:
Model Split Accuracy Weighted F1
0 Logistic Regression Validation 0.333333 0.306923
1 Logistic Regression Test 0.310811 0.264609
2 KNN (k=3) Validation 0.328829 0.324541
3 KNN (k=3) Test 0.252252 0.228784
In [170]:
top_n = 10
top_classes = y_test.value_counts().head(top_n).index.tolist()

mask = y_test.isin(top_classes)
y_test_top = y_test[mask]
y_pred_top = pd.Series(y_test_pred_knn, index=y_test.index)[mask]

cm_top = confusion_matrix(y_test_top, y_pred_top, labels=top_classes)
disp_top = ConfusionMatrixDisplay(confusion_matrix=cm_top, display_labels=top_classes)

fig, ax = plt.subplots(figsize=(8, 6))
disp_top.plot(
    ax=ax,
    cmap='viridis',
    xticks_rotation=45,
    values_format='d',
    colorbar=False
)
ax.set_title(f'KNN Confusion Matrix (Top {top_n} Classes, k={best_k})')
plt.tight_layout()
plt.show()

Part B

Q: Using (Rating + Reviews + Category + Size in Bytes + Installs_Num), using Logistic regression and KNN, find and discuss the best classification model to predict “Content Rating” (use the training/validation/test partition without cross-validation). [7 marks]

In [170]:

Part C

Q: By considering Installs_Num as categorical feature and using (Rating + Reviews + Category + Size in Bytes + Content Rating), find and discuss the best classification model to predict “Installs_Num” (using Logistic regression and KNN) (use the training/test partition without cross-validation) [8 marks]

In [170]: