Files
2026-04-26 16:01:45 +01:00

775 KiB

This is the Q2 Notebook!

It's tracked via GitHub! hence the need for this line for the init commit

In [58]:
import pandas as pd
import numpy as np
from fontTools.diff import color
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.neighbors import KNeighborsClassifier
from sklearn.metrics import accuracy_score, f1_score, confusion_matrix, ConfusionMatrixDisplay
import matplotlib.pyplot as plt
In [59]:
seed = 101
In [60]:
df = pd.read_csv('data/googleplaystore_new_new.csv')

Part A

Q: A. Using (Rating + Reviews + Content Rating + Size in Bytes + Installs_Num), using Logistic regression and KNN, find and discuss the best classification model to predict “Category” (use the training/validation/test partition without cross-validation). [8 marks]

In [61]:
columns_to_keep = ['Rating', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Category']
df_part_a = df[columns_to_keep]
In [62]:
X_raw = df_part_a.drop('Category', axis=1)
y = df_part_a['Category']
In [63]:
X = pd.get_dummies(X_raw, columns=['Content Rating'])
In [64]:
X_temp, X_test, y_temp, y_test = train_test_split(X, y, test_size=0.2, random_state=seed)
In [65]:
X_train, X_val, y_train, y_val = train_test_split(X_temp, y_temp, test_size=0.25, random_state=seed)
In [66]:
scaler = StandardScaler()
scaled_X_train = scaler.fit_transform(X_train)
scaled_X_val = scaler.transform(X_val)
scaled_X_test = scaler.transform(X_test)
In [67]:
def eval_metrics(y_true, y_pred, label="Model"):
    acc = accuracy_score(y_true, y_pred)
    f1_w = f1_score(y_true, y_pred, average='weighted')
    return acc, f1_w

def split_overview(train_rows, test_rows, val_rows=None):
    rows = [
        {'Split': 'Train', 'Rows': train_rows},
        {'Split': 'Test', 'Rows': test_rows},
    ]
    if val_rows is not None:
        rows.insert(1, {'Split': 'Validation', 'Rows': val_rows})
    return pd.DataFrame(rows)

def prediction_sheet(y_true, y_pred, y_prob, prefix, head_n=20):
    pred_idx = np.argmax(y_prob, axis=1)
    pred_conf = y_prob[np.arange(len(y_prob)), pred_idx]
    sheet = pd.DataFrame({
        f'True_{prefix}': pd.Series(y_true).reset_index(drop=True),
        f'Predicted_{prefix}': pd.Series(y_pred).reset_index(drop=True),
        'Predicted_Confidence': pred_conf
    })
    sheet['Correct'] = sheet[f'True_{prefix}'] == sheet[f'Predicted_{prefix}']
    return sheet.head(head_n)
In [68]:
split_overview(len(X_train), len(X_test), len(X_val))
Out [68]:
Split Rows
0 Train 663
1 Validation 222
2 Test 222
In [68]:

Logistic Regression

In [69]:
# Logistic Regression (tuned a bit for better performance on imbalanced classes)
log_model = LogisticRegression(
    max_iter=5000,
    random_state=seed,
    class_weight='balanced',
    C=2.0,
    solver='lbfgs'
)
log_model.fit(scaled_X_train, y_train)
Out [69]:
LogisticRegression(C=2.0, class_weight='balanced', max_iter=5000,
                   random_state=101)
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
penalty penalty: {'l1', 'l2', 'elasticnet', None}, default='l2'

Specify the norm of the penalty:

- `None`: no penalty is added;
- `'l2'`: add a L2 penalty term and it is the default choice;
- `'l1'`: add a L1 penalty term;
- `'elasticnet'`: both L1 and L2 penalty terms are added.

.. warning::
Some penalties may not work with some solvers. See the parameter
`solver` below, to know the compatibility between the penalty and
solver.

.. versionadded:: 0.19
l1 penalty with SAGA solver (allowing 'multinomial' + L1)

.. deprecated:: 1.8
`penalty` was deprecated in version 1.8 and will be removed in 1.10.
Use `l1_ratio` instead. `l1_ratio=0` for `penalty='l2'`, `l1_ratio=1` for
`penalty='l1'` and `l1_ratio` set to any float between 0 and 1 for
`'penalty='elasticnet'`.
'deprecated'
C C: float, default=1.0

Inverse of regularization strength; must be a positive float.
Like in support vector machines, smaller values specify stronger
regularization. `C=np.inf` results in unpenalized logistic regression.
For a visual example on the effect of tuning the `C` parameter
with an L1 penalty, see:
:ref:`sphx_glr_auto_examples_linear_model_plot_logistic_path.py`.
2.0
l1_ratio l1_ratio: float, default=0.0

The Elastic-Net mixing parameter, with `0 <= l1_ratio <= 1`. Setting
`l1_ratio=1` gives a pure L1-penalty, setting `l1_ratio=0` a pure L2-penalty.
Any value between 0 and 1 gives an Elastic-Net penalty of the form
`l1_ratio * L1 + (1 - l1_ratio) * L2`.

.. warning::
Certain values of `l1_ratio`, i.e. some penalties, may not work with some
solvers. See the parameter `solver` below, to know the compatibility between
the penalty and solver.

.. versionchanged:: 1.8
Default value changed from None to 0.0.

.. deprecated:: 1.8
`None` is deprecated and will be removed in version 1.10. Always use
`l1_ratio` to specify the penalty type.
0.0
dual dual: bool, default=False

Dual (constrained) or primal (regularized, see also
:ref:`this equation `) formulation. Dual formulation
is only implemented for l2 penalty with liblinear solver. Prefer `dual=False`
when n_samples > n_features.
False
tol tol: float, default=1e-4

Tolerance for stopping criteria.
0.0001
fit_intercept fit_intercept: bool, default=True

Specifies if a constant (a.k.a. bias or intercept) should be
added to the decision function.
True
intercept_scaling intercept_scaling: float, default=1

Useful only when the solver `liblinear` is used
and `self.fit_intercept` is set to `True`. In this case, `x` becomes
`[x, self.intercept_scaling]`,
i.e. a "synthetic" feature with constant value equal to
`intercept_scaling` is appended to the instance vector.
The intercept becomes
``intercept_scaling * synthetic_feature_weight``.

.. note::
The synthetic feature weight is subject to L1 or L2
regularization as all other features.
To lessen the effect of regularization on synthetic feature weight
(and therefore on the intercept) `intercept_scaling` has to be increased.
1
class_weight class_weight: dict or 'balanced', default=None

Weights associated with classes in the form ``{class_label: weight}``.
If not given, all classes are supposed to have weight one.

The "balanced" mode uses the values of y to automatically adjust
weights inversely proportional to class frequencies in the input data
as ``n_samples / (n_classes * np.bincount(y))``.

Note that these weights will be multiplied with sample_weight (passed
through the fit method) if sample_weight is specified.

.. versionadded:: 0.17
*class_weight='balanced'*
'balanced'
random_state random_state: int, RandomState instance, default=None

Used when ``solver`` == 'sag', 'saga' or 'liblinear' to shuffle the
data. See :term:`Glossary ` for details.
101
solver solver: {'lbfgs', 'liblinear', 'newton-cg', 'newton-cholesky', 'sag', 'saga'}, default='lbfgs'

Algorithm to use in the optimization problem. Default is 'lbfgs'.
To choose a solver, you might want to consider the following aspects:

- 'lbfgs' is a good default solver because it works reasonably well for a wide
class of problems.
- For :term:`multiclass` problems (`n_classes >= 3`), all solvers except
'liblinear' minimize the full multinomial loss, 'liblinear' will raise an
error.
- 'newton-cholesky' is a good choice for
`n_samples` >> `n_features * n_classes`, especially with one-hot encoded
categorical features with rare categories. Be aware that the memory usage
of this solver has a quadratic dependency on `n_features * n_classes`
because it explicitly computes the full Hessian matrix.
- For small datasets, 'liblinear' is a good choice, whereas 'sag'
and 'saga' are faster for large ones;
- 'liblinear' can only handle binary classification by default. To apply a
one-versus-rest scheme for the multiclass setting one can wrap it with the
:class:`~sklearn.multiclass.OneVsRestClassifier`.

.. warning::
The choice of the algorithm depends on the penalty chosen (`l1_ratio=0`
for L2-penalty, `l1_ratio=1` for L1-penalty and `0 < l1_ratio < 1` for
Elastic-Net) and on (multinomial) multiclass support:

================= ======================== ======================
solver l1_ratio multinomial multiclass
================= ======================== ======================
'lbfgs' l1_ratio=0 yes
'liblinear' l1_ratio=1 or l1_ratio=0 no
'newton-cg' l1_ratio=0 yes
'newton-cholesky' l1_ratio=0 yes
'sag' l1_ratio=0 yes
'saga' 0<=l1_ratio<=1 yes
================= ======================== ======================

.. note::
'sag' and 'saga' fast convergence is only guaranteed on features
with approximately the same scale. You can preprocess the data with
a scaler from :mod:`sklearn.preprocessing`.

.. seealso::
Refer to the :ref:`User Guide ` for more
information regarding :class:`LogisticRegression` and more specifically the
:ref:`Table `
summarizing solver/penalty supports.

.. versionadded:: 0.17
Stochastic Average Gradient (SAG) descent solver. Multinomial support in
version 0.18.
.. versionadded:: 0.19
SAGA solver.
.. versionchanged:: 0.22
The default solver changed from 'liblinear' to 'lbfgs' in 0.22.
.. versionadded:: 1.2
newton-cholesky solver. Multinomial support in version 1.6.
'lbfgs'
max_iter max_iter: int, default=100

Maximum number of iterations taken for the solvers to converge.
5000
verbose verbose: int, default=0

For the liblinear and lbfgs solvers set verbose to any positive
number for verbosity.
0
warm_start warm_start: bool, default=False

When set to True, reuse the solution of the previous call to fit as
initialization, otherwise, just erase the previous solution.
Useless for liblinear solver. See :term:`the Glossary `.

.. versionadded:: 0.17
*warm_start* to support *lbfgs*, *newton-cg*, *sag*, *saga* solvers.
False
n_jobs n_jobs: int, default=None

Does not have any effect.

.. deprecated:: 1.8
`n_jobs` is deprecated in version 1.8 and will be removed in 1.10.
None
In [70]:
y_val_pred_log = log_model.predict(scaled_X_val)
log_val_acc, log_val_f1 = eval_metrics(y_val, y_val_pred_log, "LogReg (Validation)")
In [71]:
y_test_pred_log = log_model.predict(scaled_X_test)
log_test_acc, log_test_f1 = eval_metrics(y_test, y_test_pred_log, "LogReg (Test)")
In [72]:
log_probs = log_model.predict_proba(scaled_X_test)
log_pred_df = prediction_sheet(y_test, y_test_pred_log, log_probs, 'Category', head_n=15)
log_pred_df
Out [72]:
True_Category Predicted_Category Predicted_Confidence Correct
0 FINANCE FINANCE 0.144609 True
1 COMMUNICATION BUSINESS 0.094501 False
2 LIBRARIES_AND_DEMO BEAUTY 0.086577 False
3 LIFESTYLE EVENTS 0.105029 False
4 EDUCATION EVENTS 0.087598 False
5 ART_AND_DESIGN HOUSE_AND_HOME 0.189590 False
6 BEAUTY ART_AND_DESIGN 0.102613 False
7 DATING DATING 0.699165 True
8 HOUSE_AND_HOME HOUSE_AND_HOME 0.131807 True
9 FINANCE HEALTH_AND_FITNESS 0.183727 False
10 HEALTH_AND_FITNESS GAME 0.348999 False
11 EDUCATION ART_AND_DESIGN 0.112001 False
12 HEALTH_AND_FITNESS HEALTH_AND_FITNESS 0.118400 True
13 FINANCE BEAUTY 0.088732 False
14 FINANCE FINANCE 0.124280 True

KNN

In [73]:
best_k = 1
best_f1 = 0.0

test_error_rates = []
knn_val_acc_scores = []
knn_val_f1_scores = []

k_candidates = list(range(1, 51))  # 1..50

for k in k_candidates:
    knn_model = KNeighborsClassifier(n_neighbors=k, weights='distance')
    knn_model.fit(scaled_X_train, y_train)

    # test error (for elbow)
    y_pred_test_k = knn_model.predict(scaled_X_test)
    test_error_rates.append(1 - accuracy_score(y_test, y_pred_test_k))

    # validation metrics (for selection + plot)
    y_pred_val_k = knn_model.predict(scaled_X_val)
    val_acc_k = accuracy_score(y_val, y_pred_val_k)
    val_f1_k = f1_score(y_val, y_pred_val_k, average='weighted')

    knn_val_acc_scores.append(val_acc_k)
    knn_val_f1_scores.append(val_f1_k)

    if val_f1_k > best_f1:
        best_f1 = val_f1_k
        best_k = k
In [74]:
plt.figure(figsize=(10, 6), dpi=120)
plt.plot(k_candidates, test_error_rates, label='test error')
plt.xlabel('k value')
plt.title('Elbow method for KNN')
plt.legend()
plt.grid(alpha=0.3)
plt.show()
In [75]:
knn_final = KNeighborsClassifier(n_neighbors=best_k, weights='distance')
knn_final.fit(scaled_X_train, y_train)
Out [75]:
KNeighborsClassifier(n_neighbors=19, weights='distance')
In a Jupyter environment, please rerun this cell to show the HTML representation or trust the notebook.
On GitHub, the HTML representation is unable to render, please try loading this page with nbviewer.org.
Parameters
n_neighbors n_neighbors: int, default=5

Number of neighbors to use by default for :meth:`kneighbors` queries.
19
weights weights: {'uniform', 'distance'}, callable or None, default='uniform'

Weight function used in prediction. Possible values:

- 'uniform' : uniform weights. All points in each neighborhood
are weighted equally.
- 'distance' : weight points by the inverse of their distance.
in this case, closer neighbors of a query point will have a
greater influence than neighbors which are further away.
- [callable] : a user-defined function which accepts an
array of distances, and returns an array of the same shape
containing the weights.

Refer to the example entitled
:ref:`sphx_glr_auto_examples_neighbors_plot_classification.py`
showing the impact of the `weights` parameter on the decision
boundary.
'distance'
algorithm algorithm: {'auto', 'ball_tree', 'kd_tree', 'brute'}, default='auto'

Algorithm used to compute the nearest neighbors:

- 'ball_tree' will use :class:`BallTree`
- 'kd_tree' will use :class:`KDTree`
- 'brute' will use a brute-force search.
- 'auto' will attempt to decide the most appropriate algorithm
based on the values passed to :meth:`fit` method.

Note: fitting on sparse input will override the setting of
this parameter, using brute force.
'auto'
leaf_size leaf_size: int, default=30

Leaf size passed to BallTree or KDTree. This can affect the
speed of the construction and query, as well as the memory
required to store the tree. The optimal value depends on the
nature of the problem.
30
p p: float, default=2

Power parameter for the Minkowski metric. When p = 1, this is equivalent
to using manhattan_distance (l1), and euclidean_distance (l2) for p = 2.
For arbitrary p, minkowski_distance (l_p) is used. This parameter is expected
to be positive.
2
metric metric: str or callable, default='minkowski'

Metric to use for distance computation. Default is "minkowski", which
results in the standard Euclidean distance when p = 2. See the
documentation of `scipy.spatial.distance
`_ and
the metrics listed in
:class:`~sklearn.metrics.pairwise.distance_metrics` for valid metric
values.

If metric is "precomputed", X is assumed to be a distance matrix and
must be square during fit. X may be a :term:`sparse graph`, in which
case only "nonzero" elements may be considered neighbors.

If metric is a callable function, it takes two arrays representing 1D
vectors as inputs and must return one value indicating the distance
between those vectors. This works for Scipy's metrics, but is less
efficient than passing the metric name as a string.
'minkowski'
metric_params metric_params: dict, default=None

Additional keyword arguments for the metric function.
None
n_jobs n_jobs: int, default=None

The number of parallel jobs to run for neighbors search.
``None`` means 1 unless in a :obj:`joblib.parallel_backend` context.
``-1`` means using all processors. See :term:`Glossary `
for more details.
Doesn't affect :meth:`fit` method.
None
In [76]:
y_val_pred_knn = knn_final.predict(scaled_X_val)
knn_val_acc, knn_val_f1 = eval_metrics(y_val, y_val_pred_knn, f"KNN k={best_k} (Validation)")
In [77]:
y_test_pred_knn = knn_final.predict(scaled_X_test)
knn_test_acc, knn_test_f1 = eval_metrics(y_test, y_test_pred_knn, f"KNN k={best_k} (Test)")
In [78]:
knn_probs = knn_final.predict_proba(scaled_X_test)
knn_pred_df = prediction_sheet(y_test, y_test_pred_knn, knn_probs, 'Category', head_n=15)
knn_pred_df
Out [78]:
True_Category Predicted_Category Predicted_Confidence Correct
0 FINANCE FINANCE 0.322858 True
1 COMMUNICATION FINANCE 0.149371 False
2 LIBRARIES_AND_DEMO HEALTH_AND_FITNESS 0.675556 False
3 LIFESTYLE FINANCE 0.220342 False
4 EDUCATION GAME 0.982053 False
5 ART_AND_DESIGN LIFESTYLE 0.696031 False
6 BEAUTY BEAUTY 0.331699 True
7 DATING DATING 0.921884 True
8 HOUSE_AND_HOME LIFESTYLE 0.187210 False
9 FINANCE HEALTH_AND_FITNESS 0.577385 False
10 HEALTH_AND_FITNESS GAME 0.308132 False
11 EDUCATION HEALTH_AND_FITNESS 0.346198 False
12 HEALTH_AND_FITNESS HEALTH_AND_FITNESS 0.185119 True
13 FINANCE ART_AND_DESIGN 0.198243 False
14 FINANCE FINANCE 0.365431 True
In [79]:
best_acc_idx = int(np.argmax(knn_val_acc_scores))
best_f1_idx = int(np.argmax(knn_val_f1_scores))

plt.figure(figsize=(9, 5))
plt.plot(k_candidates, knn_val_acc_scores, marker='o', linewidth=2, label='Validation Accuracy')
plt.plot(k_candidates, knn_val_f1_scores, marker='s', linewidth=2, label='Validation Weighted F1')

plt.scatter(
    k_candidates[best_acc_idx],
    knn_val_acc_scores[best_acc_idx],
    s=120, color='green', zorder=5,
    label=f'Best Accuracy (k={k_candidates[best_acc_idx]})'
)
plt.scatter(
    k_candidates[best_f1_idx],
    knn_val_f1_scores[best_f1_idx],
    s=120, color='red', zorder=5,
    label=f'Best Weighted F1 (k={k_candidates[best_f1_idx]})'
)

plt.axvline(best_k, color='red', linestyle='--', alpha=0.5, label=f'Selected k = {best_k}')
plt.title('KNN Validation Performance vs K (Dynamic)')
plt.xlabel('Number of Neighbors (k)')
plt.ylabel('Score')
plt.xticks(range(1, 51, 2))
plt.grid(alpha=0.3)
plt.legend()
plt.show()
In [80]:
results_df = pd.DataFrame([
    {"Model": "Logistic Regression", "Split": "Validation", "Accuracy": log_val_acc, "Weighted F1": log_val_f1},
    {"Model": "Logistic Regression", "Split": "Test",       "Accuracy": log_test_acc, "Weighted F1": log_test_f1},
    {"Model": f"KNN (k={best_k})",   "Split": "Validation", "Accuracy": knn_val_acc, "Weighted F1": knn_val_f1},
    {"Model": f"KNN (k={best_k})",   "Split": "Test",       "Accuracy": knn_test_acc, "Weighted F1": knn_test_f1},
])
In [81]:
results_df
Out [81]:
Model Split Accuracy Weighted F1
0 Logistic Regression Validation 0.301802 0.302937
1 Logistic Regression Test 0.270270 0.262607
2 KNN (k=19) Validation 0.360360 0.354603
3 KNN (k=19) Test 0.310811 0.290187
In [82]:
top_n = 10
top_classes = y_test.value_counts().head(top_n).index.tolist()

mask = y_test.isin(top_classes)
y_test_top = y_test[mask]
y_log_top = pd.Series(y_test_pred_log, index=y_test.index)[mask]
y_knn_top = pd.Series(y_test_pred_knn, index=y_test.index)[mask]

cm_log_top = confusion_matrix(y_test_top, y_log_top, labels=top_classes)
cm_knn_top = confusion_matrix(y_test_top, y_knn_top, labels=top_classes)

fig, ax = plt.subplots(1, 2, figsize=(15, 6))
ConfusionMatrixDisplay(confusion_matrix=cm_log_top, display_labels=top_classes).plot(
    ax=ax[0],
    cmap='Blues',
    xticks_rotation=45,
    values_format='d',
    colorbar=False
)
ax[0].set_title(f'Part A: Logistic Regression CM (Top {top_n} classes)')

ConfusionMatrixDisplay(confusion_matrix=cm_knn_top, display_labels=top_classes).plot(
    ax=ax[1],
    cmap='Greens',
    xticks_rotation=45,
    values_format='d',
    colorbar=False
)
ax[1].set_title(f'Part A: KNN CM (Top {top_n} classes, k={best_k})')

plt.tight_layout()
plt.show()
In [83]:
actual_counts_a = y_test.value_counts().reindex(top_classes, fill_value=0)
log_pred_counts_a = pd.Series(y_test_pred_log).value_counts().reindex(top_classes, fill_value=0)
knn_pred_counts_a = pd.Series(y_test_pred_knn).value_counts().reindex(top_classes, fill_value=0)

x = np.arange(len(top_classes))
width = 0.26

fig, ax = plt.subplots(figsize=(13, 5))
ax.bar(x - width, actual_counts_a.values, width, label='Actual (Test)')
ax.bar(x,         log_pred_counts_a.values, width, label='LogReg Predictions')
ax.bar(x + width, knn_pred_counts_a.values, width, label=f'KNN Predictions (k={best_k})')
ax.set_xticks(x)
ax.set_xticklabels(top_classes, rotation=45, ha='right')
ax.set_ylabel('Count')
ax.set_title(f'Part A: Actual vs Predicted Category Distribution (Top {top_n} classes)')
ax.grid(True, axis='y', alpha=0.3)
ax.legend()
plt.tight_layout()
plt.show()

Part B

Q: Using (Rating + Reviews + Category + Size in Bytes + Installs_Num), using Logistic regression and KNN, find and discuss the best classification model to predict “Content Rating” (use the training/validation/test partition without cross-validation). [7 marks]

For this part, I'll be training Logistic Regression and KNN using the inputs mentioned above. I'll compare how they perform using Accuracy, Weighted F1, and Confusion Matrices. I also added a table showing the raw predicted outputs from each model on the test set so it's clear what the models are actually predicting.

In [84]:
cols_part_b = [
    'Rating',
    'Reviews',
    'Category',
    'Size in bytes',
    'Numeric Installs',
    'Content Rating'
]

df_part_b = df[cols_part_b].dropna().copy()
In [85]:
app_inputs_b = df_part_b.drop('Content Rating', axis=1)
true_content_rating = df_part_b['Content Rating']
app_inputs_b_encoded = pd.get_dummies(app_inputs_b, columns=['Category'])
In [86]:
class_counts_b = true_content_rating.value_counts()
valid_mask_b = true_content_rating.map(class_counts_b) >= 2

X_b = app_inputs_b_encoded.loc[valid_mask_b]
y_b = true_content_rating.loc[valid_mask_b]

dropped_classes_b = class_counts_b[class_counts_b < 2]

X_temp_b, X_test_b, y_temp_b, y_test_b = train_test_split(
    X_b, y_b, test_size=0.2, random_state=seed, stratify=y_b
)

stratify_second_b = y_temp_b if y_temp_b.value_counts().min() >= 2 else None
part_b_fallback = stratify_second_b is None

X_train_b, X_val_b, y_train_b, y_val_b = train_test_split(
    X_temp_b, y_temp_b, test_size=0.25, random_state=seed, stratify=stratify_second_b
)

part_b_split_info = pd.DataFrame([
    {'Split': 'Train', 'Rows': len(X_train_b)},
    {'Split': 'Validation', 'Rows': len(X_val_b)},
    {'Split': 'Test', 'Rows': len(X_test_b)},
])
part_b_split_info
Out [86]:
Split Rows
0 Train 663
1 Validation 221
2 Test 222
In [87]:
scaler_b = StandardScaler()
X_train_b_scaled = scaler_b.fit_transform(X_train_b)
X_val_b_scaled = scaler_b.transform(X_val_b)
X_test_b_scaled = scaler_b.transform(X_test_b)

display(pd.DataFrame([{
    'Dropped rare classes': int(dropped_classes_b.shape[0]),
    'Second split fallback used': part_b_fallback
}]))
Dropped rare classes Second split fallback used
0 1 False

Logistic Regression

In [88]:
logreg_b = LogisticRegression(
    max_iter=5000,
    random_state=seed,
    class_weight='balanced',
    C=2.0,
    solver='lbfgs'
)
logreg_b.fit(X_train_b_scaled, y_train_b)

logreg_b_val_pred = logreg_b.predict(X_val_b_scaled)
logreg_b_test_pred = logreg_b.predict(X_test_b_scaled)

logreg_b_val_acc, logreg_b_val_f1 = eval_metrics(y_val_b, logreg_b_val_pred, "LogReg (Part B, Validation)")
logreg_b_test_acc, logreg_b_test_f1 = eval_metrics(y_test_b, logreg_b_test_pred, "LogReg (Part B, Test)")

KNN

In [89]:
k_values_b = list(range(1, 51, 2))
knn_val_acc_scores_b = []
knn_val_f1_scores_b = []
In [90]:
for k in k_values_b:
    knn_b = KNeighborsClassifier(n_neighbors=k, weights='distance')
    knn_b.fit(X_train_b_scaled, y_train_b)
    y_val_pred_k = knn_b.predict(X_val_b_scaled)
    knn_val_acc_scores_b.append(accuracy_score(y_val_b, y_val_pred_k))
    knn_val_f1_scores_b.append(f1_score(y_val_b, y_val_pred_k, average='weighted'))

best_k_b = int(k_values_b[int(np.argmax(knn_val_f1_scores_b))])

knn_final_b = KNeighborsClassifier(n_neighbors=best_k_b, weights='distance')
knn_final_b.fit(X_train_b_scaled, y_train_b)

knn_b_val_pred = knn_final_b.predict(X_val_b_scaled)
knn_b_test_pred = knn_final_b.predict(X_test_b_scaled)

knn_b_val_acc, knn_b_val_f1 = eval_metrics(y_val_b, knn_b_val_pred, f"KNN k={best_k_b} (Part B, Validation)")
knn_b_test_acc, knn_b_test_f1 = eval_metrics(y_test_b, knn_b_test_pred, f"KNN k={best_k_b} (Part B, Test)")
In [91]:
comparison_b = pd.DataFrame([
    {"Model": "Logistic Regression", "Split": "Validation", "Accuracy": logreg_b_val_acc, "Weighted F1": logreg_b_val_f1},
    {"Model": "Logistic Regression", "Split": "Test",       "Accuracy": logreg_b_test_acc, "Weighted F1": logreg_b_test_f1},
    {"Model": f"KNN (k={best_k_b})",   "Split": "Validation", "Accuracy": knn_b_val_acc, "Weighted F1": knn_b_val_f1},
    {"Model": f"KNN (k={best_k_b})",   "Split": "Test",       "Accuracy": knn_b_test_acc, "Weighted F1": knn_b_test_f1},
])

display(comparison_b)
Model Split Accuracy Weighted F1
0 Logistic Regression Validation 0.583710 0.659486
1 Logistic Regression Test 0.621622 0.690548
2 KNN (k=15) Validation 0.873303 0.865222
3 KNN (k=15) Test 0.864865 0.856732
In [92]:
fig, ax = plt.subplots(1, 2, figsize=(12, 4))

for i, metric in enumerate(["Accuracy", "Weighted F1"]):
    for model_name in comparison_b["Model"].unique():
        subset = comparison_b[comparison_b["Model"] == model_name]
        ax[i].plot(subset["Split"], subset[metric], marker='o', linewidth=2, label=model_name)
    ax[i].set_title(f"Part B: {metric} (Validation vs Test)")
    ax[i].set_ylim(0, 1)
    ax[i].grid(True, alpha=0.3)
    ax[i].legend()

plt.tight_layout()
plt.show()
In [93]:
fig, ax = plt.subplots(1, 2, figsize=(14, 5))

ConfusionMatrixDisplay.from_predictions(
    y_test_b,
    logreg_b_test_pred,
    ax=ax[0],
    cmap='Blues',
    xticks_rotation=45,
    colorbar=False
)
ax[0].set_title("Part B: Logistic Regression Confusion Matrix (Test)")

ConfusionMatrixDisplay.from_predictions(
    y_test_b,
    knn_b_test_pred,
    ax=ax[1],
    cmap='Greens',
    xticks_rotation=45,
    colorbar=False
)
ax[1].set_title(f"Part B: KNN Confusion Matrix (Test, k={best_k_b})")

plt.tight_layout()
plt.show()

Displaying the predicted data

In [94]:
logreg_b_probs = logreg_b.predict_proba(X_test_b_scaled)
logreg_b_pred_df = prediction_sheet(y_test_b, logreg_b_test_pred, logreg_b_probs, 'Content_Rating', head_n=20)
logreg_b_pred_df
Out [94]:
True_Content_Rating Predicted_Content_Rating Predicted_Confidence Correct
0 Everyone Everyone 0.993702 True
1 Everyone Everyone 0.482609 True
2 Everyone Everyone 0.994017 True
3 Everyone Everyone 10+ 0.522200 False
4 Everyone Teen 0.767573 False
5 Everyone Everyone 0.612729 True
6 Everyone Everyone 0.994448 True
7 Everyone Everyone 0.993697 True
8 Everyone Everyone 10+ 0.530097 False
9 Everyone Everyone 0.749141 True
10 Mature 17+ Mature 17+ 0.940680 True
11 Everyone Everyone 0.991359 True
12 Teen Everyone 0.639161 False
13 Everyone Everyone 0.993842 True
14 Everyone Everyone 0.993532 True
15 Everyone Everyone 10+ 0.480918 False
16 Everyone Everyone 0.760410 True
17 Everyone Everyone 10+ 0.504585 False
18 Everyone Everyone 0.460006 True
19 Everyone Everyone 0.993453 True
In [95]:
knn_b_probs = knn_final_b.predict_proba(X_test_b_scaled)
knn_b_pred_df = prediction_sheet(y_test_b, knn_b_test_pred, knn_b_probs, 'Content_Rating', head_n=20)
knn_b_pred_df
Out [95]:
True_Content_Rating Predicted_Content_Rating Predicted_Confidence Correct
0 Everyone Everyone 1.000000 True
1 Everyone Everyone 1.000000 True
2 Everyone Everyone 1.000000 True
3 Everyone Everyone 0.564860 True
4 Everyone Teen 0.640995 False
5 Everyone Everyone 0.998311 True
6 Everyone Everyone 1.000000 True
7 Everyone Everyone 1.000000 True
8 Everyone Everyone 0.875223 True
9 Everyone Everyone 1.000000 True
10 Mature 17+ Mature 17+ 1.000000 True
11 Everyone Everyone 1.000000 True
12 Teen Teen 0.998994 True
13 Everyone Everyone 1.000000 True
14 Everyone Everyone 1.000000 True
15 Everyone Everyone 0.911639 True
16 Everyone Everyone 1.000000 True
17 Everyone Everyone 0.919221 True
18 Everyone Everyone 0.886516 True
19 Everyone Everyone 1.000000 True
In [96]:
actual_counts_b = pd.Series(y_test_b).value_counts().sort_index()
logreg_pred_counts_b = pd.Series(logreg_b_test_pred).value_counts().reindex(actual_counts_b.index, fill_value=0)
knn_pred_counts_b = pd.Series(knn_b_test_pred).value_counts().reindex(actual_counts_b.index, fill_value=0)

x = np.arange(len(actual_counts_b.index))
width = 0.28

fig, ax = plt.subplots(figsize=(12, 4))
ax.bar(x - width, actual_counts_b.values, width, label="Actual (Test)")
ax.bar(x,         logreg_pred_counts_b.values, width, label="LogReg Predictions")
ax.bar(x + width, knn_pred_counts_b.values, width, label=f"KNN Predictions (k={best_k_b})")
ax.set_xticks(x)
ax.set_xticklabels(actual_counts_b.index, rotation=45, ha='right')
ax.set_title("Part B: Actual vs Predicted Content Rating distribution (Test set)")
ax.set_ylabel("Number of apps")
ax.grid(True, axis='y', alpha=0.3)
ax.legend()
plt.tight_layout()
plt.show()
In [97]:
best_model_b = (
    f"KNN (k={best_k_b})" if knn_b_test_f1 > logreg_b_test_f1 else "Logistic Regression"
)
In [98]:
part_b_final_summary = pd.DataFrame([
    {'Model': 'Logistic Regression', 'Accuracy': logreg_b_test_acc, 'Weighted F1': logreg_b_test_f1},
    {'Model': f'KNN (k={best_k_b})', 'Accuracy': knn_b_test_acc, 'Weighted F1': knn_b_test_f1},
]).sort_values('Weighted F1', ascending=False).reset_index(drop=True)

display(part_b_final_summary)
display(pd.DataFrame([{'Part B Best Model': best_model_b}]))
Model Accuracy Weighted F1
0 KNN (k=15) 0.864865 0.856732
1 Logistic Regression 0.621622 0.690548
Part B Best Model
0 KNN (k=15)

Answer for Part B:

The best model to fit the prediction of the "Category" is the KNN model with it's K being set to 15, it seems to provide the highest accuracy and weighted F1 score on the test set, and also has a better confusion matrix distribution compared to Logistic Regression. The data provided by the KNN model also shows a better distribution of predicted classes that is closer to the actual distribution of the test set.

Therefore, KNN is the best classification model for predicting "Content Rating" in Part B.

Part C

Q: By considering Installs_Num as categorical feature and using (Rating + Reviews + Category + Size in Bytes + Content Rating), find and discuss the best classification model to predict “Installs_Num” (using Logistic regression and KNN) (use the training/test partition without cross-validation) [8 marks]

Here I'm treating Installs_Num as the categorical label we want to predict. I'll train Logistic Regression and KNN using only the requested inputs and compare them on the test set using Accuracy, Weighted F1, and Confusion Matrices. I'll also include a table showing the actual per-row predictions from both models.

In [99]:
cols_part_c = [
    'Rating',
    'Reviews',
    'Category',
    'Size in bytes',
    'Content Rating',
    'Numeric Installs'
]

df_part_c = df[cols_part_c].dropna().copy()
In [100]:
app_inputs_c = df_part_c.drop('Numeric Installs', axis=1)
true_install_bucket = df_part_c['Numeric Installs'].astype(str)
app_inputs_c_encoded = pd.get_dummies(app_inputs_c, columns=['Category', 'Content Rating'])
In [101]:
class_counts_c = true_install_bucket.value_counts()
valid_mask_c = true_install_bucket.map(class_counts_c) >= 2

X_c = app_inputs_c_encoded.loc[valid_mask_c]
y_c = true_install_bucket.loc[valid_mask_c]

dropped_classes_c = class_counts_c[class_counts_c < 2]

X_train_c, X_test_c, y_train_c, y_test_c = train_test_split(
    X_c,
    y_c,
    test_size=0.2,
    random_state=seed,
    stratify=y_c
)

display(pd.DataFrame([
    {'Split': 'Train', 'Rows': len(X_train_c)},
    {'Split': 'Test', 'Rows': len(X_test_c)},
]))
display(pd.DataFrame([{
    'Part C classes after filtering': int(y_c.nunique()),
    'Dropped rare classes': int(dropped_classes_c.shape[0])
}]))
Split Rows
0 Train 884
1 Test 222
Part C classes after filtering Dropped rare classes
0 15 1
In [102]:
scaler_c = StandardScaler()
X_train_c_scaled = scaler_c.fit_transform(X_train_c)
X_test_c_scaled = scaler_c.transform(X_test_c)
display(y_c.value_counts().head(15).rename('Class_Count').reset_index().rename(columns={'index': 'Numeric_Installs'}))
Numeric Installs Class_Count
0 1000000 302
1 100000 191
2 10000000 134
3 500000 125
4 5000000 99
5 10000 66
6 50000 56
7 100000000 51
8 50000000 21
9 1000 16
10 5000 12
11 100 11
12 500000000 10
13 500 9
14 1000000000 3

Logistic Regression

In [103]:
logreg_c = LogisticRegression(
    max_iter=5000,
    random_state=seed,
    class_weight='balanced',
    C=2.0,
    solver='lbfgs'
)
logreg_c.fit(X_train_c_scaled, y_train_c)

logreg_c_test_pred = logreg_c.predict(X_test_c_scaled)
logreg_c_test_acc, logreg_c_test_f1 = eval_metrics(y_test_c, logreg_c_test_pred, "LogReg (Part C, Test)")

KNN

In [104]:
stratify_inner_c = y_train_c if y_train_c.value_counts().min() >= 2 else None
X_train_c_inner, X_val_c_inner, y_train_c_inner, y_val_c_inner = train_test_split(
    X_train_c, y_train_c, test_size=0.25, random_state=seed, stratify=stratify_inner_c
)
display(pd.DataFrame([{'Part C inner split fallback used': stratify_inner_c is None}]))

scaler_c_inner = StandardScaler()
X_train_c_inner_scaled = scaler_c_inner.fit_transform(X_train_c_inner)
X_val_c_inner_scaled = scaler_c_inner.transform(X_val_c_inner)
X_test_c_scaled_inner = scaler_c_inner.transform(X_test_c)
Part C inner split fallback used
0 False
In [105]:
k_values_c = list(range(1, 51, 2))
knn_val_acc_scores_c = []
knn_val_f1_scores_c = []
In [106]:
for k in k_values_c:
    knn_c = KNeighborsClassifier(n_neighbors=k, weights='distance')
    knn_c.fit(X_train_c_inner_scaled, y_train_c_inner)
    y_val_pred_k = knn_c.predict(X_val_c_inner_scaled)
    knn_val_acc_scores_c.append(accuracy_score(y_val_c_inner, y_val_pred_k))
    knn_val_f1_scores_c.append(f1_score(y_val_c_inner, y_val_pred_k, average='weighted'))
In [107]:
best_k_c = int(k_values_c[int(np.argmax(knn_val_f1_scores_c))])
knn_final_c = KNeighborsClassifier(n_neighbors=best_k_c, weights='distance')
knn_final_c.fit(X_train_c_inner_scaled, y_train_c_inner)
knn_c_test_pred = knn_final_c.predict(X_test_c_scaled_inner)
knn_c_test_acc, knn_c_test_f1 = eval_metrics(y_test_c, knn_c_test_pred, f"KNN k={best_k_c} (Part C, Test)")

Visual comparison

In [116]:
comparison_c = pd.DataFrame([
    {"Model": "Logistic Regression", "Accuracy": logreg_c_test_acc, "Weighted F1": logreg_c_test_f1},
    {"Model": f"KNN (k={best_k_c})", "Accuracy": knn_c_test_acc, "Weighted F1": knn_c_test_f1},
])

display(comparison_c)
fig, ax = plt.subplots(figsize=(7, 4))
ax.bar(comparison_c["Model"], comparison_c["Weighted F1"], label="Weighted F1")
ax.plot(comparison_c["Model"], comparison_c["Accuracy"], marker='o', linewidth=2, label="Accuracy", color='orange')
ax.set_title("Part C: Model comparison on Test set")
ax.set_ylim(0, 1)
ax.grid(True, axis='y', alpha=0.3)
ax.legend()
plt.tight_layout()
plt.show()
Model Accuracy Weighted F1
0 Logistic Regression 0.171171 0.182992
1 KNN (k=41) 0.324324 0.285951
In [109]:
top_n_c = 12

top_classes_c = pd.Series(y_test_c).value_counts().head(top_n_c).index.tolist()
mask_c = pd.Series(y_test_c).isin(top_classes_c)

y_test_c_top = pd.Series(y_test_c).reset_index(drop=True)[mask_c.reset_index(drop=True)]
logreg_c_pred_top = pd.Series(logreg_c_test_pred).reset_index(drop=True)[mask_c.reset_index(drop=True)]
knn_c_pred_top = pd.Series(knn_c_test_pred).reset_index(drop=True)[mask_c.reset_index(drop=True)]

fig, ax = plt.subplots(1, 2, figsize=(14, 5))

cm_log_top = confusion_matrix(y_test_c_top, logreg_c_pred_top, labels=top_classes_c)
ConfusionMatrixDisplay(confusion_matrix=cm_log_top, display_labels=top_classes_c).plot(
    ax=ax[0], cmap='Blues', xticks_rotation=45, values_format='d', colorbar=False
)
ax[0].set_title(f"Part C: Logistic Regression Confusion Matrix (Top {top_n_c} classes, Test)")

cm_knn_top = confusion_matrix(y_test_c_top, knn_c_pred_top, labels=top_classes_c)
ConfusionMatrixDisplay(confusion_matrix=cm_knn_top, display_labels=top_classes_c).plot(
    ax=ax[1], cmap='Greens', xticks_rotation=45, values_format='d', colorbar=False
)
ax[1].set_title(f"Part C: KNN Confusion Matrix (Top {top_n_c} classes, Test, k={best_k_c})")

plt.tight_layout()
plt.show()

Predicted Data Outputs:

In [110]:
logreg_c_probs = logreg_c.predict_proba(X_test_c_scaled)
logreg_c_pred_df = prediction_sheet(y_test_c, logreg_c_test_pred, logreg_c_probs, 'Installs_Num', head_n=20)
logreg_c_pred_df
Out [110]:
True_Installs_Num Predicted_Installs_Num Predicted_Confidence Correct
0 1000000 5000 0.197939 False
1 1000000 50000 0.469158 False
2 10000000 50000000 0.425556 False
3 500000 1000000 0.145711 False
4 100000000 500000000 0.470446 False
5 100000000 50000000 0.400632 False
6 10000000 50000000 0.492270 False
7 10000000 5000000 0.250524 False
8 10000 500 0.647713 False
9 1000000 1000000 0.171517 True
10 100000 500000 0.334840 False
11 10000 50000 0.212020 False
12 5000000 100 0.291118 False
13 1000000 500000 0.241188 False
14 5000 100 0.207994 False
15 500000 50000 0.244060 False
16 5000000 10000000 0.237635 False
17 10000000 50000000 0.373763 False
18 1000000 1000000 0.196770 True
19 100000 5000000 0.161215 False
In [111]:
knn_c_probs = knn_final_c.predict_proba(X_test_c_scaled_inner)
knn_c_pred_df = prediction_sheet(y_test_c, knn_c_test_pred, knn_c_probs, 'Installs_Num', head_n=20)
knn_c_pred_df
Out [111]:
True_Installs_Num Predicted_Installs_Num Predicted_Confidence Correct
0 1000000 1000000 0.293952 True
1 1000000 1000000 0.446849 True
2 10000000 1000000 0.522130 False
3 500000 1000000 0.287865 False
4 100000000 1000000 0.376263 False
5 100000000 100000000 0.995326 True
6 10000000 10000000 0.459153 True
7 10000000 1000000 0.304257 False
8 10000 100000 0.291979 False
9 1000000 100000 0.427634 False
10 100000 500000 0.259322 False
11 10000 50000 0.322032 False
12 5000000 1000000 0.234653 False
13 1000000 500000 0.431720 False
14 5000 100000 0.377805 False
15 500000 50000 0.431116 False
16 5000000 1000000 0.443130 False
17 10000000 10000000 0.353770 True
18 1000000 100000 0.376758 False
19 100000 1000000 0.595938 False

comparing distribution of predicted labels vs actual labels on the test set

In [112]:
dist_top_n_c = 12
most_common_classes_c = pd.Series(y_test_c).value_counts().head(dist_top_n_c).index

actual_counts_c = pd.Series(y_test_c).value_counts().reindex(most_common_classes_c)
logreg_pred_counts_c = pd.Series(logreg_c_test_pred).value_counts().reindex(most_common_classes_c, fill_value=0)
knn_pred_counts_c = pd.Series(knn_c_test_pred).value_counts().reindex(most_common_classes_c, fill_value=0)

x = np.arange(len(most_common_classes_c))
width = 0.28

fig, ax = plt.subplots(figsize=(12, 4))
ax.bar(x - width, actual_counts_c.values, width, label="Actual (Test)")
ax.bar(x,         logreg_pred_counts_c.values, width, label="LogReg Predictions")
ax.bar(x + width, knn_pred_counts_c.values, width, label=f"KNN Predictions (k={best_k_c})")
ax.set_xticks(x)
ax.set_xticklabels(most_common_classes_c, rotation=45, ha='right')
ax.set_title(f"Part C: Actual vs Predicted Installs_Num distribution (Top {dist_top_n_c} classes, Test set)")
ax.set_ylabel("Number of apps")
ax.grid(True, axis='y', alpha=0.3)
ax.legend()
plt.tight_layout()
plt.show()
In [113]:
best_model_c = (
    f"KNN (k={best_k_c})" if knn_c_test_f1 > logreg_c_test_f1 else "Logistic Regression"
)
In [114]:
part_c_final_summary = pd.DataFrame([
    {'Model': 'Logistic Regression', 'Accuracy': logreg_c_test_acc, 'Weighted F1': logreg_c_test_f1},
    {'Model': f'KNN (k={best_k_c})', 'Accuracy': knn_c_test_acc, 'Weighted F1': knn_c_test_f1},
]).sort_values('Weighted F1', ascending=False).reset_index(drop=True)

display(part_c_final_summary)
display(pd.DataFrame([{'Part C Best Model': best_model_c}]))
Model Accuracy Weighted F1
0 KNN (k=41) 0.324324 0.285951
1 Logistic Regression 0.171171 0.182992
Part C Best Model
0 KNN (k=41)

Answer for Part C:

The KNN Model with the K value being set to 41, provides the closest predicted data distribution to the actual distribution of the test set, and also has a higher weighted F1 score compared to Logistic Regression. Therefore, KNN is the best classification model for predicting "Installs_Num" in Part C.

In [ ]: