Files
UNI-PROG3-CW2-MLWP/Q1.ipynb
T
2026-04-26 15:57:09 +01:00

734 KiB

This is the Q1 Notebook!

It's tracked via GitHub! hence the need for this line for the init commit

In [71]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
In [72]:
seed = 101
In [73]:
df = pd.read_csv('data/googleplaystore_new.csv')

Part A

Q: Check for any missing values. If any, delete all rows that contain null values Then, check for duplicates (similar contents for every feature). If any, delete the redundant rows. [1 mark]

In [74]:
df = df.dropna()
df = df.drop_duplicates()
In [75]:
df.head(10)
Out [75]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up

Part B

Q: Create a new column called 'Size in bytes' (numeric) and convert the entries from 'Size' column (M means megabyte and k means kilobytes). [1 mark]

In [76]:
def parse_size(size_str):
    if isinstance(size_str, str):
        if size_str.endswith('M'):
            return float(size_str[:-1]) * 1024 * 1024
        elif size_str.endswith('k'):
            return float(size_str[:-1]) * 1024
    return size_str
In [77]:
df['Size in bytes'] = df['Size'].apply(parse_size)
In [78]:
df[['Size', 'Size in bytes']].head(10)
Out [78]:
Size Size in bytes
0 11k 11264.0
1 21M 22020096.0
2 41k 41984.0
3 292k 299008.0
4 3.0M 3145728.0
5 2.5M 2621440.0
6 8.2M 8598323.2
7 695k 711680.0
8 2.3M 2411724.8
9 1.4M 1468006.4
In [79]:
conversion_check = pd.DataFrame({
    'Example Input': ['11k', '21M', '1.4M'],
    'Expected Bytes': [11 * 1024, 21 * 1024 * 1024, 1.4 * 1024 * 1024]
})
conversion_check
Out [79]:
Example Input Expected Bytes
0 11k 11264.0
1 21M 22020096.0
2 1.4M 1468006.4
In [80]:
df.head(10)
Out [80]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver Size in bytes
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 11264.0
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up 22020096.0
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 41984.0
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up 299008.0
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up 3145728.0
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up 2621440.0
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up 8598323.2
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up 711680.0
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up 2411724.8
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up 1468006.4

Part C

Q: Create a new column called ‘Numeric_installs' (numeric) and convert the entries from ‘Installs’ column (remove ‘+’ and ‘,’; for example, “5,000,000+” becomes “5000000” as an Integer). [1 mark]

In [81]:
df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)
In [82]:
df.head(10)
Out [82]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver Size in bytes Numeric Installs
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 11264.0 1000000
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up 22020096.0 1000000
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 41984.0 5000000
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up 299008.0 1000000
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up 3145728.0 1000000
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up 2621440.0 1000
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up 8598323.2 100000
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up 711680.0 10000000
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up 2411724.8 100
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up 1468006.4 1000000

Part D

Q: Save the updated dataset as “googleplaystore_new_new.csv” [1 mark]

In [83]:
df.to_csv('data/googleplaystore_new_new.csv', index=False)

Part E

Q: Using (Category + Reviews + Content Rating + Size in Bytes + Installs_Num), find and discuss the best regression model to predict “rating” (use the standard training/test partition without cross-validation). [7 marks]

In [84]:
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.preprocessing import PolynomialFeatures
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
In [85]:
def split_overview(train_rows, test_rows, val_rows=None):
    rows = [
        {'Split': 'Train', 'Rows': train_rows},
        {'Split': 'Test', 'Rows': test_rows},
    ]
    if val_rows is not None:
        rows.insert(1, {'Split': 'Validation', 'Rows': val_rows})

    out = pd.DataFrame(rows)
    out['Pct_of_Total'] = (out['Rows'] / out['Rows'].sum() * 100).round(1)
    return out

def evaluate_model(model, X_train, X_test, y_train, y_test, name):
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)

    results = {
        'Model': name,
        'MAE': mean_absolute_error(y_test, y_pred),
        'MSE': mean_squared_error(y_test, y_pred),
        'RMSE': np.sqrt(mean_squared_error(y_test, y_pred)),
        'R2': r2_score(y_test, y_pred)
    }
    return results, y_pred


def regression_prediction_sheet(y_true, y_pred, model_name, head_n=20):
    sheet = pd.DataFrame({
        'True': y_true.reset_index(drop=True),
        'Predicted': pd.Series(y_pred).reset_index(drop=True)
    })
    sheet['Residual'] = sheet['True'] - sheet['Predicted']
    sheet['Abs_Error'] = sheet['Residual'].abs()
    sheet['Model'] = model_name
    return sheet.head(head_n)
In [86]:
df_new = pd.read_csv('data/googleplaystore_new_new.csv')
In [87]:
columns_to_keep = ['Category', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Rating']
df_min = df_new[columns_to_keep]
In [88]:
df_min.head(20)
Out [88]:
Category Reviews Content Rating Size in bytes Numeric Installs Rating
0 LIBRARIES_AND_DEMO 20145 Everyone 11264.0 1000000 4.1
1 BUSINESS 46353 Everyone 22020096.0 1000000 4.3
2 LIBRARIES_AND_DEMO 58055 Everyone 41984.0 5000000 3.9
3 LIBRARIES_AND_DEMO 7750 Everyone 299008.0 1000000 3.8
4 EDUCATION 1619 Everyone 3145728.0 1000000 4.4
5 LIBRARIES_AND_DEMO 26 Everyone 2621440.0 1000 5.0
6 BEAUTY 473 Mature 17+ 8598323.2 100000 4.5
7 COMMUNICATION 124346 Everyone 711680.0 10000000 4.2
8 EVENTS 16 Everyone 2411724.8 100 5.0
9 FINANCE 13868 Everyone 1468006.4 1000000 4.1
10 COMMUNICATION 32254 Everyone 5767168.0 1000000 4.4
11 COMMUNICATION 125232 Everyone 2831155.2 10000000 4.2
12 EDUCATION 430 Everyone 538624.0 10000 4.0
13 EDUCATION 275 Everyone 2411724.8 50000 4.0
14 BOOKS_AND_REFERENCE 1778 Mature 17+ 5138022.4 500000 3.9
15 BUSINESS 2287 Everyone 1572864.0 1000000 4.4
16 LIBRARIES_AND_DEMO 126862 Everyone 638976.0 10000000 3.5
17 COMMUNICATION 255 Everyone 1677721.6 10000 4.1
18 LIFESTYLE 360 Everyone 4823449.6 10000 4.1
19 EDUCATION 656 Everyone 569344.0 10000 4.3
In [89]:
df_encoded = pd.get_dummies(df_min, columns=['Category', 'Content Rating'])
In [90]:
X = df_encoded.drop('Rating', axis=1)
y = df_encoded['Rating']
In [91]:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=seed)
In [92]:
split_overview(len(X_train), len(X_test))
Out [92]:
Split Rows Pct_of_Total
0 Train 774 69.9
1 Test 333 30.1
In [93]:
results = []

Trying Linear Regression

In [94]:
lin_model = LinearRegression()
lin_res, y_pred_lin = evaluate_model(lin_model, X_train, X_test, y_train, y_test, "Linear Regression")
results.append(lin_res)
In [95]:
residual_linear = y_test - y_pred_lin

plt.figure(figsize=(8, 5))
plt.scatter(y_pred_lin, residual_linear, alpha=0.5)
plt.axhline(y=0, color='r', linestyle='--')
plt.title("Part E - Linear Regression: Residual Plot")
plt.xlabel("Predicted Rating")
plt.ylabel("Residual (y - y_hat)")
plt.tight_layout()
plt.show()
In [96]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred_lin, alpha=0.5)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.title("Part E - Linear Regression: Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Trying Polynomial Regression

In [97]:
best_degree = None
best_poly_rmse = float('inf')
best_poly_converter = None
best_poly_model = None
best_y_pred_poly = None

for degree in [2, 3]:
    poly_converter = PolynomialFeatures(degree=degree, include_bias=False)
    X_train_p = poly_converter.fit_transform(X_train)
    X_test_p = poly_converter.transform(X_test)

    poly_model = LinearRegression()
    poly_res_tmp, y_pred_poly_tmp = evaluate_model(poly_model, X_train_p, X_test_p, y_train, y_test, f"Polynomial Regression (degree={degree})")

    if poly_res_tmp['RMSE'] < best_poly_rmse:
        best_poly_rmse = poly_res_tmp['RMSE']
        best_degree = degree
        best_poly_converter = poly_converter
        best_poly_model = poly_model
        best_y_pred_poly = y_pred_poly_tmp

display(pd.DataFrame([{
    'Selection': 'Polynomial best degree by RMSE',
    'Best degree': best_degree,
    'RMSE': best_poly_rmse
}]))

# Set these names so the rest of the notebook works as before
poly_converter = best_poly_converter
poly_model = best_poly_model
X_train_p = poly_converter.fit_transform(X_train)
X_test_p = poly_converter.transform(X_test)
y_pred_poly = best_y_pred_poly
Selection Best degree RMSE
0 Polynomial best degree by RMSE 2 0.415806
In [98]:
poly_model = LinearRegression()
poly_res, y_pred_poly = evaluate_model(poly_model, X_train_p, X_test_p, y_train, y_test, "Polynomial Regression")
results.append(poly_res)
In [99]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred_poly, alpha=0.5)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.title("Part E - Polynomial Regression: Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Trying Ridge Regression

In [100]:
best_ridge_alpha = None
best_ridge_rmse = float('inf')
best_ridge_model = None
best_y_pred_ridge = None
best_ridge_res = None

for alpha in [0.1, 1.0, 10.0, 50.0]:
    ridge_model_tmp = Ridge(alpha=alpha, solver='lsqr', random_state=seed)
    ridge_res_tmp, y_pred_ridge_tmp = evaluate_model(ridge_model_tmp, X_train, X_test, y_train, y_test, f"Ridge Regression (alpha={alpha})")

    if ridge_res_tmp['RMSE'] < best_ridge_rmse:
        best_ridge_rmse = ridge_res_tmp['RMSE']
        best_ridge_alpha = alpha
        best_ridge_model = ridge_model_tmp
        best_y_pred_ridge = y_pred_ridge_tmp
        best_ridge_res = ridge_res_tmp

display(pd.DataFrame([{
    'Selection': 'Ridge best alpha by RMSE',
    'Best alpha': best_ridge_alpha,
    'RMSE': best_ridge_rmse
}]))

ridge_model = best_ridge_model
ridge_res = best_ridge_res
y_pred_ridge = best_y_pred_ridge
results.append(ridge_res)
Selection Best alpha RMSE
0 Ridge best alpha by RMSE 50.0 0.416581
In [101]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred_ridge, alpha=0.5)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.title("Part E - Ridge Regression: Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Trying Lasso Regression

In [102]:
best_lasso_alpha = None
best_lasso_rmse = float('inf')
best_lasso_model = None
best_y_pred_lasso = None
best_lasso_res = None

for alpha in [1e-4, 5e-4, 1e-3, 5e-3, 1e-2]:
    lasso_model_tmp = Lasso(alpha=alpha, max_iter=20000, random_state=seed)
    lasso_res_tmp, y_pred_lasso_tmp = evaluate_model(lasso_model_tmp, X_train, X_test, y_train, y_test, f"Lasso Regression (alpha={alpha})")

    if lasso_res_tmp['RMSE'] < best_lasso_rmse:
        best_lasso_rmse = lasso_res_tmp['RMSE']
        best_lasso_alpha = alpha
        best_lasso_model = lasso_model_tmp
        best_y_pred_lasso = y_pred_lasso_tmp
        best_lasso_res = lasso_res_tmp

display(pd.DataFrame([{
    'Selection': 'Lasso best alpha by RMSE',
    'Best alpha': best_lasso_alpha,
    'RMSE': best_lasso_rmse
}]))

lasso_model = best_lasso_model
lasso_res = best_lasso_res
y_pred_lasso = best_y_pred_lasso
results.append(lasso_res)
Selection Best alpha RMSE
0 Lasso best alpha by RMSE 0.0005 0.386607
In [103]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred_lasso, alpha=0.5)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.title("Part E - Lasso Regression: Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()
In [104]:
results_df = pd.DataFrame(results)
results_df
Out [104]:
Model MAE MSE RMSE R2
0 Linear Regression 0.280613 0.149546 0.386712 0.146514
1 Polynomial Regression 0.295032 0.172895 0.415806 0.013258
2 Ridge Regression (alpha=50.0) 0.302947 0.173539 0.416581 0.009580
3 Lasso Regression (alpha=0.0005) 0.280977 0.149465 0.386607 0.146976
In [105]:
pred_sheet_part_e_linear = regression_prediction_sheet(y_test, y_pred_lin, 'Linear Regression')
pred_sheet_part_e_poly = regression_prediction_sheet(y_test, y_pred_poly, 'Polynomial Regression')
pred_sheet_part_e_ridge = regression_prediction_sheet(y_test, y_pred_ridge, f'Ridge Regression (alpha={best_ridge_alpha})')
pred_sheet_part_e_lasso = regression_prediction_sheet(y_test, y_pred_lasso, f'Lasso Regression (alpha={best_lasso_alpha})')

best_model_part_e = results_df.sort_values('RMSE').iloc[0]['Model']
part_e_pred_sheet = {
    'Linear Regression': pred_sheet_part_e_linear,
    'Polynomial Regression': pred_sheet_part_e_poly,
}.get(best_model_part_e, pred_sheet_part_e_linear)

if 'Ridge Regression' in best_model_part_e:
    part_e_pred_sheet = pred_sheet_part_e_ridge
if 'Lasso Regression' in best_model_part_e:
    part_e_pred_sheet = pred_sheet_part_e_lasso

part_e_pred_sheet
Out [105]:
True Predicted Residual Abs_Error Model
0 4.1 4.306335 -0.206335 0.206335 Lasso Regression (alpha=0.0005)
1 4.1 4.287329 -0.187329 0.187329 Lasso Regression (alpha=0.0005)
2 4.4 4.297757 0.102243 0.102243 Lasso Regression (alpha=0.0005)
3 4.6 4.197993 0.402007 0.402007 Lasso Regression (alpha=0.0005)
4 4.5 4.371379 0.128621 0.128621 Lasso Regression (alpha=0.0005)
5 3.2 4.304839 -1.104839 1.104839 Lasso Regression (alpha=0.0005)
6 4.2 4.324286 -0.124286 0.124286 Lasso Regression (alpha=0.0005)
7 3.9 4.005045 -0.105045 0.105045 Lasso Regression (alpha=0.0005)
8 3.7 4.138157 -0.438157 0.438157 Lasso Regression (alpha=0.0005)
9 4.6 4.312731 0.287269 0.287269 Lasso Regression (alpha=0.0005)
10 4.6 4.468202 0.131798 0.131798 Lasso Regression (alpha=0.0005)
11 4.5 4.364965 0.135035 0.135035 Lasso Regression (alpha=0.0005)
12 4.6 4.443139 0.156861 0.156861 Lasso Regression (alpha=0.0005)
13 4.4 4.289412 0.110588 0.110588 Lasso Regression (alpha=0.0005)
14 4.0 4.302132 -0.302132 0.302132 Lasso Regression (alpha=0.0005)
15 4.3 4.373179 -0.073179 0.073179 Lasso Regression (alpha=0.0005)
16 4.4 4.172033 0.227967 0.227967 Lasso Regression (alpha=0.0005)
17 3.8 4.161517 -0.361517 0.361517 Lasso Regression (alpha=0.0005)
18 4.5 4.134825 0.365175 0.365175 Lasso Regression (alpha=0.0005)
19 4.2 4.279803 -0.079803 0.079803 Lasso Regression (alpha=0.0005)
In [106]:
part_e_prediction_compare = pd.concat([
    pred_sheet_part_e_linear.head(8),
    pred_sheet_part_e_poly.head(8),
    pred_sheet_part_e_ridge.head(8),
    pred_sheet_part_e_lasso.head(8),
], ignore_index=True)

part_e_prediction_compare
Out [106]:
True Predicted Residual Abs_Error Model
0 4.1 4.299298 -0.199298 0.199298 Linear Regression
1 4.1 4.277702 -0.177702 0.177702 Linear Regression
2 4.4 4.314717 0.085283 0.085283 Linear Regression
3 4.6 4.189314 0.410686 0.410686 Linear Regression
4 4.5 4.376513 0.123487 0.123487 Linear Regression
5 3.2 4.317265 -1.117265 1.117265 Linear Regression
6 4.2 4.345980 -0.145980 0.145980 Linear Regression
7 3.9 4.000260 -0.100260 0.100260 Linear Regression
8 4.1 4.279363 -0.179363 0.179363 Polynomial Regression
9 4.1 4.292264 -0.192264 0.192264 Polynomial Regression
10 4.4 4.275074 0.124926 0.124926 Polynomial Regression
11 4.6 4.230751 0.369249 0.369249 Polynomial Regression
12 4.5 4.287948 0.212052 0.212052 Polynomial Regression
13 3.2 4.233728 -1.033728 1.033728 Polynomial Regression
14 4.2 4.233425 -0.033425 0.033425 Polynomial Regression
15 3.9 4.081532 -0.181532 0.181532 Polynomial Regression
16 4.1 4.290152 -0.190152 0.190152 Ridge Regression (alpha=50.0)
17 4.1 4.261329 -0.161329 0.161329 Ridge Regression (alpha=50.0)
18 4.4 4.250952 0.149048 0.149048 Ridge Regression (alpha=50.0)
19 4.6 4.245164 0.354836 0.354836 Ridge Regression (alpha=50.0)
20 4.5 4.254136 0.245864 0.245864 Ridge Regression (alpha=50.0)
21 3.2 4.237515 -1.037515 1.037515 Ridge Regression (alpha=50.0)
22 4.2 4.237951 -0.037951 0.037951 Ridge Regression (alpha=50.0)
23 3.9 4.262971 -0.362971 0.362971 Ridge Regression (alpha=50.0)
24 4.1 4.306335 -0.206335 0.206335 Lasso Regression (alpha=0.0005)
25 4.1 4.287329 -0.187329 0.187329 Lasso Regression (alpha=0.0005)
26 4.4 4.297757 0.102243 0.102243 Lasso Regression (alpha=0.0005)
27 4.6 4.197993 0.402007 0.402007 Lasso Regression (alpha=0.0005)
28 4.5 4.371379 0.128621 0.128621 Lasso Regression (alpha=0.0005)
29 3.2 4.304839 -1.104839 1.104839 Lasso Regression (alpha=0.0005)
30 4.2 4.324286 -0.124286 0.124286 Lasso Regression (alpha=0.0005)
31 3.9 4.005045 -0.105045 0.105045 Lasso Regression (alpha=0.0005)
In [107]:
results_melted = results_df.melt(
    id_vars='Model',
    value_vars=['MAE', 'MSE', 'RMSE', 'R2'],
    var_name='Metric',
    value_name='Score'
)

plt.figure(figsize=(9, 5))
sns.barplot(data=results_melted, x='Model', y='Score', hue='Metric')
plt.title("Part E: Regression Model Comparison (MAE, RMSE, R2)")
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()

Model Performance Analysis

The plot above compares the regression models using MAE, RMSE, and R². I'm looking at these metrics to figure out which model is actually the best at predicting the rating:

  • RMSE (Root Mean Squared Error) and MAE (Mean Absolute Error): these basically tell us how far off our predictions are on average. Lower values mean the model is closer to the true ratings.
  • R²: this tells us how much of the variation in the ratings our model actually explains. We want this to be as close to 1 as possible. If it's near 0 (or negative), the model is just guessing or doing worse than a simple average.

Looking at the results, Linear Regression gives the best overall performance for this task since it has the lowest error (RMSE/MAE) and the highest R² out of all the models I tried. The polynomial model actually performs a bit worse here, and Ridge/Lasso are super close to the linear model but don't really improve it for this specific dataset.

So, based on the metrics, I'll go with Linear Regression as the best-performing regression model for predicting app ratings.

Part F

Q: Train and test all necessary model(s) to show and discuss the most predictive feature in the previous question [8 marks]

In [108]:
features_to_test = ['Reviews', 'Size in bytes', 'Numeric Installs']
feature_results = []
In [109]:
for feature in features_to_test:
    # Use a single input at a time to see how predictive it is for app rating
    single_input = df_encoded[[feature]]

    X_train_s, X_test_s, y_train_s, y_test_s = train_test_split(
        single_input, y, test_size=0.3, random_state=seed
    )

    simple_model = LinearRegression()
    simple_model.fit(X_train_s, y_train_s)
    y_pred_s = simple_model.predict(X_test_s)

    feature_results.append({
        'Feature': feature,
        'MAE': mean_absolute_error(y_test_s, y_pred_s),
        'MSE': mean_squared_error(y_test_s, y_pred_s),
        'RMSE': np.sqrt(mean_squared_error(y_test_s, y_pred_s)),
        'R2': r2_score(y_test_s, y_pred_s)
    })
In [110]:
feature_results_df = pd.DataFrame(feature_results).sort_values('RMSE')
feature_results_df
Out [110]:
Feature MAE MSE RMSE R2
1 Size in bytes 0.304495 0.173620 0.416677 0.009120
0 Reviews 0.304707 0.173865 0.416971 0.007722
2 Numeric Installs 0.307489 0.174643 0.417903 0.003281
In [111]:
feature_melted = feature_results_df.melt(
    id_vars='Feature',
    value_vars=['MAE', 'MSE', 'RMSE', 'R2'],
    var_name='Metric',
    value_name='Score'
)
In [112]:
plt.figure(figsize=(9, 5))
sns.barplot(data=feature_melted, x='Feature', y='Score', hue='Metric')
plt.title("Part F: Single-Feature Comparison (Linear Regression)")
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()
In [113]:
best_feature = feature_results_df.iloc[0]['Feature']

best_row = feature_results_df.iloc[0]
display(pd.DataFrame([{
    'Best Feature': best_row['Feature'],
    'MAE': best_row['MAE'],
    'MSE': best_row['MSE'],
    'RMSE': best_row['RMSE'],
    'R2': best_row['R2']
}]))
Best Feature MAE MSE RMSE R2
0 Size in bytes 0.304495 0.17362 0.416677 0.00912
In [114]:
X_best = df_encoded[[best_feature]]
X_train_b, X_test_b, y_train_b, y_test_b = train_test_split(
    X_best, y, test_size=0.3, random_state=seed
)
In [115]:
best_model = LinearRegression()
best_model.fit(X_train_b, y_train_b)
y_pred_b = best_model.predict(X_test_b)
In [116]:
pred_sheet_part_f = regression_prediction_sheet(y_test_b, y_pred_b, f'Linear Regression ({best_feature})')
pred_sheet_part_f
Out [116]:
True Predicted Residual Abs_Error Model
0 4.1 4.307502 -0.207502 0.207502 Linear Regression (Size in bytes)
1 4.1 4.267637 -0.167637 0.167637 Linear Regression (Size in bytes)
2 4.4 4.251691 0.148309 0.148309 Linear Regression (Size in bytes)
3 4.6 4.243240 0.356760 0.356760 Linear Regression (Size in bytes)
4 4.5 4.256475 0.243525 0.243525 Linear Regression (Size in bytes)
5 3.2 4.232078 -1.032078 1.032078 Linear Regression (Size in bytes)
6 4.2 4.232716 -0.032716 0.032716 Linear Regression (Size in bytes)
7 3.9 4.269232 -0.369232 0.369232 Linear Regression (Size in bytes)
8 3.7 4.251691 -0.551691 0.551691 Linear Regression (Size in bytes)
9 4.6 4.336205 0.263795 0.263795 Linear Regression (Size in bytes)
10 4.6 4.376070 0.223930 0.223930 Linear Regression (Size in bytes)
11 4.5 4.231759 0.268241 0.268241 Linear Regression (Size in bytes)
12 4.6 4.289962 0.310038 0.310038 Linear Regression (Size in bytes)
13 4.4 4.245313 0.154687 0.154687 Linear Regression (Size in bytes)
14 4.0 4.294746 -0.294746 0.294746 Linear Regression (Size in bytes)
15 4.3 4.261259 0.038741 0.038741 Linear Regression (Size in bytes)
16 4.4 4.275610 0.124390 0.124390 Linear Regression (Size in bytes)
17 3.8 4.241167 -0.441167 0.441167 Linear Regression (Size in bytes)
18 4.5 4.239732 0.260268 0.260268 Linear Regression (Size in bytes)
19 4.2 4.232237 -0.032237 0.032237 Linear Regression (Size in bytes)
In [117]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test_b, y_pred_b, alpha=0.5)
plt.plot([y_test_b.min(), y_test_b.max()], [y_test_b.min(), y_test_b.max()], 'r--')
plt.title(f"Part F - Best Feature ({best_feature}): Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Part F Analysis

To find out which single input is the best at predicting the app rating, I trained three separate simple linear regression models using just one input at a time:

  • Reviews (number of reviews)
  • Size in bytes (app size)
  • Numeric Installs (install count)

The bar chart above compares each model using MAE, RMSE, and R². The feature with the lowest RMSE is our most predictive single input. I printed out the best feature in the code cell above.

Looking at the "Actual vs Predicted" scatter plot for the best single feature, you can see the typical spread you get when trying to guess the rating from just one piece of info. This is why using multiple inputs together (like in the previous part) generally gives much better predictions than relying on just one.

Part G

Q: Using (Category + Reviews + Content Rating + Rating + Installs_Num), find and discuss the best regression model to predict “Size in Bytes” (use the standard training/test partition with cross-validation). [8 marks]

In [118]:
from sklearn.model_selection import cross_val_score
In [119]:
columns_to_keep_g = ['Category', 'Reviews', 'Content Rating', 'Rating', 'Numeric Installs', 'Size in bytes']
df_g = df_new[columns_to_keep_g]
In [120]:
df_g_encoded = pd.get_dummies(df_g, columns=['Category', 'Content Rating'])
In [121]:
X_g = df_g_encoded.drop('Size in bytes', axis=1)
y_g = df_g_encoded['Size in bytes']
In [122]:
X_train_g, X_test_g, y_train_g, y_test_g = train_test_split(X_g, y_g, test_size=0.3, random_state=seed)
In [123]:
split_overview(len(X_train_g), len(X_test_g))
Out [123]:
Split Rows Pct_of_Total
0 Train 774 69.9
1 Test 333 30.1
In [124]:
def evaluate_model_cv(model, X_train, X_test, y_train, y_test, name):
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)

    cv_scores = cross_val_score(model, X_train, y_train, cv=5, scoring='r2')

    results = {
        'Model': name,
        'MAE (test)': mean_absolute_error(y_test, y_pred),
        'MSE (test)': mean_squared_error(y_test, y_pred),
        'RMSE (test)': np.sqrt(mean_squared_error(y_test, y_pred)),
        'R2 (test)': r2_score(y_test, y_pred),
        'CV R2 (mean)': cv_scores.mean(),
        'CV R2 (std)': cv_scores.std()
    }
    return results, y_pred, cv_scores
In [125]:
results_g = []

Trying Linear Regression for Part G

In [126]:
lin_model_g = LinearRegression()
lin_res_g, y_pred_lin_g, cv_lin_g = evaluate_model_cv(lin_model_g, X_train_g, X_test_g, y_train_g, y_test_g, "Linear Regression")
results_g.append(lin_res_g)
In [127]:
residual_lin_g = y_test_g - y_pred_lin_g
In [128]:
plt.figure(figsize=(8, 5))
plt.scatter(y_pred_lin_g, residual_lin_g, alpha=0.5)
plt.axhline(y=0, color='r', linestyle='--')
plt.title("Part G - Linear Regression: Residual Plot")
plt.xlabel("Predicted Size in Bytes")
plt.ylabel("Residual (y - y_hat)")
plt.tight_layout()
plt.show()

Trying polynomial regression for Part G

In [129]:
poly_converter_g = PolynomialFeatures(degree=2, include_bias=False)
X_train_p_g = poly_converter_g.fit_transform(X_train_g)
X_test_p_g = poly_converter_g.transform(X_test_g)
In [130]:
poly_model_g = LinearRegression()
poly_res_g, y_pred_poly_g, cv_poly_g = evaluate_model_cv(poly_model_g, X_train_p_g, X_test_p_g, y_train_g, y_test_g, "Polynomial Regression")
results_g.append(poly_res_g)
In [131]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test_g, y_pred_poly_g, alpha=0.5)
plt.plot([y_test_g.min(), y_test_g.max()], [y_test_g.min(), y_test_g.max()], 'r--')
plt.title("Part G - Polynomial Regression: Actual vs Predicted")
plt.xlabel("True Size in Bytes")
plt.ylabel("Predicted Size in Bytes")
plt.tight_layout()
plt.show()

Trying Ridge Regression for Part G

In [132]:
best_ridge_alpha_g = None
best_ridge_cv_r2_g = float('-inf')
best_ridge_rmse_g = float('inf')
best_ridge_model_g = None
best_ridge_res_g = None
best_y_pred_ridge_g = None
best_cv_ridge_g = None

for alpha in [0.1, 1.0, 10.0, 50.0, 100.0]:
    ridge_model_tmp = Ridge(alpha=alpha, solver='lsqr', random_state=seed)
    ridge_res_tmp, y_pred_tmp, cv_tmp = evaluate_model_cv(
        ridge_model_tmp, X_train_g, X_test_g, y_train_g, y_test_g, f"Ridge Regression (alpha={alpha})"
    )

    if (ridge_res_tmp['CV R2 (mean)'] > best_ridge_cv_r2_g) or (
        ridge_res_tmp['CV R2 (mean)'] == best_ridge_cv_r2_g and ridge_res_tmp['RMSE (test)'] < best_ridge_rmse_g
    ):
        best_ridge_cv_r2_g = ridge_res_tmp['CV R2 (mean)']
        best_ridge_rmse_g = ridge_res_tmp['RMSE (test)']
        best_ridge_alpha_g = alpha
        best_ridge_model_g = ridge_model_tmp
        best_ridge_res_g = ridge_res_tmp
        best_y_pred_ridge_g = y_pred_tmp
        best_cv_ridge_g = cv_tmp

display(pd.DataFrame([{
    'Selection': 'Part G Ridge best alpha by CV R2 mean then RMSE',
    'Best alpha': best_ridge_alpha_g,
    'CV R2 mean': best_ridge_cv_r2_g,
    'RMSE (test)': best_ridge_rmse_g
}]))

ridge_model_g = best_ridge_model_g
ridge_res_g = best_ridge_res_g
y_pred_ridge_g = best_y_pred_ridge_g
cv_ridge_g = best_cv_ridge_g

results_g.append(ridge_res_g)
Selection Best alpha CV R2 mean RMSE (test)
0 Part G Ridge best alpha by CV R2 mean then RMSE 1.0 0.090846 2.162718e+07
In [133]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test_g, y_pred_ridge_g, alpha=0.5)
plt.plot([y_test_g.min(), y_test_g.max()], [y_test_g.min(), y_test_g.max()], 'r--')
plt.title("Part G - Ridge Regression: Actual vs Predicted")
plt.xlabel("True Size in Bytes")
plt.ylabel("Predicted Size in Bytes")
plt.tight_layout()
plt.show()

Regression Model Results

In [134]:
results_df_g = pd.DataFrame(results_g)
results_df_g
Out [134]:
Model MAE (test) MSE (test) RMSE (test) R2 (test) CV R2 (mean) CV R2 (std)
0 Linear Regression 1.458116e+07 3.634642e+14 1.906474e+07 0.292880 0.245829 0.067041
1 Polynomial Regression 1.708302e+07 7.093839e+14 2.663426e+07 -0.380108 -0.452607 0.900419
2 Ridge Regression (alpha=1.0) 1.629031e+07 4.677350e+14 2.162718e+07 0.090021 0.090846 0.056075
In [135]:
results_melted_g = results_df_g.melt(
    id_vars='Model',
    value_vars=['MAE (test)', 'MSE (test)', 'RMSE (test)', 'R2 (test)', 'CV R2 (mean)'],
    var_name='Metric',
    value_name='Score'
)
In [136]:
plt.figure(figsize=(10, 6))
sns.barplot(data=results_melted_g, x='Model', y='Score', hue='Metric')
plt.title("Part G: Regression Model Comparison for Size in Bytes (with Cross-Validation)")
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()

ranked_g = results_df_g.sort_values(by=['CV R2 (mean)', 'RMSE (test)'], ascending=[False, True]).reset_index(drop=True)
display(ranked_g)

best_model_g = ranked_g.iloc[0]['Model']
Model MAE (test) MSE (test) RMSE (test) R2 (test) CV R2 (mean) CV R2 (std)
0 Linear Regression 1.458116e+07 3.634642e+14 1.906474e+07 0.292880 0.245829 0.067041
1 Ridge Regression (alpha=1.0) 1.629031e+07 4.677350e+14 2.162718e+07 0.090021 0.090846 0.056075
2 Polynomial Regression 1.708302e+07 7.093839e+14 2.663426e+07 -0.380108 -0.452607 0.900419
In [137]:
display(pd.DataFrame([{'Part G Best Model': best_model_g}]))

pred_sheet_part_g_linear = regression_prediction_sheet(y_test_g, y_pred_lin_g, 'Linear Regression')
pred_sheet_part_g_poly = regression_prediction_sheet(y_test_g, y_pred_poly_g, 'Polynomial Regression')
pred_sheet_part_g_ridge = regression_prediction_sheet(y_test_g, y_pred_ridge_g, f'Ridge Regression (alpha={best_ridge_alpha_g})')

part_g_pred_sheet = {
    'Linear Regression': pred_sheet_part_g_linear,
    'Polynomial Regression': pred_sheet_part_g_poly,
}.get(best_model_g, pred_sheet_part_g_linear)

if 'Ridge Regression' in best_model_g:
    part_g_pred_sheet = pred_sheet_part_g_ridge

part_g_pred_sheet
Out [137]:
Part G Best Model
0 Linear Regression
True Predicted Residual Abs_Error Model
0 52428800.0 3.171894e+07 2.070986e+07 2.070986e+07 Linear Regression
1 26214400.0 1.364562e+07 1.256878e+07 1.256878e+07 Linear Regression
2 15728640.0 2.325830e+07 -7.529660e+06 7.529660e+06 Linear Regression
3 10171187.2 2.181280e+07 -1.164161e+07 1.164161e+07 Linear Regression
4 18874368.0 2.145150e+07 -2.577137e+06 2.577137e+06 Linear Regression
5 2831155.2 1.371277e+07 -1.088162e+07 1.088162e+07 Linear Regression
6 3250585.6 1.375252e+07 -1.050193e+07 1.050193e+07 Linear Regression
7 27262976.0 1.933636e+07 7.926614e+06 7.926614e+06 Linear Regression
8 15728640.0 1.726232e+07 -1.533682e+06 1.533682e+06 Linear Regression
9 71303168.0 3.203716e+07 3.926601e+07 3.926601e+07 Linear Regression
10 97517568.0 2.900348e+07 6.851409e+07 6.851409e+07 Linear Regression
11 2621440.0 2.145499e+07 -1.883355e+07 1.883355e+07 Linear Regression
12 40894464.0 2.853311e+07 1.236136e+07 1.236136e+07 Linear Regression
13 11534336.0 3.188308e+07 -2.034874e+07 2.034874e+07 Linear Regression
14 44040192.0 3.146167e+07 1.257852e+07 1.257852e+07 Linear Regression
15 22020096.0 2.135446e+07 6.656323e+05 6.656323e+05 Linear Regression
16 31457280.0 2.574583e+07 5.711451e+06 5.711451e+06 Linear Regression
17 8808038.4 2.481990e+07 -1.601186e+07 1.601186e+07 Linear Regression
18 7864320.0 1.801774e+07 -1.015342e+07 1.015342e+07 Linear Regression
19 2936012.8 1.400022e+07 -1.106421e+07 1.106421e+07 Linear Regression
In [138]:
part_g_prediction_compare = pd.concat([
    pred_sheet_part_g_linear.head(8),
    pred_sheet_part_g_poly.head(8),
    pred_sheet_part_g_ridge.head(8),
], ignore_index=True)

part_g_prediction_compare
Out [138]:
True Predicted Residual Abs_Error Model
0 52428800.0 3.171894e+07 2.070986e+07 2.070986e+07 Linear Regression
1 26214400.0 1.364562e+07 1.256878e+07 1.256878e+07 Linear Regression
2 15728640.0 2.325830e+07 -7.529660e+06 7.529660e+06 Linear Regression
3 10171187.2 2.181280e+07 -1.164161e+07 1.164161e+07 Linear Regression
4 18874368.0 2.145150e+07 -2.577137e+06 2.577137e+06 Linear Regression
5 2831155.2 1.371277e+07 -1.088162e+07 1.088162e+07 Linear Regression
6 3250585.6 1.375252e+07 -1.050193e+07 1.050193e+07 Linear Regression
7 27262976.0 1.933636e+07 7.926614e+06 7.926614e+06 Linear Regression
8 52428800.0 2.267715e+07 2.975165e+07 2.975165e+07 Polynomial Regression
9 26214400.0 1.925624e+07 6.958158e+06 6.958158e+06 Polynomial Regression
10 15728640.0 2.101692e+07 -5.288280e+06 5.288280e+06 Polynomial Regression
11 10171187.2 2.107864e+07 -1.090745e+07 1.090745e+07 Polynomial Regression
12 18874368.0 2.302922e+07 -4.154853e+06 4.154853e+06 Polynomial Regression
13 2831155.2 2.101016e+07 -1.817901e+07 1.817901e+07 Polynomial Regression
14 3250585.6 2.098037e+07 -1.772979e+07 1.772979e+07 Polynomial Regression
15 27262976.0 2.102464e+07 6.238332e+06 6.238332e+06 Polynomial Regression
16 52428800.0 2.375635e+07 2.867245e+07 2.867245e+07 Ridge Regression (alpha=1.0)
17 26214400.0 2.357213e+07 2.642274e+06 2.642274e+06 Ridge Regression (alpha=1.0)
18 15728640.0 2.349521e+07 -7.766567e+06 7.766567e+06 Ridge Regression (alpha=1.0)
19 10171187.2 2.349771e+07 -1.332652e+07 1.332652e+07 Ridge Regression (alpha=1.0)
20 18874368.0 2.349279e+07 -4.618422e+06 4.618422e+06 Ridge Regression (alpha=1.0)
21 2831155.2 2.349545e+07 -2.066429e+07 2.066429e+07 Ridge Regression (alpha=1.0)
22 3250585.6 2.349522e+07 -2.024464e+07 2.024464e+07 Ridge Regression (alpha=1.0)
23 27262976.0 2.349620e+07 3.766777e+06 3.766777e+06 Ridge Regression (alpha=1.0)

Part G Answer (with Cross-Validation)

For Part G, the goal is to predict Size in bytes using the other inputs:

  • Category
  • Reviews
  • Content Rating
  • Rating
  • Numeric Installs

I compared the allowed regression models using a standard train/test split, but this time I also added 5-fold cross-validation on the training set.

How I decided which model is best:

  • First, I looked for the model with the highest Cross-Validation mean R² since that tells me how well the model generalizes across different splits of the data.
  • Then, I used Test RMSE as a secondary check (lower is better).

What the results show:

  • Linear Regression is the strongest overall here: it has the best mean CV R² and the best test R², along with a lower test RMSE than Ridge. It's also way more stable than the polynomial model.
  • Polynomial Regression actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.
  • Ridge Regression is more stable than the polynomial model, but it still underperforms compared to standard Linear Regression on both metrics.

So, based on the cross-validation and test metrics, Linear Regression is the best model for predicting Size in bytes.

I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).