Files

458 KiB

This is the Q1 Notebook!

It's tracked via GitHub! hence the need for this line for the init commit

In [71]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
In [72]:
seed = 101
In [73]:
df = pd.read_csv('data/googleplaystore_new.csv')

Part A

Q: Check for any missing values. If any, delete all rows that contain null values Then, check for duplicates (similar contents for every feature). If any, delete the redundant rows. [1 mark]

In [74]:
df = df.dropna()
df = df.drop_duplicates()
In [75]:
df.head(10)
Out [75]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up

Part B

Q: Create a new column called 'Size in bytes' (numeric) and convert the entries from 'Size' column (M means megabyte and k means kilobytes). [1 mark]

In [76]:
def parse_size(size_str):
    if isinstance(size_str, str):
        if size_str.endswith('M'):
            return float(size_str[:-1]) * 1024 * 1024
        elif size_str.endswith('k'):
            return float(size_str[:-1]) * 1024
    return size_str
In [77]:
df['Size in bytes'] = df['Size'].apply(parse_size)
In [78]:
df[['Size', 'Size in bytes']].head(10)
Out [78]:
Size Size in bytes
0 11k 11264.0
1 21M 22020096.0
2 41k 41984.0
3 292k 299008.0
4 3.0M 3145728.0
5 2.5M 2621440.0
6 8.2M 8598323.2
7 695k 711680.0
8 2.3M 2411724.8
9 1.4M 1468006.4
In [79]:
conversion_check = pd.DataFrame({
    'Example Input': ['11k', '21M', '1.4M'],
    'Expected Bytes': [11 * 1024, 21 * 1024 * 1024, 1.4 * 1024 * 1024]
})
conversion_check
Out [79]:
Example Input Expected Bytes
0 11k 11264.0
1 21M 22020096.0
2 1.4M 1468006.4
In [80]:
df.head(10)
Out [80]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver Size in bytes
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 11264.0
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up 22020096.0
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 41984.0
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up 299008.0
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up 3145728.0
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up 2621440.0
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up 8598323.2
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up 711680.0
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up 2411724.8
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up 1468006.4

Part C

Q: Create a new column called ‘Numeric_installs' (numeric) and convert the entries from ‘Installs’ column (remove ‘+’ and ‘,’; for example, “5,000,000+” becomes “5000000” as an Integer). [1 mark]

In [81]:
df['Numeric Installs'] = df['Installs'].str.replace('+', '').str.replace(',', '').astype(int)
In [82]:
df.head(10)
Out [82]:
App Category Rating Reviews Size Installs Type Price Content Rating Genres Android Ver Size in bytes Numeric Installs
0 Market Update Helper LIBRARIES_AND_DEMO 4.1 20145 11k 1,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 11264.0 1000000
1 SuperLivePro BUSINESS 4.3 46353 21M 1,000,000+ Free 0 Everyone Business 1.5 and up 22020096.0 1000000
2 Wifi Connect Library LIBRARIES_AND_DEMO 3.9 58055 41k 5,000,000+ Free 0 Everyone Libraries & Demo 1.5 and up 41984.0 5000000
3 Apk Installer LIBRARIES_AND_DEMO 3.8 7750 292k 1,000,000+ Free 0 Everyone Libraries & Demo 1.6 and up 299008.0 1000000
4 English speaking texts EDUCATION 4.4 1619 3.0M 1,000,000+ Free 0 Everyone Education 1.6 and up 3145728.0 1000000
5 Eternal life LIBRARIES_AND_DEMO 5.0 26 2.5M 1,000+ Free 0 Everyone Libraries & Demo 1.6 and up 2621440.0 1000
6 Dresses Ideas & Fashions +3000 BEAUTY 4.5 473 8.2M 100,000+ Free 0 Mature 17+ Beauty 1.6 and up 8598323.2 100000
7 GO Notifier COMMUNICATION 4.2 124346 695k 10,000,000+ Free 0 Everyone Communication 2.0 and up 711680.0 10000000
8 Prosperity EVENTS 5.0 16 2.3M 100+ Free 0 Everyone Events 2.0 and up 2411724.8 100
9 NSE Mobile Trading FINANCE 4.1 13868 1.4M 1,000,000+ Free 0 Everyone Finance 2.1 and up 1468006.4 1000000

Part D

Q: Save the updated dataset as “googleplaystore_new_new.csv” [1 mark]

In [83]:
df.to_csv('data/googleplaystore_new_new.csv', index=False)

Part E

Q: Using (Category + Reviews + Content Rating + Size in Bytes + Installs_Num), find and discuss the best regression model to predict “rating” (use the standard training/test partition without cross-validation). [7 marks]

In [84]:
from sklearn.model_selection import train_test_split
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.preprocessing import PolynomialFeatures
from sklearn.metrics import mean_absolute_error, mean_squared_error, r2_score
In [85]:
def split_overview(train_rows, test_rows, val_rows=None):
    rows = [
        {'Split': 'Train', 'Rows': train_rows},
        {'Split': 'Test', 'Rows': test_rows},
    ]
    if val_rows is not None:
        rows.insert(1, {'Split': 'Validation', 'Rows': val_rows})

    out = pd.DataFrame(rows)
    out['Pct_of_Total'] = (out['Rows'] / out['Rows'].sum() * 100).round(1)
    return out

def evaluate_model(model, X_train, X_test, y_train, y_test, name):
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)

    results = {
        'Model': name,
        'MAE': mean_absolute_error(y_test, y_pred),
        'MSE': mean_squared_error(y_test, y_pred),
        'RMSE': np.sqrt(mean_squared_error(y_test, y_pred)),
        'R2': r2_score(y_test, y_pred)
    }
    return results, y_pred


def regression_prediction_sheet(y_true, y_pred, model_name, head_n=20):
    sheet = pd.DataFrame({
        'True': y_true.reset_index(drop=True),
        'Predicted': pd.Series(y_pred).reset_index(drop=True)
    })
    sheet['Residual'] = sheet['True'] - sheet['Predicted']
    sheet['Abs_Error'] = sheet['Residual'].abs()
    sheet['Model'] = model_name
    return sheet.head(head_n)
In [86]:
df_new = pd.read_csv('data/googleplaystore_new_new.csv')
In [87]:
columns_to_keep = ['Category', 'Reviews', 'Content Rating', 'Size in bytes', 'Numeric Installs', 'Rating']
df_min = df_new[columns_to_keep]
In [88]:
df_min.head(20)
Out [88]:
Category Reviews Content Rating Size in bytes Numeric Installs Rating
0 LIBRARIES_AND_DEMO 20145 Everyone 11264.0 1000000 4.1
1 BUSINESS 46353 Everyone 22020096.0 1000000 4.3
2 LIBRARIES_AND_DEMO 58055 Everyone 41984.0 5000000 3.9
3 LIBRARIES_AND_DEMO 7750 Everyone 299008.0 1000000 3.8
4 EDUCATION 1619 Everyone 3145728.0 1000000 4.4
5 LIBRARIES_AND_DEMO 26 Everyone 2621440.0 1000 5.0
6 BEAUTY 473 Mature 17+ 8598323.2 100000 4.5
7 COMMUNICATION 124346 Everyone 711680.0 10000000 4.2
8 EVENTS 16 Everyone 2411724.8 100 5.0
9 FINANCE 13868 Everyone 1468006.4 1000000 4.1
10 COMMUNICATION 32254 Everyone 5767168.0 1000000 4.4
11 COMMUNICATION 125232 Everyone 2831155.2 10000000 4.2
12 EDUCATION 430 Everyone 538624.0 10000 4.0
13 EDUCATION 275 Everyone 2411724.8 50000 4.0
14 BOOKS_AND_REFERENCE 1778 Mature 17+ 5138022.4 500000 3.9
15 BUSINESS 2287 Everyone 1572864.0 1000000 4.4
16 LIBRARIES_AND_DEMO 126862 Everyone 638976.0 10000000 3.5
17 COMMUNICATION 255 Everyone 1677721.6 10000 4.1
18 LIFESTYLE 360 Everyone 4823449.6 10000 4.1
19 EDUCATION 656 Everyone 569344.0 10000 4.3
In [89]:
df_encoded = pd.get_dummies(df_min, columns=['Category', 'Content Rating'])
In [90]:
X = df_encoded.drop('Rating', axis=1)
y = df_encoded['Rating']
In [91]:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=seed)
In [92]:
split_overview(len(X_train), len(X_test))
Out [92]:
Split Rows Pct_of_Total
0 Train 774 69.9
1 Test 333 30.1
In [93]:
results = []

Trying Linear Regression

In [94]:
lin_model = LinearRegression()
lin_res, y_pred_lin = evaluate_model(lin_model, X_train, X_test, y_train, y_test, "Linear Regression")
results.append(lin_res)
In [95]:
residual_linear = y_test - y_pred_lin

plt.figure(figsize=(8, 5))
plt.scatter(y_pred_lin, residual_linear, alpha=0.5)
plt.axhline(y=0, color='r', linestyle='--')
plt.title("Part E - Linear Regression: Residual Plot")
plt.xlabel("Predicted Rating")
plt.ylabel("Residual (y - y_hat)")
plt.tight_layout()
plt.show()
In [96]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test, y_pred_lin, alpha=0.5)
plt.plot([y_test.min(), y_test.max()], [y_test.min(), y_test.max()], 'r--')
plt.title("Part E - Linear Regression: Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Trying Polynomial Regression

In [97]:
best_degree = None
best_poly_rmse = float('inf')
best_poly_converter = None
best_poly_model = None
best_y_pred_poly = None

for degree in [2, 3]:
    poly_converter = PolynomialFeatures(degree=degree, include_bias=False)
    X_train_p = poly_converter.fit_transform(X_train)
    X_test_p = poly_converter.transform(X_test)

    poly_model = LinearRegression()
    poly_res_tmp, y_pred_poly_tmp = evaluate_model(poly_model, X_train_p, X_test_p, y_train, y_test, f"Polynomial Regression (degree={degree})")

    if poly_res_tmp['RMSE'] < best_poly_rmse:
        best_poly_rmse = poly_res_tmp['RMSE']
        best_degree = degree
        best_poly_converter = poly_converter
        best_poly_model = poly_model
        best_y_pred_poly = y_pred_poly_tmp

display(pd.DataFrame([{
    'Selection': 'Polynomial best degree by RMSE',
    'Best degree': best_degree,
    'RMSE': best_poly_rmse
}]))

# Set these names so the rest of the notebook works as before
poly_converter = best_poly_converter
poly_model = best_poly_model
X_train_p = poly_converter.fit_transform(X_train)
X_test_p = poly_converter.transform(X_test)
y_pred_poly = best_y_pred_poly
Selection Best degree RMSE
0 Polynomial best degree by RMSE 2 0.415806

Model Performance Analysis

The plot above compares the regression models using MAE, RMSE, and R². I'm looking at these metrics to figure out which model is actually the best at predicting the rating:

  • RMSE (Root Mean Squared Error) and MAE (Mean Absolute Error): these basically tell us how far off our predictions are on average. Lower values mean the model is closer to the true ratings.
  • R²: this tells us how much of the variation in the ratings our model actually explains. We want this to be as close to 1 as possible. If it's near 0 (or negative), the model is just guessing or doing worse than a simple average.

Looking at the results, Linear Regression gives the best overall performance for this task since it has the lowest error (RMSE/MAE) and the highest R² out of all the models I tried. The polynomial model actually performs a bit worse here, and Ridge/Lasso are super close to the linear model but don't really improve it for this specific dataset.

So, based on the metrics, I'll go with Linear Regression as the best-performing regression model for predicting app ratings.

Part F

Q: Train and test all necessary model(s) to show and discuss the most predictive feature in the previous question [8 marks]

In [108]:
features_to_test = ['Reviews', 'Size in bytes', 'Numeric Installs']
feature_results = []
In [109]:
for feature in features_to_test:
    # Use a single input at a time to see how predictive it is for app rating
    single_input = df_encoded[[feature]]

    X_train_s, X_test_s, y_train_s, y_test_s = train_test_split(
        single_input, y, test_size=0.3, random_state=seed
    )

    simple_model = LinearRegression()
    simple_model.fit(X_train_s, y_train_s)
    y_pred_s = simple_model.predict(X_test_s)

    feature_results.append({
        'Feature': feature,
        'MAE': mean_absolute_error(y_test_s, y_pred_s),
        'MSE': mean_squared_error(y_test_s, y_pred_s),
        'RMSE': np.sqrt(mean_squared_error(y_test_s, y_pred_s)),
        'R2': r2_score(y_test_s, y_pred_s)
    })
In [110]:
feature_results_df = pd.DataFrame(feature_results).sort_values('RMSE')
feature_results_df
Out [110]:
Feature MAE MSE RMSE R2
1 Size in bytes 0.304495 0.173620 0.416677 0.009120
0 Reviews 0.304707 0.173865 0.416971 0.007722
2 Numeric Installs 0.307489 0.174643 0.417903 0.003281
In [111]:
feature_melted = feature_results_df.melt(
    id_vars='Feature',
    value_vars=['MAE', 'MSE', 'RMSE', 'R2'],
    var_name='Metric',
    value_name='Score'
)
In [112]:
plt.figure(figsize=(9, 5))
sns.barplot(data=feature_melted, x='Feature', y='Score', hue='Metric')
plt.title("Part F: Single-Feature Comparison (Linear Regression)")
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()
In [113]:
best_feature = feature_results_df.iloc[0]['Feature']

best_row = feature_results_df.iloc[0]
display(pd.DataFrame([{
    'Best Feature': best_row['Feature'],
    'MAE': best_row['MAE'],
    'MSE': best_row['MSE'],
    'RMSE': best_row['RMSE'],
    'R2': best_row['R2']
}]))
Best Feature MAE MSE RMSE R2
0 Size in bytes 0.304495 0.17362 0.416677 0.00912
In [114]:
X_best = df_encoded[[best_feature]]
X_train_b, X_test_b, y_train_b, y_test_b = train_test_split(
    X_best, y, test_size=0.3, random_state=seed
)
In [115]:
best_model = LinearRegression()
best_model.fit(X_train_b, y_train_b)
y_pred_b = best_model.predict(X_test_b)
In [116]:
pred_sheet_part_f = regression_prediction_sheet(y_test_b, y_pred_b, f'Linear Regression ({best_feature})')
pred_sheet_part_f
Out [116]:
True Predicted Residual Abs_Error Model
0 4.1 4.307502 -0.207502 0.207502 Linear Regression (Size in bytes)
1 4.1 4.267637 -0.167637 0.167637 Linear Regression (Size in bytes)
2 4.4 4.251691 0.148309 0.148309 Linear Regression (Size in bytes)
3 4.6 4.243240 0.356760 0.356760 Linear Regression (Size in bytes)
4 4.5 4.256475 0.243525 0.243525 Linear Regression (Size in bytes)
5 3.2 4.232078 -1.032078 1.032078 Linear Regression (Size in bytes)
6 4.2 4.232716 -0.032716 0.032716 Linear Regression (Size in bytes)
7 3.9 4.269232 -0.369232 0.369232 Linear Regression (Size in bytes)
8 3.7 4.251691 -0.551691 0.551691 Linear Regression (Size in bytes)
9 4.6 4.336205 0.263795 0.263795 Linear Regression (Size in bytes)
10 4.6 4.376070 0.223930 0.223930 Linear Regression (Size in bytes)
11 4.5 4.231759 0.268241 0.268241 Linear Regression (Size in bytes)
12 4.6 4.289962 0.310038 0.310038 Linear Regression (Size in bytes)
13 4.4 4.245313 0.154687 0.154687 Linear Regression (Size in bytes)
14 4.0 4.294746 -0.294746 0.294746 Linear Regression (Size in bytes)
15 4.3 4.261259 0.038741 0.038741 Linear Regression (Size in bytes)
16 4.4 4.275610 0.124390 0.124390 Linear Regression (Size in bytes)
17 3.8 4.241167 -0.441167 0.441167 Linear Regression (Size in bytes)
18 4.5 4.239732 0.260268 0.260268 Linear Regression (Size in bytes)
19 4.2 4.232237 -0.032237 0.032237 Linear Regression (Size in bytes)
In [117]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test_b, y_pred_b, alpha=0.5)
plt.plot([y_test_b.min(), y_test_b.max()], [y_test_b.min(), y_test_b.max()], 'r--')
plt.title(f"Part F - Best Feature ({best_feature}): Actual vs Predicted")
plt.xlabel("True Rating")
plt.ylabel("Predicted Rating")
plt.tight_layout()
plt.show()

Part F Analysis

To find out which single input is the best at predicting the app rating, I trained two separate simple linear regression models using just one input at a time:

  • Reviews (number of reviews)
  • Size in bytes (app size)
  • Numeric Installs (install count)

The bar chart above compares each model using MAE, RMSE, and R². The feature with the lowest RMSE is our most predictive single input. I printed out the best feature in the code cell above.

Looking at the "Actual vs Predicted" scatter plot for the best single feature, you can see the typical spread you get when trying to guess the rating from just one piece of info.

This is why using multiple inputs together (like in the previous part) generally gives much better predictions than relying on just one.

Part G

Q: Using (Category + Reviews + Content Rating + Rating + Installs_Num), find and discuss the best regression model to predict “Size in Bytes” (use the standard training/test partition with cross-validation). [8 marks]

In [118]:
from sklearn.model_selection import cross_val_score
In [119]:
columns_to_keep_g = ['Category', 'Reviews', 'Content Rating', 'Rating', 'Numeric Installs', 'Size in bytes']
df_g = df_new[columns_to_keep_g]
In [120]:
df_g_encoded = pd.get_dummies(df_g, columns=['Category', 'Content Rating'])
In [121]:
X_g = df_g_encoded.drop('Size in bytes', axis=1)
y_g = df_g_encoded['Size in bytes']
In [122]:
X_train_g, X_test_g, y_train_g, y_test_g = train_test_split(X_g, y_g, test_size=0.3, random_state=seed)
In [123]:
split_overview(len(X_train_g), len(X_test_g))
Out [123]:
Split Rows Pct_of_Total
0 Train 774 69.9
1 Test 333 30.1
In [124]:
def evaluate_model_cv(model, X_train, X_test, y_train, y_test, name):
    model.fit(X_train, y_train)
    y_pred = model.predict(X_test)

    cv_scores = cross_val_score(model, X_train, y_train, cv=5, scoring='r2')

    results = {
        'Model': name,
        'MAE (test)': mean_absolute_error(y_test, y_pred),
        'MSE (test)': mean_squared_error(y_test, y_pred),
        'RMSE (test)': np.sqrt(mean_squared_error(y_test, y_pred)),
        'R2 (test)': r2_score(y_test, y_pred),
        'CV R2 (mean)': cv_scores.mean(),
        'CV R2 (std)': cv_scores.std()
    }
    return results, y_pred, cv_scores
In [125]:
results_g = []

Trying Linear Regression for Part G

In [126]:
lin_model_g = LinearRegression()
lin_res_g, y_pred_lin_g, cv_lin_g = evaluate_model_cv(lin_model_g, X_train_g, X_test_g, y_train_g, y_test_g, "Linear Regression")
results_g.append(lin_res_g)
In [127]:
residual_lin_g = y_test_g - y_pred_lin_g
In [128]:
plt.figure(figsize=(8, 5))
plt.scatter(y_pred_lin_g, residual_lin_g, alpha=0.5)
plt.axhline(y=0, color='r', linestyle='--')
plt.title("Part G - Linear Regression: Residual Plot")
plt.xlabel("Predicted Size in Bytes")
plt.ylabel("Residual (y - y_hat)")
plt.tight_layout()
plt.show()

Trying polynomial regression for Part G

In [129]:
poly_converter_g = PolynomialFeatures(degree=2, include_bias=False)
X_train_p_g = poly_converter_g.fit_transform(X_train_g)
X_test_p_g = poly_converter_g.transform(X_test_g)
In [130]:
poly_model_g = LinearRegression()
poly_res_g, y_pred_poly_g, cv_poly_g = evaluate_model_cv(poly_model_g, X_train_p_g, X_test_p_g, y_train_g, y_test_g, "Polynomial Regression")
results_g.append(poly_res_g)
In [131]:
plt.figure(figsize=(6, 6))
plt.scatter(y_test_g, y_pred_poly_g, alpha=0.5)
plt.plot([y_test_g.min(), y_test_g.max()], [y_test_g.min(), y_test_g.max()], 'r--')
plt.title("Part G - Polynomial Regression: Actual vs Predicted")
plt.xlabel("True Size in Bytes")
plt.ylabel("Predicted Size in Bytes")
plt.tight_layout()
plt.show()

Regression Model Results

In [134]:
results_df_g = pd.DataFrame(results_g)
results_df_g
Out [134]:
Model MAE (test) MSE (test) RMSE (test) R2 (test) CV R2 (mean) CV R2 (std)
0 Linear Regression 1.458116e+07 3.634642e+14 1.906474e+07 0.292880 0.245829 0.067041
1 Polynomial Regression 1.708302e+07 7.093839e+14 2.663426e+07 -0.380108 -0.452607 0.900419
2 Ridge Regression (alpha=1.0) 1.629031e+07 4.677350e+14 2.162718e+07 0.090021 0.090846 0.056075
In [135]:
results_melted_g = results_df_g.melt(
    id_vars='Model',
    value_vars=['MAE (test)', 'MSE (test)', 'RMSE (test)', 'R2 (test)', 'CV R2 (mean)'],
    var_name='Metric',
    value_name='Score'
)
In [136]:
plt.figure(figsize=(10, 6))
sns.barplot(data=results_melted_g, x='Model', y='Score', hue='Metric')
plt.title("Part G: Regression Model Comparison for Size in Bytes (with Cross-Validation)")
plt.xticks(rotation=15)
plt.tight_layout()
plt.show()

ranked_g = results_df_g.sort_values(by=['CV R2 (mean)', 'RMSE (test)'], ascending=[False, True]).reset_index(drop=True)
display(ranked_g)

best_model_g = ranked_g.iloc[0]['Model']
Model MAE (test) MSE (test) RMSE (test) R2 (test) CV R2 (mean) CV R2 (std)
0 Linear Regression 1.458116e+07 3.634642e+14 1.906474e+07 0.292880 0.245829 0.067041
1 Ridge Regression (alpha=1.0) 1.629031e+07 4.677350e+14 2.162718e+07 0.090021 0.090846 0.056075
2 Polynomial Regression 1.708302e+07 7.093839e+14 2.663426e+07 -0.380108 -0.452607 0.900419
In [137]:
display(pd.DataFrame([{'Part G Best Model': best_model_g}]))

pred_sheet_part_g_linear = regression_prediction_sheet(y_test_g, y_pred_lin_g, 'Linear Regression')
pred_sheet_part_g_poly = regression_prediction_sheet(y_test_g, y_pred_poly_g, 'Polynomial Regression')
pred_sheet_part_g_ridge = regression_prediction_sheet(y_test_g, y_pred_ridge_g, f'Ridge Regression (alpha={best_ridge_alpha_g})')

part_g_pred_sheet = {
    'Linear Regression': pred_sheet_part_g_linear,
    'Polynomial Regression': pred_sheet_part_g_poly,
}.get(best_model_g, pred_sheet_part_g_linear)

if 'Ridge Regression' in best_model_g:
    part_g_pred_sheet = pred_sheet_part_g_ridge

part_g_pred_sheet
Out [137]:
Part G Best Model
0 Linear Regression
True Predicted Residual Abs_Error Model
0 52428800.0 3.171894e+07 2.070986e+07 2.070986e+07 Linear Regression
1 26214400.0 1.364562e+07 1.256878e+07 1.256878e+07 Linear Regression
2 15728640.0 2.325830e+07 -7.529660e+06 7.529660e+06 Linear Regression
3 10171187.2 2.181280e+07 -1.164161e+07 1.164161e+07 Linear Regression
4 18874368.0 2.145150e+07 -2.577137e+06 2.577137e+06 Linear Regression
5 2831155.2 1.371277e+07 -1.088162e+07 1.088162e+07 Linear Regression
6 3250585.6 1.375252e+07 -1.050193e+07 1.050193e+07 Linear Regression
7 27262976.0 1.933636e+07 7.926614e+06 7.926614e+06 Linear Regression
8 15728640.0 1.726232e+07 -1.533682e+06 1.533682e+06 Linear Regression
9 71303168.0 3.203716e+07 3.926601e+07 3.926601e+07 Linear Regression
10 97517568.0 2.900348e+07 6.851409e+07 6.851409e+07 Linear Regression
11 2621440.0 2.145499e+07 -1.883355e+07 1.883355e+07 Linear Regression
12 40894464.0 2.853311e+07 1.236136e+07 1.236136e+07 Linear Regression
13 11534336.0 3.188308e+07 -2.034874e+07 2.034874e+07 Linear Regression
14 44040192.0 3.146167e+07 1.257852e+07 1.257852e+07 Linear Regression
15 22020096.0 2.135446e+07 6.656323e+05 6.656323e+05 Linear Regression
16 31457280.0 2.574583e+07 5.711451e+06 5.711451e+06 Linear Regression
17 8808038.4 2.481990e+07 -1.601186e+07 1.601186e+07 Linear Regression
18 7864320.0 1.801774e+07 -1.015342e+07 1.015342e+07 Linear Regression
19 2936012.8 1.400022e+07 -1.106421e+07 1.106421e+07 Linear Regression
In [138]:
part_g_prediction_compare = pd.concat([
    pred_sheet_part_g_linear.head(8),
    pred_sheet_part_g_poly.head(8),
    pred_sheet_part_g_ridge.head(8),
], ignore_index=True)

part_g_prediction_compare
Out [138]:
True Predicted Residual Abs_Error Model
0 52428800.0 3.171894e+07 2.070986e+07 2.070986e+07 Linear Regression
1 26214400.0 1.364562e+07 1.256878e+07 1.256878e+07 Linear Regression
2 15728640.0 2.325830e+07 -7.529660e+06 7.529660e+06 Linear Regression
3 10171187.2 2.181280e+07 -1.164161e+07 1.164161e+07 Linear Regression
4 18874368.0 2.145150e+07 -2.577137e+06 2.577137e+06 Linear Regression
5 2831155.2 1.371277e+07 -1.088162e+07 1.088162e+07 Linear Regression
6 3250585.6 1.375252e+07 -1.050193e+07 1.050193e+07 Linear Regression
7 27262976.0 1.933636e+07 7.926614e+06 7.926614e+06 Linear Regression
8 52428800.0 2.267715e+07 2.975165e+07 2.975165e+07 Polynomial Regression
9 26214400.0 1.925624e+07 6.958158e+06 6.958158e+06 Polynomial Regression
10 15728640.0 2.101692e+07 -5.288280e+06 5.288280e+06 Polynomial Regression
11 10171187.2 2.107864e+07 -1.090745e+07 1.090745e+07 Polynomial Regression
12 18874368.0 2.302922e+07 -4.154853e+06 4.154853e+06 Polynomial Regression
13 2831155.2 2.101016e+07 -1.817901e+07 1.817901e+07 Polynomial Regression
14 3250585.6 2.098037e+07 -1.772979e+07 1.772979e+07 Polynomial Regression
15 27262976.0 2.102464e+07 6.238332e+06 6.238332e+06 Polynomial Regression
16 52428800.0 2.375635e+07 2.867245e+07 2.867245e+07 Ridge Regression (alpha=1.0)
17 26214400.0 2.357213e+07 2.642274e+06 2.642274e+06 Ridge Regression (alpha=1.0)
18 15728640.0 2.349521e+07 -7.766567e+06 7.766567e+06 Ridge Regression (alpha=1.0)
19 10171187.2 2.349771e+07 -1.332652e+07 1.332652e+07 Ridge Regression (alpha=1.0)
20 18874368.0 2.349279e+07 -4.618422e+06 4.618422e+06 Ridge Regression (alpha=1.0)
21 2831155.2 2.349545e+07 -2.066429e+07 2.066429e+07 Ridge Regression (alpha=1.0)
22 3250585.6 2.349522e+07 -2.024464e+07 2.024464e+07 Ridge Regression (alpha=1.0)
23 27262976.0 2.349620e+07 3.766777e+06 3.766777e+06 Ridge Regression (alpha=1.0)

Part G Answer (with Cross-Validation)

For Part G, the goal was to predict Size in bytes using the other inputs:

  • Category
  • Reviews
  • Content Rating
  • Rating
  • Numeric Installs

I compared the regression models using a standard train/test split, but this time I also added 5-fold cross-validation on the training set.

How I decided which model is best: First, I looked for the model with the highest Cross-Validation mean R2 since that tells me how well the model generalizes across different splits of the data. Then, I used the test RMSE as a secondary check (lower is better).

What the results show:

  • Linear Regression is the strongest overall here: it has the best mean CV R² and the best test R², along with a low test RMSE. It's also way more stable than the polynomial model.
  • Polynomial Regression actually performs really poorly for this dataset (it got a negative test R² and a super unstable CV R²), meaning it's overcomplicating things and not generalizing well.

So, based on the cross-validation and test metrics, Linear Regression is the best model for predicting Size in bytes.

I also use the residual and actual-vs-predicted plots to visually confirm if the errors are randomly spread out (which is what we want) or if they show patterns (which means the model is missing something).

In [ ]: