🏡 AI-Driven Agents for Fair and Efficient Property Valuation¶

  • Student Name: Fahad Abubaker Bahashwan
  • Advisor: Dr. Ashraf Elsayed
  • Institution: MidOcean University
  • Date: 23, July 2025

What we'll do:

  1. Load and explore data
  2. Create geographic features
  3. Train different ML models
  4. Compare which model works best

Step 1: Import Libraries¶

Import all the tools we need:

In [2]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from geopy.distance import geodesic
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.compose import ColumnTransformer
from sklearn.linear_model import Ridge, Lasso, LinearRegression
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.tree import DecisionTreeRegressor
from sklearn.neighbors import KNeighborsRegressor
from sklearn.metrics import mean_squared_error, r2_score
import xgboost as xgb

sns.set(style="whitegrid")

Step 2: Load Data¶

In [3]:
df = pd.read_excel("WorkedData6.xlsx")
df['transactionDate'] = pd.to_datetime(df['transactionDate'])
df['transaction_year'] = df['transactionDate'].dt.year
df['transaction_month'] = df['transactionDate'].dt.month
df['days_since_tx'] = (df['transactionDate'].max() - df['transactionDate']).dt.days

3️⃣ Calculate Jeddah Center Point¶

In [4]:
jeddah_center = (21.543333, 39.172778)
df['dist_to_center_km'] = df.apply(
    lambda row: geodesic((row['latitude'], row['longitude']), jeddah_center).kilometers, axis=1)

4️⃣ data exploration¶

In [5]:
df.info()
<class 'pandas.core.frame.DataFrame'>
RangeIndex: 2135 entries, 0 to 2134
Data columns (total 10 columns):
 #   Column             Non-Null Count  Dtype         
---  ------             --------------  -----         
 0   district_id        2135 non-null   int64         
 1   latitude           2135 non-null   float64       
 2   longitude          2135 non-null   float64       
 3   transactionDate    2135 non-null   datetime64[ns]
 4   district_cat       2135 non-null   int64         
 5   MeterPrice         2135 non-null   int64         
 6   transaction_year   2135 non-null   int32         
 7   transaction_month  2135 non-null   int32         
 8   days_since_tx      2135 non-null   int64         
 9   dist_to_center_km  2135 non-null   float64       
dtypes: datetime64[ns](1), float64(3), int32(2), int64(4)
memory usage: 150.2 KB

5️⃣ Exploratory Data Analysis (EDA)¶

In [6]:
plt.figure(figsize=(10, 4))
sns.histplot(df['MeterPrice'], bins=50, kde=True)
plt.title("Distribution of Meter Price")
plt.xlabel("Meter Price")
plt.ylabel("Frequency")
plt.show()
No description has been provided for this image
In [7]:
plt.figure(figsize=(8, 5))
sns.scatterplot(data=df, x='dist_to_center_km', y='MeterPrice')
plt.title("Meter Price vs Distance from Center")
plt.xlabel("Distance to Jeddah Center (km)")
plt.ylabel("Meter Price")
plt.show()
No description has been provided for this image
In [8]:
plt.figure(figsize=(10, 6))
corr = df.select_dtypes(include=[np.number]).corr()
sns.heatmap(corr, annot=True, cmap="coolwarm", fmt=".2f")
plt.title("Correlation Heatmap")
plt.show()
No description has been provided for this image

6️⃣Remove Duplicates¶

In [9]:
df = df.dropna(subset=['MeterPrice', 'district_cat'])

7️⃣ Feature and Target Selection (or Feature Extraction)¶

In [10]:
features = ['latitude', 'longitude', 'district_cat', 'dist_to_center_km',
            'transaction_year', 'transaction_month', 'days_since_tx']
target = 'MeterPrice'
X = df[features]
y = df[target]

8️⃣Feature Engineering and Preprocessing Pipeline Definition¶

In [11]:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)

numeric_features = ['latitude', 'longitude', 'dist_to_center_km', 'transaction_year', 'transaction_month', 'days_since_tx']
categorical_features = ['district_cat']
numeric_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy="mean")),
    ('scaler', StandardScaler())
])
categorical_transformer = Pipeline([
    ('imputer', SimpleImputer(strategy="most_frequent")),
    ('encoder', OneHotEncoder(handle_unknown="ignore"))
])
preprocessor = ColumnTransformer([
    ('num', numeric_transformer, numeric_features),
    ('cat', categorical_transformer, categorical_features)
])

9️⃣ Training Models Used¶

In [12]:
models = {
    'LinearRegression': LinearRegression(),
    'Ridge': Ridge(),
    'Lasso': Lasso(),
    'RandomForest': RandomForestRegressor(n_estimators=50, random_state=42),
    'GradientBoosting': GradientBoostingRegressor(n_estimators=50, random_state=42),
    'KNN': KNeighborsRegressor(n_neighbors=5),
    'DecisionTree': DecisionTreeRegressor(random_state=42),
    'XGBoost': xgb.XGBRegressor(n_estimators=50, random_state=42)
}

🔟 Model Training and Evaluation¶

In [13]:
results = []
for name, model in models.items():
    pipe = Pipeline([
        ('preprocessor', preprocessor),
        ('regressor', model)
    ])
    pipe.fit(X_train, y_train)
    train_pred = pipe.predict(X_train)
    test_pred = pipe.predict(X_test)
    results.append({
        'Model': name,
        'Train_R2': r2_score(y_train, train_pred),
        'Test_R2': r2_score(y_test, test_pred),
        'Train_RMSE': np.sqrt(mean_squared_error(y_train, train_pred)),
        'Test_RMSE': np.sqrt(mean_squared_error(y_test, test_pred))
    })

1️⃣1️⃣ Model Performance Comparison and Overfitting Analysis¶

In [14]:
result_df = pd.DataFrame(results)
result_df['Overfit_Gap_R2'] = result_df['Train_R2'] - result_df['Test_R2']
result_df = result_df.sort_values(by='Overfit_Gap_R2', ascending=False)
result_df
Out[14]:
Model Train_R2 Test_R2 Train_RMSE Test_RMSE Overfit_Gap_R2
6 DecisionTree 0.999995 -0.622522 6.843859 3225.423847 1.622517
7 XGBoost 0.936102 0.401738 766.288376 1958.560568 0.534364
3 RandomForest 0.930509 0.403321 799.124167 1955.968856 0.527188
4 GradientBoosting 0.640733 0.447858 1817.014198 1881.555499 0.192876
5 KNN 0.617690 0.438314 1874.379103 1897.747303 0.179377
2 Lasso 0.408701 0.421994 2331.059078 1925.119904 -0.013293
1 Ridge 0.408772 0.422434 2330.919685 1924.385775 -0.013663
0 LinearRegression 0.408780 0.422541 2330.902482 1924.208536 -0.013760

1️⃣2️⃣ Visual Comparison of Model Generalization (Train vs. Test R²)¶

In [15]:
plt.figure(figsize=(14, 6))
bar_width = 0.35
index = np.arange(len(result_df))

# Custom colors
train_color = '#1f77b4'      # Standard blue
test_color = '#aec7e8'       # Light blue (less blue)

# Plot with custom colors
plt.bar(index, result_df["Train_R2"], bar_width, label='Train R²', color=train_color)
plt.bar(index + bar_width, result_df["Test_R2"], bar_width, label='Test R²', color=test_color)

# Labels and formatting
plt.xlabel("Model")
plt.ylabel("R² Score")
plt.title("Overfitting Analysis (Train vs Test R²)")
plt.xticks(index + bar_width / 2, result_df["Model"], rotation=90)
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image

1️⃣3️⃣ Visual Comparison of Model Generalization (Train vs. Test RMSE)¶

In [16]:
plt.figure(figsize=(14, 6))
bar_width = 0.35
index = np.arange(len(result_df))

# Custom colors
train_color = '#1f77b4'   # Blue
test_color = '#aec7e8'    # Light blue

# Plot with custom colors
plt.bar(index, result_df["Train_RMSE"], bar_width, label='Train RMSE', color=train_color)
plt.bar(index + bar_width, result_df["Test_RMSE"], bar_width, label='Test RMSE', color=test_color)

# Labels and formatting
plt.xlabel("Model")
plt.ylabel("RMSE")
plt.title("Overfitting Analysis (Train vs Test RMSE)")
plt.xticks(index + bar_width / 2, result_df["Model"], rotation=45)
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
No description has been provided for this image