🏡 AI-Driven Agents for Fair and Efficient Property Valuation¶
- Student Name: Fahad Abubaker Bahashwan
- Advisor: Dr. Ashraf Elsayed
- Institution: MidOcean University
- Date: 23, July 2025
What we'll do:
- Load and explore data
- Create geographic features
- Train different ML models
- Compare which model works best
Step 1: Import Libraries¶
Import all the tools we need:
In [2]:
import pandas as pd
import numpy as np
import matplotlib.pyplot as plt
import seaborn as sns
from geopy.distance import geodesic
from sklearn.model_selection import train_test_split
from sklearn.pipeline import Pipeline
from sklearn.preprocessing import StandardScaler, OneHotEncoder
from sklearn.impute import SimpleImputer
from sklearn.compose import ColumnTransformer
from sklearn.linear_model import Ridge, Lasso, LinearRegression
from sklearn.ensemble import RandomForestRegressor, GradientBoostingRegressor
from sklearn.tree import DecisionTreeRegressor
from sklearn.neighbors import KNeighborsRegressor
from sklearn.metrics import mean_squared_error, r2_score
import xgboost as xgb
sns.set(style="whitegrid")
Step 2: Load Data¶
In [3]:
df = pd.read_excel("WorkedData6.xlsx")
df['transactionDate'] = pd.to_datetime(df['transactionDate'])
df['transaction_year'] = df['transactionDate'].dt.year
df['transaction_month'] = df['transactionDate'].dt.month
df['days_since_tx'] = (df['transactionDate'].max() - df['transactionDate']).dt.days
3️⃣ Calculate Jeddah Center Point¶
In [4]:
jeddah_center = (21.543333, 39.172778)
df['dist_to_center_km'] = df.apply(
lambda row: geodesic((row['latitude'], row['longitude']), jeddah_center).kilometers, axis=1)
4️⃣ data exploration¶
In [5]:
df.info()
<class 'pandas.core.frame.DataFrame'> RangeIndex: 2135 entries, 0 to 2134 Data columns (total 10 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 district_id 2135 non-null int64 1 latitude 2135 non-null float64 2 longitude 2135 non-null float64 3 transactionDate 2135 non-null datetime64[ns] 4 district_cat 2135 non-null int64 5 MeterPrice 2135 non-null int64 6 transaction_year 2135 non-null int32 7 transaction_month 2135 non-null int32 8 days_since_tx 2135 non-null int64 9 dist_to_center_km 2135 non-null float64 dtypes: datetime64[ns](1), float64(3), int32(2), int64(4) memory usage: 150.2 KB
5️⃣ Exploratory Data Analysis (EDA)¶
In [6]:
plt.figure(figsize=(10, 4))
sns.histplot(df['MeterPrice'], bins=50, kde=True)
plt.title("Distribution of Meter Price")
plt.xlabel("Meter Price")
plt.ylabel("Frequency")
plt.show()
In [7]:
plt.figure(figsize=(8, 5))
sns.scatterplot(data=df, x='dist_to_center_km', y='MeterPrice')
plt.title("Meter Price vs Distance from Center")
plt.xlabel("Distance to Jeddah Center (km)")
plt.ylabel("Meter Price")
plt.show()
In [8]:
plt.figure(figsize=(10, 6))
corr = df.select_dtypes(include=[np.number]).corr()
sns.heatmap(corr, annot=True, cmap="coolwarm", fmt=".2f")
plt.title("Correlation Heatmap")
plt.show()
6️⃣Remove Duplicates¶
In [9]:
df = df.dropna(subset=['MeterPrice', 'district_cat'])
7️⃣ Feature and Target Selection (or Feature Extraction)¶
In [10]:
features = ['latitude', 'longitude', 'district_cat', 'dist_to_center_km',
'transaction_year', 'transaction_month', 'days_since_tx']
target = 'MeterPrice'
X = df[features]
y = df[target]
8️⃣Feature Engineering and Preprocessing Pipeline Definition¶
In [11]:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)
numeric_features = ['latitude', 'longitude', 'dist_to_center_km', 'transaction_year', 'transaction_month', 'days_since_tx']
categorical_features = ['district_cat']
numeric_transformer = Pipeline([
('imputer', SimpleImputer(strategy="mean")),
('scaler', StandardScaler())
])
categorical_transformer = Pipeline([
('imputer', SimpleImputer(strategy="most_frequent")),
('encoder', OneHotEncoder(handle_unknown="ignore"))
])
preprocessor = ColumnTransformer([
('num', numeric_transformer, numeric_features),
('cat', categorical_transformer, categorical_features)
])
9️⃣ Training Models Used¶
In [12]:
models = {
'LinearRegression': LinearRegression(),
'Ridge': Ridge(),
'Lasso': Lasso(),
'RandomForest': RandomForestRegressor(n_estimators=50, random_state=42),
'GradientBoosting': GradientBoostingRegressor(n_estimators=50, random_state=42),
'KNN': KNeighborsRegressor(n_neighbors=5),
'DecisionTree': DecisionTreeRegressor(random_state=42),
'XGBoost': xgb.XGBRegressor(n_estimators=50, random_state=42)
}
🔟 Model Training and Evaluation¶
In [13]:
results = []
for name, model in models.items():
pipe = Pipeline([
('preprocessor', preprocessor),
('regressor', model)
])
pipe.fit(X_train, y_train)
train_pred = pipe.predict(X_train)
test_pred = pipe.predict(X_test)
results.append({
'Model': name,
'Train_R2': r2_score(y_train, train_pred),
'Test_R2': r2_score(y_test, test_pred),
'Train_RMSE': np.sqrt(mean_squared_error(y_train, train_pred)),
'Test_RMSE': np.sqrt(mean_squared_error(y_test, test_pred))
})
1️⃣1️⃣ Model Performance Comparison and Overfitting Analysis¶
In [14]:
result_df = pd.DataFrame(results)
result_df['Overfit_Gap_R2'] = result_df['Train_R2'] - result_df['Test_R2']
result_df = result_df.sort_values(by='Overfit_Gap_R2', ascending=False)
result_df
Out[14]:
| Model | Train_R2 | Test_R2 | Train_RMSE | Test_RMSE | Overfit_Gap_R2 | |
|---|---|---|---|---|---|---|
| 6 | DecisionTree | 0.999995 | -0.622522 | 6.843859 | 3225.423847 | 1.622517 |
| 7 | XGBoost | 0.936102 | 0.401738 | 766.288376 | 1958.560568 | 0.534364 |
| 3 | RandomForest | 0.930509 | 0.403321 | 799.124167 | 1955.968856 | 0.527188 |
| 4 | GradientBoosting | 0.640733 | 0.447858 | 1817.014198 | 1881.555499 | 0.192876 |
| 5 | KNN | 0.617690 | 0.438314 | 1874.379103 | 1897.747303 | 0.179377 |
| 2 | Lasso | 0.408701 | 0.421994 | 2331.059078 | 1925.119904 | -0.013293 |
| 1 | Ridge | 0.408772 | 0.422434 | 2330.919685 | 1924.385775 | -0.013663 |
| 0 | LinearRegression | 0.408780 | 0.422541 | 2330.902482 | 1924.208536 | -0.013760 |
1️⃣2️⃣ Visual Comparison of Model Generalization (Train vs. Test R²)¶
In [15]:
plt.figure(figsize=(14, 6))
bar_width = 0.35
index = np.arange(len(result_df))
# Custom colors
train_color = '#1f77b4' # Standard blue
test_color = '#aec7e8' # Light blue (less blue)
# Plot with custom colors
plt.bar(index, result_df["Train_R2"], bar_width, label='Train R²', color=train_color)
plt.bar(index + bar_width, result_df["Test_R2"], bar_width, label='Test R²', color=test_color)
# Labels and formatting
plt.xlabel("Model")
plt.ylabel("R² Score")
plt.title("Overfitting Analysis (Train vs Test R²)")
plt.xticks(index + bar_width / 2, result_df["Model"], rotation=90)
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()
1️⃣3️⃣ Visual Comparison of Model Generalization (Train vs. Test RMSE)¶
In [16]:
plt.figure(figsize=(14, 6))
bar_width = 0.35
index = np.arange(len(result_df))
# Custom colors
train_color = '#1f77b4' # Blue
test_color = '#aec7e8' # Light blue
# Plot with custom colors
plt.bar(index, result_df["Train_RMSE"], bar_width, label='Train RMSE', color=train_color)
plt.bar(index + bar_width, result_df["Test_RMSE"], bar_width, label='Test RMSE', color=test_color)
# Labels and formatting
plt.xlabel("Model")
plt.ylabel("RMSE")
plt.title("Overfitting Analysis (Train vs Test RMSE)")
plt.xticks(index + bar_width / 2, result_df["Model"], rotation=45)
plt.legend()
plt.grid(True)
plt.tight_layout()
plt.show()