{
  "cells": [
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "96uNKNLST9pb"
      },
      "source": [
        "# IRIS Dataset – Machine Learning Assignment\n",
        "\n",
        "This notebook completes the required IRIS dataset tasks:\n",
        "\n",
        "- Connect Google Colab to Google Drive\n",
        "- Load the dataset using Pandas\n",
        "- Clean and explore the dataset\n",
        "- Display class counts and missing values\n",
        "- Plot histograms and scatterplots\n",
        "- Compute correlation matrix and heatmap\n",
        "- Encode the class variable\n",
        "- Split data into 70% training and 30% testing\n",
        "- Train Logistic Regression, KNN, and Decision Tree models\n",
        "- Compare model accuracies\n",
        "- Change hyperparameters and compare results\n",
        "\n",
        "> **Note:** The notebook first tries to load `Iris.csv` from Google Drive.  \n",
        "> If the file is not found, it automatically loads the standard Iris dataset from Scikit-learn."
      ],
      "id": "96uNKNLST9pb"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "5Cy01AsUT9pg"
      },
      "source": [
        "## 1. Create a folder in Google Drive and upload the dataset\n",
        "\n",
        "Create a folder named:\n",
        "\n",
        "`IRIS_ML_Assignment`\n",
        "\n",
        "Upload the dataset file into that folder and name it:\n",
        "\n",
        "`Iris.csv`\n",
        "\n",
        "Expected path:\n",
        "\n",
        "`/content/drive/MyDrive/IRIS_ML_Assignment/Iris.csv`"
      ],
      "id": "5Cy01AsUT9pg"
    },
    {
      "cell_type": "code",
      "execution_count": 1,
      "metadata": {
        "id": "gx_fxNfdT9ph",
        "colab": {
          "base_uri": "https://localhost:8080/",
          "height": 327
        },
        "outputId": "2fb82546-673a-4abc-9c24-fa3da1f5b5ef"
      },
      "outputs": [
        {
          "output_type": "error",
          "ename": "MessageError",
          "evalue": "Error: credential propagation was unsuccessful",
          "traceback": [
            "\u001b[0;31m---------------------------------------------------------------------------\u001b[0m",
            "\u001b[0;31mMessageError\u001b[0m                              Traceback (most recent call last)",
            "\u001b[0;32m/tmp/ipykernel_2092/3470177505.py\u001b[0m in \u001b[0;36m<cell line: 0>\u001b[0;34m()\u001b[0m\n\u001b[1;32m      1\u001b[0m \u001b[0;31m# Connect Google Colab to Google Drive\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m      2\u001b[0m \u001b[0;32mfrom\u001b[0m \u001b[0mgoogle\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mcolab\u001b[0m \u001b[0;32mimport\u001b[0m \u001b[0mdrive\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m----> 3\u001b[0;31m \u001b[0mdrive\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mmount\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m'/content/drive'\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
            "\u001b[0;32m/usr/local/lib/python3.12/dist-packages/google/colab/drive.py\u001b[0m in \u001b[0;36mmount\u001b[0;34m(mountpoint, force_remount, timeout_ms, readonly)\u001b[0m\n\u001b[1;32m     95\u001b[0m \u001b[0;32mdef\u001b[0m \u001b[0mmount\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mmountpoint\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mforce_remount\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mFalse\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mtimeout_ms\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;36m120000\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mreadonly\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mFalse\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     96\u001b[0m   \u001b[0;34m\"\"\"Mount your Google Drive at the specified mountpoint path.\"\"\"\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m---> 97\u001b[0;31m   return _mount(\n\u001b[0m\u001b[1;32m     98\u001b[0m       \u001b[0mmountpoint\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m     99\u001b[0m       \u001b[0mforce_remount\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mforce_remount\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n",
            "\u001b[0;32m/usr/local/lib/python3.12/dist-packages/google/colab/drive.py\u001b[0m in \u001b[0;36m_mount\u001b[0;34m(mountpoint, force_remount, timeout_ms, ephemeral, readonly)\u001b[0m\n\u001b[1;32m    132\u001b[0m   )\n\u001b[1;32m    133\u001b[0m   \u001b[0;32mif\u001b[0m \u001b[0mephemeral\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 134\u001b[0;31m     _message.blocking_request(\n\u001b[0m\u001b[1;32m    135\u001b[0m         \u001b[0;34m'request_auth'\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    136\u001b[0m         \u001b[0mrequest\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;34m{\u001b[0m\u001b[0;34m'authType'\u001b[0m\u001b[0;34m:\u001b[0m \u001b[0;34m'dfs_ephemeral'\u001b[0m\u001b[0;34m}\u001b[0m\u001b[0;34m,\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n",
            "\u001b[0;32m/usr/local/lib/python3.12/dist-packages/google/colab/_message.py\u001b[0m in \u001b[0;36mblocking_request\u001b[0;34m(request_type, request, timeout_sec, parent)\u001b[0m\n\u001b[1;32m    174\u001b[0m       \u001b[0mrequest_type\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mrequest\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mparent\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0mparent\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mexpect_reply\u001b[0m\u001b[0;34m=\u001b[0m\u001b[0;32mTrue\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    175\u001b[0m   )\n\u001b[0;32m--> 176\u001b[0;31m   \u001b[0;32mreturn\u001b[0m \u001b[0mread_reply_from_input\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mrequest_id\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0mtimeout_sec\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m",
            "\u001b[0;32m/usr/local/lib/python3.12/dist-packages/google/colab/_message.py\u001b[0m in \u001b[0;36mread_reply_from_input\u001b[0;34m(message_id, timeout_sec)\u001b[0m\n\u001b[1;32m    101\u001b[0m     ):\n\u001b[1;32m    102\u001b[0m       \u001b[0;32mif\u001b[0m \u001b[0;34m'error'\u001b[0m \u001b[0;32min\u001b[0m \u001b[0mreply\u001b[0m\u001b[0;34m:\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0;32m--> 103\u001b[0;31m         \u001b[0;32mraise\u001b[0m \u001b[0mMessageError\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0mreply\u001b[0m\u001b[0;34m[\u001b[0m\u001b[0;34m'error'\u001b[0m\u001b[0;34m]\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[0m\u001b[1;32m    104\u001b[0m       \u001b[0;32mreturn\u001b[0m \u001b[0mreply\u001b[0m\u001b[0;34m.\u001b[0m\u001b[0mget\u001b[0m\u001b[0;34m(\u001b[0m\u001b[0;34m'data'\u001b[0m\u001b[0;34m,\u001b[0m \u001b[0;32mNone\u001b[0m\u001b[0;34m)\u001b[0m\u001b[0;34m\u001b[0m\u001b[0;34m\u001b[0m\u001b[0m\n\u001b[1;32m    105\u001b[0m \u001b[0;34m\u001b[0m\u001b[0m\n",
            "\u001b[0;31mMessageError\u001b[0m: Error: credential propagation was unsuccessful"
          ]
        }
      ],
      "source": [
        "# Connect Google Colab to Google Drive\n",
        "from google.colab import drive\n",
        "drive.mount('/content/drive')"
      ],
      "id": "gx_fxNfdT9ph"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "mtUHpgRpT9pj"
      },
      "source": [
        "## 2. Import the necessary modules"
      ],
      "id": "mtUHpgRpT9pj"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "V8fA6puXT9pj"
      },
      "outputs": [],
      "source": [
        "import os\n",
        "import warnings\n",
        "warnings.filterwarnings('ignore')\n",
        "\n",
        "import numpy as np\n",
        "import pandas as pd\n",
        "import matplotlib.pyplot as plt\n",
        "import seaborn as sns\n",
        "\n",
        "from sklearn.datasets import load_iris\n",
        "from sklearn.preprocessing import LabelEncoder, StandardScaler\n",
        "from sklearn.model_selection import train_test_split\n",
        "from sklearn.pipeline import Pipeline\n",
        "from sklearn.linear_model import LogisticRegression\n",
        "from sklearn.neighbors import KNeighborsClassifier\n",
        "from sklearn.tree import DecisionTreeClassifier, plot_tree\n",
        "from sklearn.metrics import accuracy_score, classification_report, confusion_matrix\n",
        "\n",
        "pd.set_option('display.max_columns', None)\n",
        "sns.set_context('notebook')"
      ],
      "id": "V8fA6puXT9pj"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "g5h9iEaiT9pk"
      },
      "source": [
        "## 3. Load the dataset using Pandas"
      ],
      "id": "g5h9iEaiT9pk"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "Qj5xvkSaT9pk"
      },
      "outputs": [],
      "source": [
        "# Path to the uploaded dataset\n",
        "file_path = '/content/drive/MyDrive/IRIS_ML_Assignment/Iris.csv'\n",
        "\n",
        "if os.path.exists(file_path):\n",
        "    df = pd.read_csv(file_path)\n",
        "    print('Dataset loaded successfully from Google Drive.')\n",
        "else:\n",
        "    print('Iris.csv was not found in Google Drive.')\n",
        "    print('The standard Iris dataset will be loaded from Scikit-learn instead.')\n",
        "\n",
        "    iris_data = load_iris(as_frame=True)\n",
        "    df = iris_data.frame.copy()\n",
        "\n",
        "    # Rename columns to commonly used Iris dataset names\n",
        "    df.columns = [\n",
        "        'SepalLengthCm',\n",
        "        'SepalWidthCm',\n",
        "        'PetalLengthCm',\n",
        "        'PetalWidthCm',\n",
        "        'target'\n",
        "    ]\n",
        "\n",
        "    # Convert numeric target to class names\n",
        "    class_map = {\n",
        "        0: 'Iris-setosa',\n",
        "        1: 'Iris-versicolor',\n",
        "        2: 'Iris-virginica'\n",
        "    }\n",
        "    df['Species'] = df['target'].map(class_map)\n",
        "    df.drop(columns=['target'], inplace=True)\n",
        "\n",
        "    # Add an ID column to match common Iris.csv versions\n",
        "    df.insert(0, 'Id', range(1, len(df) + 1))\n",
        "\n",
        "print('Dataset shape:', df.shape)\n",
        "df.head()"
      ],
      "id": "Qj5xvkSaT9pk"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "yf3MUShmT9pl"
      },
      "source": [
        "## 4. Delete the ID column"
      ],
      "id": "yf3MUShmT9pl"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "trN6u_yBT9pl"
      },
      "outputs": [],
      "source": [
        "# Delete the ID column if it exists\n",
        "id_columns = [col for col in df.columns if col.strip().lower() in ['id', 'index']]\n",
        "\n",
        "if id_columns:\n",
        "    df.drop(columns=id_columns, inplace=True)\n",
        "    print('Deleted ID column(s):', id_columns)\n",
        "else:\n",
        "    print('No ID column found.')\n",
        "\n",
        "print('Current columns:')\n",
        "print(df.columns.tolist())"
      ],
      "id": "trN6u_yBT9pl"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "okWk_8JJT9pm"
      },
      "source": [
        "## 5. Display the first 10 rows"
      ],
      "id": "okWk_8JJT9pm"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "VezNnLfpT9pm"
      },
      "outputs": [],
      "source": [
        "df.head(10)"
      ],
      "id": "VezNnLfpT9pm"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "wKYJXnJnT9pm"
      },
      "source": [
        "## 6. Display a description of the dataset using Pandas"
      ],
      "id": "wKYJXnJnT9pm"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "kAG376OWT9pm"
      },
      "outputs": [],
      "source": [
        "df.describe(include='all')"
      ],
      "id": "kAG376OWT9pm"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "o3VvjtXsT9pm"
      },
      "source": [
        "## 7. Display information about the attributes"
      ],
      "id": "o3VvjtXsT9pm"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "0OhbU9GIT9pn"
      },
      "outputs": [],
      "source": [
        "df.info()"
      ],
      "id": "0OhbU9GIT9pn"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "wYbp48xpT9pn"
      },
      "source": [
        "## 8. Identify the class column"
      ],
      "id": "wYbp48xpT9pn"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "YgCtwXgxT9pn"
      },
      "outputs": [],
      "source": [
        "# Detect the class column automatically\n",
        "possible_class_columns = ['Species', 'species', 'Class', 'class', 'target', 'Target']\n",
        "\n",
        "class_column = None\n",
        "for col in possible_class_columns:\n",
        "    if col in df.columns:\n",
        "        class_column = col\n",
        "        break\n",
        "\n",
        "if class_column is None:\n",
        "    # Use the last non-numeric column as a fallback\n",
        "    non_numeric_columns = df.select_dtypes(exclude=np.number).columns.tolist()\n",
        "    if non_numeric_columns:\n",
        "        class_column = non_numeric_columns[-1]\n",
        "    else:\n",
        "        class_column = df.columns[-1]\n",
        "\n",
        "print('Class column:', class_column)\n",
        "print('Unique classes:', df[class_column].unique())"
      ],
      "id": "YgCtwXgxT9pn"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "KCbs4xLeT9pn"
      },
      "source": [
        "## 9. Display the number of samples in each class"
      ],
      "id": "KCbs4xLeT9pn"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "FxK0j8fTT9pn"
      },
      "outputs": [],
      "source": [
        "class_counts = df[class_column].value_counts()\n",
        "print(class_counts)"
      ],
      "id": "FxK0j8fTT9pn"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "rC2WPr8uT9pn"
      },
      "source": [
        "## 10. Check for null values"
      ],
      "id": "rC2WPr8uT9pn"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "EcQhOdozT9pn"
      },
      "outputs": [],
      "source": [
        "null_values = df.isnull().sum()\n",
        "print('Null values in each column:')\n",
        "print(null_values)\n",
        "\n",
        "print('\\nTotal null values:', df.isnull().sum().sum())"
      ],
      "id": "EcQhOdozT9pn"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "-mkQZ6UvT9po"
      },
      "source": [
        "### Interpretation of null values\n",
        "\n",
        "If the total number of null values is **0**, the dataset is complete and no missing-value treatment is required."
      ],
      "id": "-mkQZ6UvT9po"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "Dw9-Jg5pT9po"
      },
      "source": [
        "## 11. Display a histogram for each numerical attribute"
      ],
      "id": "Dw9-Jg5pT9po"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "GXUngRSQT9po"
      },
      "outputs": [],
      "source": [
        "numeric_columns = df.select_dtypes(include=np.number).columns.tolist()\n",
        "\n",
        "df[numeric_columns].hist(\n",
        "    bins=15,\n",
        "    figsize=(12, 8),\n",
        "    edgecolor='black'\n",
        ")\n",
        "\n",
        "plt.suptitle('Histograms of Iris Attributes', fontsize=16)\n",
        "plt.tight_layout()\n",
        "plt.show()"
      ],
      "id": "GXUngRSQT9po"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "7Ra2AaKOT9po"
      },
      "source": [
        "## 12. Prepare standardized class names and colors"
      ],
      "id": "7Ra2AaKOT9po"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "nEVbOBTsT9po"
      },
      "outputs": [],
      "source": [
        "# Convert class names to lowercase text for reliable matching\n",
        "class_text = df[class_column].astype(str).str.lower()\n",
        "\n",
        "def normalize_class_name(value):\n",
        "    value = str(value).lower()\n",
        "    if 'virginica' in value:\n",
        "        return 'Virginica'\n",
        "    elif 'versicolor' in value or 'versicolour' in value:\n",
        "        return 'Versicolor'\n",
        "    elif 'setosa' in value:\n",
        "        return 'Setosa'\n",
        "    return str(value)\n",
        "\n",
        "df['Class_Name'] = df[class_column].apply(normalize_class_name)\n",
        "\n",
        "class_colors = {\n",
        "    'Virginica': 'red',\n",
        "    'Versicolor': 'orange',\n",
        "    'Setosa': 'blue'\n",
        "}\n",
        "\n",
        "print(df['Class_Name'].value_counts())"
      ],
      "id": "nEVbOBTsT9po"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "QqnUvJaLT9po"
      },
      "source": [
        "## 13. Display scatterplots for each class"
      ],
      "id": "QqnUvJaLT9po"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "_ddJPLGkT9po"
      },
      "outputs": [],
      "source": [
        "# Detect the four feature columns\n",
        "feature_lookup = {}\n",
        "\n",
        "for col in numeric_columns:\n",
        "    name = col.lower().replace(' ', '').replace('_', '')\n",
        "\n",
        "    if 'sepallength' in name:\n",
        "        feature_lookup['sepal_length'] = col\n",
        "    elif 'sepalwidth' in name:\n",
        "        feature_lookup['sepal_width'] = col\n",
        "    elif 'petallength' in name:\n",
        "        feature_lookup['petal_length'] = col\n",
        "    elif 'petalwidth' in name:\n",
        "        feature_lookup['petal_width'] = col\n",
        "\n",
        "required_features = ['sepal_length', 'sepal_width', 'petal_length', 'petal_width']\n",
        "\n",
        "missing_features = [f for f in required_features if f not in feature_lookup]\n",
        "\n",
        "if missing_features:\n",
        "    raise ValueError(f'Could not detect these feature columns: {missing_features}')\n",
        "\n",
        "print('Detected features:')\n",
        "for key, value in feature_lookup.items():\n",
        "    print(f'{key}: {value}')"
      ],
      "id": "_ddJPLGkT9po"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "OnK-AYFST9po"
      },
      "outputs": [],
      "source": [
        "scatter_pairs = [\n",
        "    ('sepal_length', 'sepal_width', 'Sepal Length vs Sepal Width'),\n",
        "    ('petal_length', 'petal_width', 'Petal Length vs Petal Width'),\n",
        "    ('sepal_length', 'petal_length', 'Sepal Length vs Petal Length'),\n",
        "    ('sepal_width', 'petal_width', 'Sepal Width vs Petal Width')\n",
        "]\n",
        "\n",
        "for x_key, y_key, title in scatter_pairs:\n",
        "    plt.figure(figsize=(8, 6))\n",
        "\n",
        "    for class_name in ['Virginica', 'Versicolor', 'Setosa']:\n",
        "        subset = df[df['Class_Name'] == class_name]\n",
        "\n",
        "        plt.scatter(\n",
        "            subset[feature_lookup[x_key]],\n",
        "            subset[feature_lookup[y_key]],\n",
        "            label=class_name,\n",
        "            color=class_colors[class_name],\n",
        "            alpha=0.75\n",
        "        )\n",
        "\n",
        "    plt.xlabel(feature_lookup[x_key])\n",
        "    plt.ylabel(feature_lookup[y_key])\n",
        "    plt.title(title)\n",
        "    plt.legend()\n",
        "    plt.grid(alpha=0.25)\n",
        "    plt.show()"
      ],
      "id": "OnK-AYFST9po"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "TSTJaubMT9po"
      },
      "source": [
        "## 14. Interpretation of the scatterplots\n",
        "\n",
        "The scatterplots show that:\n",
        "\n",
        "- **Iris Setosa** is clearly separated from the other two classes, especially when petal length or petal width is used.\n",
        "- **Iris Versicolor** and **Iris Virginica** overlap in several plots.\n",
        "- Petal measurements provide stronger class separation than sepal measurements.\n",
        "- Sepal length versus sepal width shows more overlap, so these two features alone are less effective for classification.\n",
        "- Petal length versus petal width gives the clearest separation among the three classes."
      ],
      "id": "TSTJaubMT9po"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "gqO4zg9HT9po"
      },
      "source": [
        "## 15. Display the correlation matrix"
      ],
      "id": "gqO4zg9HT9po"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "djl8yCZ6T9pp"
      },
      "outputs": [],
      "source": [
        "correlation_matrix = df[numeric_columns].corr()\n",
        "correlation_matrix"
      ],
      "id": "djl8yCZ6T9pp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "Ufg9teDuT9pp"
      },
      "source": [
        "## 16. Display a heatmap for the correlation matrix"
      ],
      "id": "Ufg9teDuT9pp"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "MWLF26N3T9pp"
      },
      "outputs": [],
      "source": [
        "plt.figure(figsize=(9, 7))\n",
        "sns.heatmap(\n",
        "    correlation_matrix,\n",
        "    annot=True,\n",
        "    fmt='.2f',\n",
        "    cmap='coolwarm',\n",
        "    square=True,\n",
        "    linewidths=0.5\n",
        ")\n",
        "\n",
        "plt.title('Correlation Matrix Heatmap')\n",
        "plt.show()"
      ],
      "id": "MWLF26N3T9pp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "tw-VUI4zT9pp"
      },
      "source": [
        "## 17. Interpretation of the correlation matrix\n",
        "\n",
        "- Correlation values close to **+1** indicate a strong positive relationship.\n",
        "- Values close to **-1** indicate a strong negative relationship.\n",
        "- Values close to **0** indicate a weak linear relationship.\n",
        "- Petal length and petal width usually have a very strong positive correlation.\n",
        "- Sepal length is also positively correlated with petal length and petal width.\n",
        "- Sepal width usually has weaker or negative correlations with some other attributes.\n",
        "- Strongly correlated petal features are useful for separating the Iris classes."
      ],
      "id": "tw-VUI4zT9pp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "ArpSZgANT9pp"
      },
      "source": [
        "## 18. Encode the class variable using LabelEncoder"
      ],
      "id": "ArpSZgANT9pp"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "DyhYZS7CT9pp"
      },
      "outputs": [],
      "source": [
        "label_encoder = LabelEncoder()\n",
        "\n",
        "df['Class_Encoded'] = label_encoder.fit_transform(df[class_column])\n",
        "\n",
        "encoding_table = pd.DataFrame({\n",
        "    'Class': label_encoder.classes_,\n",
        "    'Code': label_encoder.transform(label_encoder.classes_)\n",
        "})\n",
        "\n",
        "encoding_table"
      ],
      "id": "DyhYZS7CT9pp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "45NJQn1ET9pp"
      },
      "source": [
        "## 19. What is the code of the third class?"
      ],
      "id": "45NJQn1ET9pp"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "MZ7XNHNbT9pp"
      },
      "outputs": [],
      "source": [
        "third_class = label_encoder.classes_[2]\n",
        "third_class_code = label_encoder.transform([third_class])[0]\n",
        "\n",
        "print('Third class:', third_class)\n",
        "print('Code of the third class:', third_class_code)"
      ],
      "id": "MZ7XNHNbT9pp"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "RzOle25NT9pq"
      },
      "source": [
        "### Answer\n",
        "\n",
        "With Scikit-learn `LabelEncoder`, classes are sorted alphabetically before encoding.  \n",
        "Therefore, the third class receives the code shown in the output above, which is normally **2**."
      ],
      "id": "RzOle25NT9pq"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "SIMEydibT9pq"
      },
      "source": [
        "## 20. Split the dataset into 70% training and 30% testing"
      ],
      "id": "SIMEydibT9pq"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "QcbhTp-ST9pr"
      },
      "outputs": [],
      "source": [
        "X = df[numeric_columns]\n",
        "y = df['Class_Encoded']\n",
        "\n",
        "X_train, X_test, y_train, y_test = train_test_split(\n",
        "    X,\n",
        "    y,\n",
        "    test_size=0.30,\n",
        "    random_state=42,\n",
        "    stratify=y\n",
        ")\n",
        "\n",
        "print('Training samples:', X_train.shape[0])\n",
        "print('Testing samples:', X_test.shape[0])"
      ],
      "id": "QcbhTp-ST9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "PVkC2G2nT9pr"
      },
      "source": [
        "## 21. Train a Logistic Regression model"
      ],
      "id": "PVkC2G2nT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "NUaCeWG9T9pr"
      },
      "outputs": [],
      "source": [
        "logistic_model = Pipeline([\n",
        "    ('scaler', StandardScaler()),\n",
        "    ('classifier', LogisticRegression(max_iter=1000, random_state=42))\n",
        "])\n",
        "\n",
        "logistic_model.fit(X_train, y_train)\n",
        "logistic_predictions = logistic_model.predict(X_test)\n",
        "logistic_accuracy = accuracy_score(y_test, logistic_predictions)\n",
        "\n",
        "print('Logistic Regression Accuracy:', logistic_accuracy)\n",
        "print('\\nClassification Report:')\n",
        "print(classification_report(y_test, logistic_predictions, target_names=label_encoder.classes_))"
      ],
      "id": "NUaCeWG9T9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "IZJhm8rzT9pr"
      },
      "source": [
        "## 22. Train a KNN model"
      ],
      "id": "IZJhm8rzT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "wGEhH6AET9pr"
      },
      "outputs": [],
      "source": [
        "knn_model = Pipeline([\n",
        "    ('scaler', StandardScaler()),\n",
        "    ('classifier', KNeighborsClassifier(n_neighbors=5))\n",
        "])\n",
        "\n",
        "knn_model.fit(X_train, y_train)\n",
        "knn_predictions = knn_model.predict(X_test)\n",
        "knn_accuracy = accuracy_score(y_test, knn_predictions)\n",
        "\n",
        "print('KNN Accuracy:', knn_accuracy)\n",
        "print('\\nClassification Report:')\n",
        "print(classification_report(y_test, knn_predictions, target_names=label_encoder.classes_))"
      ],
      "id": "wGEhH6AET9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "Ast12yFVT9pr"
      },
      "source": [
        "## 23. Train a Decision Tree model"
      ],
      "id": "Ast12yFVT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "FVUSg9euT9pr"
      },
      "outputs": [],
      "source": [
        "tree_model = DecisionTreeClassifier(\n",
        "    random_state=42\n",
        ")\n",
        "\n",
        "tree_model.fit(X_train, y_train)\n",
        "tree_predictions = tree_model.predict(X_test)\n",
        "tree_accuracy = accuracy_score(y_test, tree_predictions)\n",
        "\n",
        "print('Decision Tree Accuracy:', tree_accuracy)\n",
        "print('\\nClassification Report:')\n",
        "print(classification_report(y_test, tree_predictions, target_names=label_encoder.classes_))"
      ],
      "id": "FVUSg9euT9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "D1UTB5uRT9pr"
      },
      "source": [
        "## 24. Compare the accuracies of the three models"
      ],
      "id": "D1UTB5uRT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "jxR7rMxhT9pr"
      },
      "outputs": [],
      "source": [
        "baseline_results = pd.DataFrame({\n",
        "    'Model': [\n",
        "        'Logistic Regression',\n",
        "        'KNN',\n",
        "        'Decision Tree'\n",
        "    ],\n",
        "    'Accuracy': [\n",
        "        logistic_accuracy,\n",
        "        knn_accuracy,\n",
        "        tree_accuracy\n",
        "    ]\n",
        "}).sort_values(by='Accuracy', ascending=False)\n",
        "\n",
        "baseline_results"
      ],
      "id": "jxR7rMxhT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "wTrKqxz4T9pr"
      },
      "outputs": [],
      "source": [
        "plt.figure(figsize=(8, 5))\n",
        "sns.barplot(\n",
        "    data=baseline_results,\n",
        "    x='Model',\n",
        "    y='Accuracy'\n",
        ")\n",
        "\n",
        "plt.ylim(0, 1.05)\n",
        "plt.title('Baseline Model Accuracy Comparison')\n",
        "plt.xticks(rotation=15)\n",
        "plt.show()"
      ],
      "id": "wTrKqxz4T9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "uP8APg6BT9pr"
      },
      "source": [
        "## 25. Change the hyperparameters of each model"
      ],
      "id": "uP8APg6BT9pr"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "1JTrDmHST9pr"
      },
      "outputs": [],
      "source": [
        "# Tuned Logistic Regression\n",
        "tuned_logistic = Pipeline([\n",
        "    ('scaler', StandardScaler()),\n",
        "    ('classifier', LogisticRegression(\n",
        "        C=10,\n",
        "        solver='liblinear',\n",
        "        max_iter=2000,\n",
        "        random_state=42\n",
        "    ))\n",
        "])\n",
        "\n",
        "# Tuned KNN\n",
        "tuned_knn = Pipeline([\n",
        "    ('scaler', StandardScaler()),\n",
        "    ('classifier', KNeighborsClassifier(\n",
        "        n_neighbors=7,\n",
        "        weights='distance',\n",
        "        metric='minkowski',\n",
        "        p=2\n",
        "    ))\n",
        "])\n",
        "\n",
        "# Tuned Decision Tree\n",
        "tuned_tree = DecisionTreeClassifier(\n",
        "    max_depth=3,\n",
        "    min_samples_split=4,\n",
        "    min_samples_leaf=2,\n",
        "    criterion='entropy',\n",
        "    random_state=42\n",
        ")\n",
        "\n",
        "tuned_models = {\n",
        "    'Tuned Logistic Regression': tuned_logistic,\n",
        "    'Tuned KNN': tuned_knn,\n",
        "    'Tuned Decision Tree': tuned_tree\n",
        "}\n",
        "\n",
        "tuned_accuracies = {}\n",
        "\n",
        "for model_name, model in tuned_models.items():\n",
        "    model.fit(X_train, y_train)\n",
        "    predictions = model.predict(X_test)\n",
        "    tuned_accuracies[model_name] = accuracy_score(y_test, predictions)\n",
        "\n",
        "tuned_results = pd.DataFrame({\n",
        "    'Model': list(tuned_accuracies.keys()),\n",
        "    'Accuracy': list(tuned_accuracies.values())\n",
        "}).sort_values(by='Accuracy', ascending=False)\n",
        "\n",
        "tuned_results"
      ],
      "id": "1JTrDmHST9pr"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "9DjPp5_3T9ps"
      },
      "source": [
        "## 26. Compare baseline and tuned accuracies"
      ],
      "id": "9DjPp5_3T9ps"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "i1nqxShBT9ps"
      },
      "outputs": [],
      "source": [
        "comparison_results = pd.DataFrame({\n",
        "    'Model': [\n",
        "        'Logistic Regression',\n",
        "        'KNN',\n",
        "        'Decision Tree'\n",
        "    ],\n",
        "    'Baseline Accuracy': [\n",
        "        logistic_accuracy,\n",
        "        knn_accuracy,\n",
        "        tree_accuracy\n",
        "    ],\n",
        "    'Tuned Accuracy': [\n",
        "        tuned_accuracies['Tuned Logistic Regression'],\n",
        "        tuned_accuracies['Tuned KNN'],\n",
        "        tuned_accuracies['Tuned Decision Tree']\n",
        "    ]\n",
        "})\n",
        "\n",
        "comparison_results['Accuracy Change'] = (\n",
        "    comparison_results['Tuned Accuracy']\n",
        "    - comparison_results['Baseline Accuracy']\n",
        ")\n",
        "\n",
        "comparison_results"
      ],
      "id": "i1nqxShBT9ps"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "bTKJVRU2T9ps"
      },
      "outputs": [],
      "source": [
        "comparison_plot = comparison_results.melt(\n",
        "    id_vars='Model',\n",
        "    value_vars=['Baseline Accuracy', 'Tuned Accuracy'],\n",
        "    var_name='Version',\n",
        "    value_name='Accuracy'\n",
        ")\n",
        "\n",
        "plt.figure(figsize=(9, 5))\n",
        "sns.barplot(\n",
        "    data=comparison_plot,\n",
        "    x='Model',\n",
        "    y='Accuracy',\n",
        "    hue='Version'\n",
        ")\n",
        "\n",
        "plt.ylim(0, 1.05)\n",
        "plt.title('Baseline vs Tuned Model Accuracy')\n",
        "plt.xticks(rotation=15)\n",
        "plt.show()"
      ],
      "id": "bTKJVRU2T9ps"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "nNFkVhCAT9ps"
      },
      "source": [
        "## 27. Did the accuracies change?\n",
        "\n",
        "The table above shows whether changing the hyperparameters improved, reduced, or maintained the accuracy.\n",
        "\n",
        "Possible observations:\n",
        "\n",
        "- Logistic Regression may change slightly after changing `C` and the solver.\n",
        "- KNN may improve or decrease depending on `n_neighbors`, distance weighting, and feature scaling.\n",
        "- Decision Tree accuracy may change after limiting the tree depth and controlling minimum samples.\n",
        "- Hyperparameter tuning does not always guarantee higher accuracy, especially with a small dataset such as Iris.\n",
        "- A simpler model can sometimes perform as well as or better than a more complex model."
      ],
      "id": "nNFkVhCAT9ps"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "G8p_WGVfT9ps"
      },
      "source": [
        "## 28. Confusion matrices for the baseline models"
      ],
      "id": "G8p_WGVfT9ps"
    },
    {
      "cell_type": "code",
      "execution_count": null,
      "metadata": {
        "id": "Supy0yWzT9ps"
      },
      "outputs": [],
      "source": [
        "models_and_predictions = {\n",
        "    'Logistic Regression': logistic_predictions,\n",
        "    'KNN': knn_predictions,\n",
        "    'Decision Tree': tree_predictions\n",
        "}\n",
        "\n",
        "for model_name, predictions in models_and_predictions.items():\n",
        "    matrix = confusion_matrix(y_test, predictions)\n",
        "\n",
        "    plt.figure(figsize=(6, 5))\n",
        "    sns.heatmap(\n",
        "        matrix,\n",
        "        annot=True,\n",
        "        fmt='d',\n",
        "        cmap='Blues',\n",
        "        xticklabels=label_encoder.classes_,\n",
        "        yticklabels=label_encoder.classes_\n",
        "    )\n",
        "\n",
        "    plt.title(f'Confusion Matrix - {model_name}')\n",
        "    plt.xlabel('Predicted Class')\n",
        "    plt.ylabel('Actual Class')\n",
        "    plt.show()"
      ],
      "id": "Supy0yWzT9ps"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "ixB6tAV1T9ps"
      },
      "source": [
        "## 29. Final conclusion"
      ],
      "id": "ixB6tAV1T9ps"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "zCYXOyXqT9ps"
      },
      "source": [
        "The Iris dataset contains three classes and four numerical features.  \n",
        "Exploratory analysis shows that petal length and petal width provide the clearest class separation.\n",
        "\n",
        "Three classification models were trained:\n",
        "\n",
        "1. Logistic Regression\n",
        "2. K-Nearest Neighbors\n",
        "3. Decision Tree\n",
        "\n",
        "The models were evaluated using a 70% training and 30% testing split.  \n",
        "Their baseline and tuned accuracies were compared. The best model is the one with the highest testing accuracy shown in the final comparison table.\n",
        "\n",
        "Because the Iris dataset is small and relatively clean, all three models can achieve high accuracy. However, their exact results may vary depending on the train-test split and selected hyperparameters."
      ],
      "id": "zCYXOyXqT9ps"
    },
    {
      "cell_type": "markdown",
      "metadata": {
        "id": "UiDbueOnT9ps"
      },
      "source": [
        "## 30. Download the completed notebook\n",
        "\n",
        "After running all cells in Google Colab:\n",
        "\n",
        "1. Click **File**\n",
        "2. Select **Download**\n",
        "3. Choose **Download .ipynb**\n",
        "4. Upload the downloaded `.ipynb` file to your university platform"
      ],
      "id": "UiDbueOnT9ps"
    }
  ],
  "metadata": {
    "colab": {
      "provenance": []
    },
    "kernelspec": {
      "display_name": "Python 3",
      "language": "python",
      "name": "python3"
    },
    "language_info": {
      "name": "python",
      "version": "3.x"
    }
  },
  "nbformat": 4,
  "nbformat_minor": 5
}