Supervised Learning

Supervised learning is a type of machine learning in which a model is trained using labeled data, where each input has a known output or target value.

The model learns the relationship between input features (X) and the target/output (Y) and uses this knowledge to predict the output for new, unseen data.

FeatureID3CART
Split criterionEntropy / Information GainGini impurity (default)
SplitsUsually multiwayBinary
ClassificationYesYes
RegressionNoYes
sklearn built-inNot directly as “ID3”DecisionTreeClassifier

Steps Involved in Supervised Learning

  1. Collect Data
    Gather a dataset containing input features and their corresponding output labels.
  2. Prepare the Data
    Clean the data, handle missing values, remove unnecessary data, and convert categorical data into a suitable format.
  3. Split the Dataset
    Divide the data into:
    • Training data (X)– used to train the model.
    • Testing data (y)– used to evaluate the model.
  4. Select a Model
    Choose a suitable supervised learning algorithm, such as:
    • Linear Regression
    • Decision Tree
    • K-Nearest Neighbors (KNN)
    • Support Vector Machine (SVM)
    • Naive Bayes
  5. Train the Model
    Feed the training data to the algorithm so that it learns patterns and relationships between inputs and outputs.
  6. Evaluate the Model
    Test the trained model using unseen test data and measure its performance using suitable metrics such as accuracy, precision, recall, or mean squared error.
  7. Tune the Model
    Adjust model parameters or features to improve performance and reduce overfitting.
  8. Make Predictions
    Use the trained model to predict outputs for new, unseen data.

Confusion Matrix