Designing a K-Nearest Neighbour (KNN) Exam-Failure Classifier

Designing a K-Nearest Neighbour (KNN) Exam-Failure Classifier

Verified Sources
Oct 5, 2026

To predict whether a student will pass or fail the exam using K-Nearest Neighbour (K-Nearest Neighbour), we treat each student as a feature vector and the exam outcome as a class label. KNN is non-parametric: for a new student, it finds the kk closest training students using a distance metric and assigns the class most common among those neighbours (optionally using distance-weighting). This overall decision rule is consistent with KNN classifier descriptions in common ML references and KNN implementations. 2

From the table, we build inputs from:

  • Language (Java/C++/Python)
  • Passed all Assignments (Yes/No)
  • GPA

The target is:

  • Passed Exam (Yes/No)

KNN design depends heavily on (i) how we encode categorical variables, (ii) feature scaling, and (iii) selecting kk and the distance metric. Since KNN uses distances in feature space, features on different scales can dominate the distance calculations; thus scaling is a key preprocessing step.

Footnotes

  1. K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩

  2. KNeighborsClassifier — scikit-learn predict/parameter behavior - Covers weights (uniform vs distance) and distance metric details including Minkowski. ↩

  3. What is Feature Scaling and Why is it Important? - Analytics Vidhya - Discusses scaling importance for KNN and distance-based algorithms. ↩

KNN design workflow for this exam dataset

Define features and label

1. Data & target

Use Language, Passed all Assignments, GPA → predict Passed Exam."

Encode & scale

2. Preprocessing

Encode Language; keep assignments as binary; scale GPA."

Distance metric + vote rule

3. Distance & voting

Choose metric (e.g., Euclidean via Minkowski) and weights (uniform vs distance)."

Select k

4. Hyperparameters

Evaluate candidate k values using cross-validation."

Fit + inference

5. Train and predict

Store training vectors; for each query, compute nearest neighbours and vote."

Data preparation and feature engineering

1) Define the supervised learning problem

Let each student ii be represented by a feature vector xi∈Rdx_i \in \mathbb{R}^d and label yi∈{0,1}y_i \in \{0,1\}:

  • yi=1y_i = 1 if Passed Exam = Yes, otherwise 00.

KNN stores the training pairs (xi,yi)(x_i, y_i); during prediction it uses the labels of the nearest stored vectors.

2) Encode categorical features (Language)

KNN requires numeric vectors. Encode Language using one-hot encoding into three binary features:

  • IJavaI_{\text{Java}}, IC++I_{\text{C++}}, IPythonI_{\text{Python}} (only one of them is 1 for each student).

Keep Passed all Assignments as a binary indicator:

  • IAssignmentsYes∈{0,1}I_{\text{AssignmentsYes}} \in \{0,1\}

Now each student becomes: xi=[IJava,IC++,IPython,IAssignmentsYes,GPA]x_i = [I_{\text{Java}}, I_{\text{C++}}, I_{\text{Python}}, I_{\text{AssignmentsYes}}, \text{GPA}] so here d=5d=5.

Pro Tip: With KNN, categorical encoding choices can change the geometry of the feature space. One-hot encoding is standard for Euclidean-style distances, because categories become orthogonal indicator directions.

3) Scale numeric features (especially GPA)

Because KNN relies on distances, features with larger numeric ranges can disproportionately influence the result; scaling reduces this effect and improves distance fairness. A typical approach is standardization or normalization applied to numeric features like GPA before training and prediction.

In this dataset, GPA is numeric while the other features are 0/1. Even so, scaling GPA is still recommended to ensure stable behaviour and prevent it from overwhelming or being overwhelmed by binary dimensions.

Footnotes

  1. K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩

  2. What is Feature Scaling and Why is it Important? - Analytics Vidhya - Discusses scaling importance for KNN and distance-based algorithms. ↩ ↩2

KNN prediction algorithm (Pass/Fail) for a new student

  1. 1
    Step 1

    Encode Language with the same one-hot scheme as training and convert Passed all Assignments to 0/1; scale GPA using the training-fitted scaler.

  2. 2
    Step 2

    For each training vector xix_i, compute dist(xi,xquery)dist(x_i, x_{query}) using the chosen metric (e.g., Euclidean is Minkowski with p=2p=2).

  3. 3
    Step 3

    Sort (or partial-select) by distance and keep the kk closest training students.

  4. 4
    Step 4

    If using uniform weights, predict the majority label among the kk neighbours. If using distance weights, assign larger influence to closer points (e.g., inverse-distance weighting).

  5. 5
    Step 5

    Return PassPass if predicted class is 1, else FailFail.

Distance metric and voting design

Distance metric

A common default in KNN implementations is Minkowski distance; scikit-style KNN uses Minkowski and typically supports Euclidean by setting p=2p=2. 2

For two students xx and zz:

  • Choose p=2p=2 for Euclidean distance: dist(x,z)=∑j=1d(xj−zj)2dist(x,z)=\sqrt{\sum_{j=1}^{d}(x_j-z_j)^2} Other norms (e.g., Manhattan, p=1p=1) can be explored, but the Euclidean choice is a strong baseline.

Voting rule and the role of the weights parameter

KNN classification is often implemented as a “vote” over the neighbour labels. The weights parameter typically supports:

  • uniform: all neighbours vote equally
  • distance: closer neighbours vote more (often using inverse distance)

Given this dataset is small, distance weighting can reduce sensitivity to noisy neighbours (e.g., a single close but contradictory example).

Warning: If you set kk too small, KNN becomes sensitive to noise; if kk is too large, it may smooth over local patterns. This bias–variance trade-off is why kk must be tuned using validation.

Footnotes

  1. K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩

  2. KNeighborsClassifier — scikit-learn predict/parameter behavior - Covers weights (uniform vs distance) and distance metric details including Minkowski. ↩ ↩2

  3. A Gentle Introduction to k-fold Cross-Validation - MachineLearningMastery - Explains cross-validation for hyperparameter selection and general evaluation stability. ↩

Selecting hyperparameter kk (and evaluating the design)

Cross-validation for choosing kk

Since the dataset in the prompt is small, performance estimates from a single train/test split may vary a lot. Use k-fold cross-validation to select kk (and optionally the distance metric or weights). Cross-validation is widely used to choose hyperparameters by averaging scores over multiple splits. A typical setup is:

  • Candidate values like k∈{1,3,5,7}k \in \{1,3,5,7\}
  • Use stratification so each fold keeps similar proportions of Pass/Fail labels (important for classification).

Tie-handling

For classification, if votes tie, some implementations may depend on neighbour ordering; it’s another reason to prefer odd kk for binary classification.

Footnotes

  1. A Gentle Introduction to k-fold Cross-Validation - MachineLearningMastery - Explains cross-validation for hyperparameter selection and general evaluation stability. ↩ ↩2

  2. K-Nearest Neighbours - GeeksforGeeks - Notes tie/odd-k considerations for classification and cross-validation for choosing k. ↩

Illustrative feature composition for the KNN input vector

Binary-encoded categories + scaled numeric GPA

Common edge cases & design choices

KNN Algorithm (intuition, distance, k, voting)

Knowledge Check

Question 1 of 4
Q1Single choice

In a KNN classifier, the predicted class for a query point is typically determined by: