Designing a K-Nearest Neighbour (KNN) Exam-Failure Classifier
To predict whether a student will pass or fail the exam using K-Nearest Neighbour (K-Nearest Neighbour), we treat each student as a feature vector and the exam outcome as a class label. KNN is non-parametric: for a new student, it finds the closest training students using a distance metric and assigns the class most common among those neighbours (optionally using distance-weighting). This overall decision rule is consistent with KNN classifier descriptions in common ML references and KNN implementations. 2
From the table, we build inputs from:
- Language (Java/C++/Python)
- Passed all Assignments (Yes/No)
- GPA
The target is:
- Passed Exam (Yes/No)
KNN design depends heavily on (i) how we encode categorical variables, (ii) feature scaling, and (iii) selecting and the distance metric. Since KNN uses distances in feature space, features on different scales can dominate the distance calculations; thus scaling is a key preprocessing step.
Footnotes
-
K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩
-
KNeighborsClassifier — scikit-learn predict/parameter behavior - Covers
weights(uniform vs distance) and distance metric details including Minkowski. ↩ -
What is Feature Scaling and Why is it Important? - Analytics Vidhya - Discusses scaling importance for KNN and distance-based algorithms. ↩
KNN design workflow for this exam dataset
Define features and label
1. Data & targetUse Language, Passed all Assignments, GPA → predict Passed Exam."
Encode & scale
2. PreprocessingEncode Language; keep assignments as binary; scale GPA."
Distance metric + vote rule
3. Distance & votingChoose metric (e.g., Euclidean via Minkowski) and weights (uniform vs distance)."
Select k
4. HyperparametersEvaluate candidate k values using cross-validation."
Fit + inference
5. Train and predictStore training vectors; for each query, compute nearest neighbours and vote."
Data preparation and feature engineering
1) Define the supervised learning problem
Let each student be represented by a feature vector and label :
- if Passed Exam = Yes, otherwise .
KNN stores the training pairs ; during prediction it uses the labels of the nearest stored vectors.
2) Encode categorical features (Language)
KNN requires numeric vectors. Encode Language using one-hot encoding into three binary features:
- , , (only one of them is 1 for each student).
Keep Passed all Assignments as a binary indicator:
Now each student becomes: so here .
Pro Tip: With KNN, categorical encoding choices can change the geometry of the feature space. One-hot encoding is standard for Euclidean-style distances, because categories become orthogonal indicator directions.
3) Scale numeric features (especially GPA)
Because KNN relies on distances, features with larger numeric ranges can disproportionately influence the result; scaling reduces this effect and improves distance fairness. A typical approach is standardization or normalization applied to numeric features like GPA before training and prediction.
In this dataset, GPA is numeric while the other features are 0/1. Even so, scaling GPA is still recommended to ensure stable behaviour and prevent it from overwhelming or being overwhelmed by binary dimensions.
Footnotes
-
K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩
-
What is Feature Scaling and Why is it Important? - Analytics Vidhya - Discusses scaling importance for KNN and distance-based algorithms. ↩ ↩2
KNN prediction algorithm (Pass/Fail) for a new student
- 1Step 1
Encode Language with the same one-hot scheme as training and convert Passed all Assignments to 0/1; scale GPA using the training-fitted scaler.
- 2Step 2
For each training vector , compute using the chosen metric (e.g., Euclidean is Minkowski with ).
- 3Step 3
Sort (or partial-select) by distance and keep the closest training students.
- 4Step 4
If using uniform weights, predict the majority label among the neighbours. If using distance weights, assign larger influence to closer points (e.g., inverse-distance weighting).
- 5Step 5
Return if predicted class is 1, else .
Distance metric and voting design
Distance metric
A common default in KNN implementations is Minkowski distance; scikit-style KNN uses Minkowski and typically supports Euclidean by setting . 2
For two students and :
- Choose for Euclidean distance: Other norms (e.g., Manhattan, ) can be explored, but the Euclidean choice is a strong baseline.
Voting rule and the role of the weights parameter
KNN classification is often implemented as a “vote” over the neighbour labels. The weights parameter typically supports:
uniform: all neighbours vote equallydistance: closer neighbours vote more (often using inverse distance)
Given this dataset is small, distance weighting can reduce sensitivity to noisy neighbours (e.g., a single close but contradictory example).
Warning: If you set too small, KNN becomes sensitive to noise; if is too large, it may smooth over local patterns. This bias–variance trade-off is why must be tuned using validation.
Footnotes
-
K-Nearest Neighbor (KNN) Algorithm - CODERCOPS - Explains KNN classification rule: choose k closest and use majority label for classification. ↩
-
KNeighborsClassifier — scikit-learn predict/parameter behavior - Covers
weights(uniform vs distance) and distance metric details including Minkowski. ↩ ↩2 -
A Gentle Introduction to k-fold Cross-Validation - MachineLearningMastery - Explains cross-validation for hyperparameter selection and general evaluation stability. ↩
Selecting hyperparameter (and evaluating the design)
Cross-validation for choosing
Since the dataset in the prompt is small, performance estimates from a single train/test split may vary a lot. Use k-fold cross-validation to select (and optionally the distance metric or weights). Cross-validation is widely used to choose hyperparameters by averaging scores over multiple splits. A typical setup is:
- Candidate values like
- Use stratification so each fold keeps similar proportions of Pass/Fail labels (important for classification).
Tie-handling
For classification, if votes tie, some implementations may depend on neighbour ordering; it’s another reason to prefer odd for binary classification.
Footnotes
-
A Gentle Introduction to k-fold Cross-Validation - MachineLearningMastery - Explains cross-validation for hyperparameter selection and general evaluation stability. ↩ ↩2
-
K-Nearest Neighbours - GeeksforGeeks - Notes tie/odd-k considerations for classification and cross-validation for choosing k. ↩
Illustrative feature composition for the KNN input vector
Binary-encoded categories + scaled numeric GPA
Common edge cases & design choices
KNN Algorithm (intuition, distance, k, voting)
Knowledge Check
In a KNN classifier, the predicted class for a query point is typically determined by: