A machine-learning classification pipeline that categorizes biometric records into risk categories using a Random Forest classifier (Python, Pandas, NumPy, Scikit-learn).
Please read first: the original dataset (200K+ biometric records) is not included in this repository. The code is dataset-agnostic: you supply a CSV and the name of its target column. A small synthetic dataset is bundled for demonstration only — results from it are not project results.
The repository contains a modular, reproducible pipeline: load a CSV, clean it, impute missing values, split into train/test sets, train a Random Forest, evaluate it, save the model with joblib, and run predictions on new records from the command line.
Categorize biometric records into risk categories with a supervised classifier, and measure how reliably it does so using accuracy, precision, recall, F1-score, a confusion matrix and ROC-AUC.
- Build a clean, end-to-end preprocessing → training → evaluation → inference pipeline.
- Keep the code modular and runnable from the command line.
- Report evaluation results honestly: documented results and reproduced results are shown separately.
Python · Pandas · NumPy · Scikit-learn (Random Forest) · joblib · Matplotlib (notebook plots) · pytest
Raw CSV
→ Data loading (src/data_preprocessing.py)
→ Cleaning (drop rows with missing target; optional duplicate removal)
→ Feature/target split
→ Stratified train/test split (fixed seed)
→ [Pipeline] Missing-value imputation + categorical encoding (src/feature_engineering.py)
→ [Pipeline] Random Forest training (src/train.py)
→ Saved model (joblib) + held-out test set (src/model_io.py)
→ Prediction / Evaluation (src/predict.py, src/evaluate.py)
→ Report (reports/evaluation_results.md)
Imputation and encoding live inside one scikit-learn Pipeline, so the exact same preparation is applied at training time and at prediction time.
Known information (provided by the author):
- Size: 200K+ biometric records.
Not documented in this repository (fill in before publishing, if you wish):
| Item | Value |
|---|---|
| Dataset source | TODO |
| Feature names / descriptions | TODO |
| Target column name | TODO |
| Class labels and meaning | TODO |
Place your CSV in data/raw/ (git-ignored). See data/README.md.
Demo data: data/sample/sample_data.csv is synthetic (generated by src/generate_sample_data.py with scikit-learn's make_classification). Its column names (sample_feature_1…, sample_category) and target (risk_category) are generic placeholders with no biometric meaning. It exists only to show that the pipeline runs.
Supervised classification with a Random Forest. The held-out test set (20% by default, stratified) is used only for evaluation.
Column types are inferred from the DataFrame dtypes, so no column names are hard-coded:
| Step | Applies to | Method |
|---|---|---|
| Drop rows with missing target | all | clean_data |
| Drop exact duplicates | all | optional, off by default (--drop-duplicates) |
| Missing-value imputation | numeric columns | median |
| Missing-value imputation | non-numeric columns | most frequent |
| Encoding | non-numeric columns | ordinal encoding (unseen categories → -1) |
| Scaling | — | not applied (not needed for tree-based models) |
These are default choices, not findings about your data — review them against your dataset.
RandomForestClassifier from scikit-learn:
| Parameter | Value |
|---|---|
n_estimators |
100 (override with --n-estimators) |
max_depth |
unlimited (override with --max-depth) |
n_jobs |
-1 (use all CPU cores) |
random_state |
42 (override with --random-state) |
| everything else | scikit-learn defaults |
No hyperparameter tuning or cross-validation is implemented.
python -m src.evaluate computes accuracy, precision, recall, F1-score, a confusion matrix and ROC-AUC on the held-out test set. For binary targets, precision/recall/F1 use a positive class (default: last class in sorted order; set with --positive-label when training). For multiclass targets they are macro-averaged.
It writes reports/evaluation_results.md (and a .json copy), which separates:
- Documented project result — author-supplied, copied from project documentation.
- Current reproducible result — computed by the run, plus an automatic check of whether the documented numbers were reproduced.
| Metric | Documented project result | Current reproducible result |
|---|---|---|
| Accuracy | 99.31% | Run on the original dataset to fill in |
| Recall | 99.79% | Run on the original dataset to fill in |
| ROC-AUC | 0.9997 | Run on the original dataset to fill in |
The documented numbers are not verified by this repository; they can only be confirmed by running the pipeline on the original dataset. Update the right-hand column only with values printed by python -m src.evaluate.
Distributed-Biometric-Risk-Categorization/
├── data/
│ ├── raw/ # your CSV goes here (git-ignored)
│ ├── processed/ # held-out test set written by training (git-ignored)
│ ├── sample/sample_data.csv # synthetic demo data
│ └── README.md
├── notebooks/biometric_risk_analysis.ipynb
├── src/
│ ├── config.py # shared defaults and paths
│ ├── data_preprocessing.py
│ ├── feature_engineering.py
│ ├── model_io.py # joblib save/load
│ ├── train.py
│ ├── evaluate.py
│ ├── predict.py
│ └── generate_sample_data.py
├── models/README.md # trained models are git-ignored
├── reports/evaluation_results.md
├── tests/
│ ├── test_preprocessing.py
│ └── test_prediction.py
├── main.py # runs train → evaluate → example prediction
├── pytest.ini
├── requirements.txt
├── requirements-dev.txt
└── .gitignore
git clone https://github.com/Prajwalgrathish/Distributed-Biometric-Risk-Categorization.git
cd Distributed-Biometric-Risk-Categorization
python -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txt
pip install -r requirements-dev.txt # only needed for tests and the notebookQuick demo (synthetic data, no setup):
python main.py --demoDemo outputs are saved with a demo_ prefix so they never overwrite real results.
Full pipeline on your own data:
python main.py --data data/raw/your_data.csv --target YOUR_TARGET_COLUMNStep by step:
python -m src.train --data data/raw/your_data.csv --target YOUR_TARGET_COLUMN
python -m src.evaluate
python -m src.predict --input new_records.csv --output predictions.csvUseful training options: --test-size, --random-state, --n-estimators, --max-depth, --positive-label, --drop-duplicates. Run python -m src.train --help for all.
After python main.py --demo, predict a single synthetic-demo record:
python -m src.predict --model-path models/demo_random_forest.joblib \
--json '{"sample_feature_1":0.5,"sample_feature_2":-0.2,"sample_feature_3":1.1,"sample_feature_4":0.0,"sample_feature_5":0.3,"sample_feature_6":-0.7,"sample_category":"B"}'The output has a prediction column and one probability_<class> column per class. Inputs must contain the same feature columns the model was trained on; extra columns (such as the target) are ignored.
pip install -r requirements-dev.txt
pytestThe tests cover data loading and cleaning, splitting, imputation/encoding (including unseen categories), model saving/loading, prediction (shape, probabilities, missing/extra columns), evaluation metrics, and a small end-to-end training run.
- The original dataset is not included; results here cannot be confirmed without it.
- Single train/test split; no cross-validation and no hyperparameter tuning.
- Class imbalance is not handled (no class weights or resampling).
- Very high scores (such as those documented) should be checked for target leakage and duplicate records across train/test before being trusted; this repo does not automate that check.
- Ordinal encoding implies an arbitrary order for categorical values; fine for tree models, but not a general-purpose choice.
- "Distributed" in the project name: this implementation trains with scikit-learn on a single machine (
n_jobs=-1parallelizes across CPU cores). It does not implement distributed computing. - Biometric data is sensitive; keep raw data and trained models out of version control.
- Cross-validation and hyperparameter search (e.g.
GridSearchCV). - Class-imbalance handling and threshold analysis.
- Feature-importance analysis and leakage checks.
- Compare against other classifiers as baselines.
Athish Prajwal G R GitHub: https://github.com/Prajwalgrathish