Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

3 Commits

Folders and files

Repository files navigation

Distributed Biometric Risk Categorization

A machine-learning classification pipeline that categorizes biometric records into risk categories using a Random Forest classifier (Python, Pandas, NumPy, Scikit-learn).

Please read first: the original dataset (200K+ biometric records) is not included in this repository. The code is dataset-agnostic: you supply a CSV and the name of its target column. A small synthetic dataset is bundled for demonstration only — results from it are not project results.

Overview

The repository contains a modular, reproducible pipeline: load a CSV, clean it, impute missing values, split into train/test sets, train a Random Forest, evaluate it, save the model with joblib, and run predictions on new records from the command line.

Problem Statement

Categorize biometric records into risk categories with a supervised classifier, and measure how reliably it does so using accuracy, precision, recall, F1-score, a confusion matrix and ROC-AUC.

Objectives

  • Build a clean, end-to-end preprocessing → training → evaluation → inference pipeline.
  • Keep the code modular and runnable from the command line.
  • Report evaluation results honestly: documented results and reproduced results are shown separately.

Tech Stack

Python · Pandas · NumPy · Scikit-learn (Random Forest) · joblib · Matplotlib (notebook plots) · pytest

Architecture / Workflow

Raw CSV
  → Data loading            (src/data_preprocessing.py)
  → Cleaning                (drop rows with missing target; optional duplicate removal)
  → Feature/target split
  → Stratified train/test split (fixed seed)
  → [Pipeline] Missing-value imputation + categorical encoding   (src/feature_engineering.py)
  → [Pipeline] Random Forest training                            (src/train.py)
  → Saved model (joblib) + held-out test set                     (src/model_io.py)
  → Prediction / Evaluation                                      (src/predict.py, src/evaluate.py)
  → Report                                                       (reports/evaluation_results.md)

Imputation and encoding live inside one scikit-learn Pipeline, so the exact same preparation is applied at training time and at prediction time.

Dataset

Known information (provided by the author):

  • Size: 200K+ biometric records.

Not documented in this repository (fill in before publishing, if you wish):

Item Value
Dataset source TODO
Feature names / descriptions TODO
Target column name TODO
Class labels and meaning TODO

Place your CSV in data/raw/ (git-ignored). See data/README.md.

Demo data: data/sample/sample_data.csv is synthetic (generated by src/generate_sample_data.py with scikit-learn's make_classification). Its column names (sample_feature_1…, sample_category) and target (risk_category) are generic placeholders with no biometric meaning. It exists only to show that the pipeline runs.

Machine Learning Approach

Supervised classification with a Random Forest. The held-out test set (20% by default, stratified) is used only for evaluation.

Preprocessing

Column types are inferred from the DataFrame dtypes, so no column names are hard-coded:

Step Applies to Method
Drop rows with missing target all clean_data
Drop exact duplicates all optional, off by default (--drop-duplicates)
Missing-value imputation numeric columns median
Missing-value imputation non-numeric columns most frequent
Encoding non-numeric columns ordinal encoding (unseen categories → -1)
Scaling — not applied (not needed for tree-based models)

These are default choices, not findings about your data — review them against your dataset.

Model

RandomForestClassifier from scikit-learn:

Parameter Value
n_estimators 100 (override with --n-estimators)
max_depth unlimited (override with --max-depth)
n_jobs -1 (use all CPU cores)
random_state 42 (override with --random-state)
everything else scikit-learn defaults

No hyperparameter tuning or cross-validation is implemented.

Evaluation

python -m src.evaluate computes accuracy, precision, recall, F1-score, a confusion matrix and ROC-AUC on the held-out test set. For binary targets, precision/recall/F1 use a positive class (default: last class in sorted order; set with --positive-label when training). For multiclass targets they are macro-averaged.

It writes reports/evaluation_results.md (and a .json copy), which separates:

  • Documented project result — author-supplied, copied from project documentation.
  • Current reproducible result — computed by the run, plus an automatic check of whether the documented numbers were reproduced.

Results

Metric Documented project result Current reproducible result
Accuracy 99.31% Run on the original dataset to fill in
Recall 99.79% Run on the original dataset to fill in
ROC-AUC 0.9997 Run on the original dataset to fill in

The documented numbers are not verified by this repository; they can only be confirmed by running the pipeline on the original dataset. Update the right-hand column only with values printed by python -m src.evaluate.

Project Structure

Distributed-Biometric-Risk-Categorization/
├── data/
│   ├── raw/                 # your CSV goes here (git-ignored)
│   ├── processed/           # held-out test set written by training (git-ignored)
│   ├── sample/sample_data.csv   # synthetic demo data
│   └── README.md
├── notebooks/biometric_risk_analysis.ipynb
├── src/
│   ├── config.py            # shared defaults and paths
│   ├── data_preprocessing.py
│   ├── feature_engineering.py
│   ├── model_io.py          # joblib save/load
│   ├── train.py
│   ├── evaluate.py
│   ├── predict.py
│   └── generate_sample_data.py
├── models/README.md         # trained models are git-ignored
├── reports/evaluation_results.md
├── tests/
│   ├── test_preprocessing.py
│   └── test_prediction.py
├── main.py                  # runs train → evaluate → example prediction
├── pytest.ini
├── requirements.txt
├── requirements-dev.txt
└── .gitignore

Installation

git clone https://github.com/Prajwalgrathish/Distributed-Biometric-Risk-Categorization.git
cd Distributed-Biometric-Risk-Categorization
python -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt
pip install -r requirements-dev.txt   # only needed for tests and the notebook

Usage

Quick demo (synthetic data, no setup):

python main.py --demo

Demo outputs are saved with a demo_ prefix so they never overwrite real results.

Full pipeline on your own data:

python main.py --data data/raw/your_data.csv --target YOUR_TARGET_COLUMN

Step by step:

python -m src.train    --data data/raw/your_data.csv --target YOUR_TARGET_COLUMN
python -m src.evaluate
python -m src.predict  --input new_records.csv --output predictions.csv

Useful training options: --test-size, --random-state, --n-estimators, --max-depth, --positive-label, --drop-duplicates. Run python -m src.train --help for all.

Example Prediction

After python main.py --demo, predict a single synthetic-demo record:

python -m src.predict --model-path models/demo_random_forest.joblib \
  --json '{"sample_feature_1":0.5,"sample_feature_2":-0.2,"sample_feature_3":1.1,"sample_feature_4":0.0,"sample_feature_5":0.3,"sample_feature_6":-0.7,"sample_category":"B"}'

The output has a prediction column and one probability_<class> column per class. Inputs must contain the same feature columns the model was trained on; extra columns (such as the target) are ignored.

Testing

pip install -r requirements-dev.txt
pytest

The tests cover data loading and cleaning, splitting, imputation/encoding (including unseen categories), model saving/loading, prediction (shape, probabilities, missing/extra columns), evaluation metrics, and a small end-to-end training run.

Limitations

  • The original dataset is not included; results here cannot be confirmed without it.
  • Single train/test split; no cross-validation and no hyperparameter tuning.
  • Class imbalance is not handled (no class weights or resampling).
  • Very high scores (such as those documented) should be checked for target leakage and duplicate records across train/test before being trusted; this repo does not automate that check.
  • Ordinal encoding implies an arbitrary order for categorical values; fine for tree models, but not a general-purpose choice.
  • "Distributed" in the project name: this implementation trains with scikit-learn on a single machine (n_jobs=-1 parallelizes across CPU cores). It does not implement distributed computing.
  • Biometric data is sensitive; keep raw data and trained models out of version control.

Future Improvements

  • Cross-validation and hyperparameter search (e.g. GridSearchCV).
  • Class-imbalance handling and threshold analysis.
  • Feature-importance analysis and leakage checks.
  • Compare against other classifiers as baselines.

Author

Athish Prajwal G R GitHub: https://github.com/Prajwalgrathish

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages