Skip to content
View AwonAziz's full-sized avatar
🥨
funny
🥨
funny

Block or report AwonAziz

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
AwonAziz/README.md

Awon Aziz

AI/MLOps Engineer — I build the operational half of machine learning: evaluation, drift detection, and promotion gates.

→ awonaziz.github.io — the portfolio, with a written case study for each system below.

A model isn't finished when it trains well. It's finished when you can prove it still works, detect when it stops, and refuse to promote a replacement that only looks better. That's the part I work on.

Every repository below ships with tests, a CI run, and numbers I measured myself — including the ones that came out worse than expected.


Pinned work

Six evaluation questions that are painful in a notebook and trivial in SQL, over a DuckDB star schema fed by every project's artefacts. 141 hermetic tests, 5 versioned migrations, 17 charts.

Every row carries a data_origin label, so a generated figure can never be read as a measured one — filter to data_origin = 'real' and only observed numbers remain. The quality gate fails CI rather than warning in it, and its fifth check is the one that matters: analytical coverage. Fewer than two arms, fewer than two quantisations, or a missing slice family makes all six analyses quietly meaningless while every individual query still returns rows.

Rejected rows are counted with machine-readable error codes rather than tolerated, and editing an applied migration raises instead of letting the schema drift away from what CI built.

LoRA fine-tuning built on a hand-written NumPy autodiff, exported to ONNX, quantised to INT8, served behind FastAPI. Every layer proved against its reference implementation:

Checked against Result
Central differences, every op worst error 5.9e-09
torch.optim, same trajectories agree to 1e-09
nn.MultiheadAttention max diff 1.19e-07
Hugging Face peft, same weights and factors exactly 0.0
ONNX vs PyTorch predictions 1.0000 agreement
INT8 vs fp32 67.8 MB vs 268.6 MB, 98.4% agreement, −0.002 accuracy

LoRA reached 0.9217 against 0.9295 for full fine-tuning while training 1.1% of the weights (McNemar p=0.006). The finding I did not expect: LoRA is better calibrated than full fine-tuning (ECE 0.0180 vs 0.0306) — a low-rank update keeps weights near the pretrained solution, so if you threshold on a predicted probability the cheap method is the safer one to ship. 80 tests, one CI job, runs offline on CPU.

Embedding drift, output-quality drift and LLM-as-judge regression tracking for a production-shaped LLM application. Six drift signals — permutation-calibrated MMD, sliced Wasserstein, domain-classifier AUC, normalised Fréchet, k-NN novelty, concept gap — where effect size sets severity and significance only confirms.

Output quality covers accuracy, macro-F1, ECE/MCE/adaptive-ECE, Brier, AUC, reliability curves, abstention and per-intent damage, with separate in-scope and out-of-scope baselines. Quality is attributed to the window that served the request, not the window the label arrived in. The triage policy refuses to retrain on input drift alone and weights quality above input shift.

Platform: SQLite system of record with schema migration and upserts, 17 API endpoints, a 5-tab dashboard, Docker and Makefile. Six calibration bugs found and pinned by regression tests — an inflated domain-classifier null, a mismatched MMD permutation null, a novelty cutoff without leave-one-out, tautological accuracy, a missing judge veto cap, and quality baselines averaged over out-of-scope traffic. 184 tests, four CI jobs.

Sparse TF-IDF and dense retrieval fused with Reciprocal Rank Fusion over Chroma, feeding a two-agent pipeline that proposes root-cause hypotheses for human review — advisory only, never auto-executes.

Ships an 8-case golden harness scoring Hit@1/Hit@3, hypothesis correctness, confidence calibration, and appropriate uncertainty on a deliberate no-match case. The harness caught a regression in my own fusion method: the LSA fallback ranked an unrelated incident first on 2 of 8 cases that sparse retrieval alone solved. Diagnosed as insufficient co-occurrence data at small corpus size, fixed with a corpus-size trust gate that falls back to sparse-only below 50 documents. 38 tests, no API key required.

MLflow-tracked training and model registry, FastAPI serving (/predict, /health, /metrics, /drift/status), and a model-health dashboard. Champion/challenger promotion is gated on a measured F1 gain, and drift is watched with Population Stability Index, Kolmogorov–Smirnov and Jensen–Shannon divergence — mean PSI 1.66 across five features after injected drift. Champion at F1 0.873 / AUC 0.937 on the project's own synthetic data, under 5 ms p99.


Also here

  • ai-incident-response-system — multi-cloud Isolation Forest anomaly detection over AWS, Azure and GCP telemetry into a rule engine, triage engine, and notification router. Extended by agentic-incident-copilot, which adds a Chroma-backed postmortem knowledge base, hybrid TF-IDF + dense retrieval, two CrewAI agents, a human-review UI and Kubernetes manifests.
  • AI-Pair-Engineer — a four-stage review pipeline where each stage receives only the upstream findings relevant to its own job, cutting token cost and blocking context propagation. Every model response is schema-validated with retry-and-correct; submitted code is never executed or interpolated into a shell command.
  • cleanjobfunnel — has run unattended across 18 job boards reading Greenhouse, Lever, Ashby and SmartRecruiters directly, refreshed every ~20 minutes by a scheduled workflow and published to a live dashboard. 590 commits, 587 of them automated, 3 human. Built to run my own search, no scraping, no middleman board.
  • Infrastructure labs — 200+ hands-on commits across DevOps & CI/CD, Red Hat Linux and Cybersecurity.

How I work

I would rather ship a small system I can measure than a large one I can't. Every project here reports what broke and why — the LSA regression, the six calibration bugs, the three bugs inside from-scratch-to-served — because a repository that only shows its successes teaches the wrong lesson.

The three in from-scratch-to-served are the ones I would point at first. A cross-entropy implementation used logits - log(softmax(logits)), which collapses to a per-row constant: loss fell smoothly while accuracy stayed at chance, because argmax is scale-invariant. Gradient checking did not catch it — my analytic and numerical gradients agreed to 1e-10, because both were differentiating the same wrong function. Then signed INT8 quantisation took the model to 0.2610 accuracy, which is chance, until sweeping eight configurations found unsigned per-channel holding 0.9280. And inject_lora(targets=('q_proj','v_proj')) silently matched nothing on DistilBERT, which calls them q_lin and v_lin — the run completed, reported 92% validation accuracy, and was in fact a linear probe wearing a LoRA label.

Background

Diploma in Artificial Intelligence Operations, EduQual (UK), RQF Level 6 (bachelor's-level equivalent), Al Nafi International Colleges, 2026 — 90%.

Contact

Islamabad, Pakistan — open to remote (EU or US overlap) and relocation to UAE, Saudi Arabia, Qatar, UK or EU.

awonaziz786@gmail.com · linkedin.com/in/awonaziz · +92 335 5528211

Pinned Loading

  1. llm-drift-monitor llm-drift-monitor Public

    LLM drift & quality monitoring: embedding drift, output-quality drift, and LLM-as-judge regression tracking

    Python 1

  2. Cybersecurity-labs Cybersecurity-labs Public

    hands on showcase of what i've learned/ worked with across various domains

    Shell

  3. Devops-CICD-labs Devops-CICD-labs Public

    ill be showcasing concepts and tools that i have learned from my cloud labs practice at alnafi

    Shell

  4. ml-lifecycle-platform ml-lifecycle-platform Public

    A FastAPI-served ML model with MLflow experiment tracking, automated data drift detection (using Evidently AI), GitHub Actions CI/CD that auto-retriggers retraining when drift is detected, and a St…

    Python

  5. RedHat-Linux-Labs RedHat-Linux-Labs Public

    Here I'm showcasing most of the projects and tools that i have worked with hands on inside alnafi cloud labs, I have done these labs quite a long time ago but I'll showcase what I've learned

    Shell

  6. Hybrid-retrieval Hybrid-retrieval Public

    Hybrid sparse+dense retrieval with Reciprocal Rank Fusion over Chroma, feeding a two-agent RAG copilot for incident root-cause analysis. Ships an 8-case golden eval harness with Hit@k, hypothesis s…

    Python