I'm Pavan. I work the full data stack, and I like building things that actually run — ingestion pipelines that don't break at 2am, streaming systems that keep up, dashboards people actually open.
These days I'm into batch + streaming pipelines (Spark, Kafka), lakehouses (Delta Lake, Databricks), analytics engineering (dbt, SQL), and BI (Power BI, DAX) — mostly on AWS.
| Project | What it does |
|---|---|
| Real-Time Retail Intelligence | Kafka → Spark Structured Streaming → Delta Lake platform with ML forecasting and real-time fraud alerts, dbt-tested |
| Retail Sales Analytics | Batch lakehouse (bronze/silver/gold) processing 56K+ orders ($80.96M), 10/10 data-quality checks, Power BI dashboard |
| Cloud Analytics Knowledge Search | Hybrid Retrieval and Grounded QA over 300 synthetic enterprise documents — TF-IDF + dense (MiniLM) + hybrid retrieval, measured recall/MRR, extractive answers with citations, Streamlit UI |
| Portfolio | Personal portfolio website — Data Engineer | Data Analyst |
- AWS Certified Data Engineer – Associate
- Microsoft Power BI Data Analyst Associate (PL-300)
- Databricks Certified Data Analyst Associate
- Databricks Fundamentals Accreditation
- Microsoft Azure AI Engineer Associate (AI-102)
