I build data pipelines that have to be correct, not just running CDC and streaming into OLAP warehouses, with idempotency and replay worked out rather than assumed.
Currently a Data Engineer Intern at RIPT-PTIT (Hanoi), building a real-time CDC pipeline streaming MySQL/MongoDB into ClickHouse and StarRocks. BSc in Information Systems at PTIT, expected 2028.
Selected work
| Project | What it demonstrates |
|---|---|
| Real-Time Clickstream Analytics Pipeline | Exactly-once results from at-least-once delivery Kafka → Flink → ClickHouse. Verified by killing Flink mid-stream, not by assertion. |
| GDELT Medallion Pipeline | Batch lakehouse end to end Polars / DuckDB / Postgres / MinIO. Manifest state machine where replaying a failed batch is a no-op; 160+ tests at 90%+ coverage. |
Working with
Kafka Polars DuckDB ClickHouse StarRocks Parquet
MinIO / S3 PostgreSQL MySQL MongoDB Python SQL Docker
Currently going deeper on: data warehouse modelling, pipeline correctness under failure, and the trade-offs between table formats.
Hanoi, Vietnam · English (TOEIC 945 LR) · Open to Data Engineer internships and junior roles

