Data engineer, 7+ years in. Most of my work is Python talking to PostgreSQL and TimescaleDB: pipelines that move tens of millions of rows a day, with statistical models sitting on top.
Day job: I lead data science in the public sphere, mostly election estimation with multilevel regression and poststratification (MRP) over large survey and unstructured opinion datasets. Everything else on this profile is side projects I build to learn: quant research, and spatial analytics over Spanish open data (Catastro, INE, MITMA).
I've also started contributing upstream: merged changes in taskiq and numba, a docs PR in pandas under review, and bug reports in pandas, pydantic and langgraph. The contribution graph below has the rest.
| Repo | Description |
|---|---|
| quant-market-data-forensics | Point-in-time market data platform: bitemporal ingestion (Sharadar, FRED vintages), TimescaleDB, SQL feature factory, Numba kernels. Its data-quality instruments found 3 biases in the feed and killed the strategy built on it. |
| sqlpush | prisma db push for SQLAlchemy: push, diff and check PostgreSQL/TimescaleDB schemas straight from your models. Risk-classified changes, CI exit codes, no migration files. |
| ine-api | Python client for the Tempus API of the INE, the Spanish national statistics office. |


