I am a Data Engineer with a background in IT infrastructure, cloud engineering, and network engineering. I specialize in building scalable cloud-based data platforms using Azure Data Factory, Databricks, Python, and SQL. My focus is on designing reliable data pipelines that transform raw data into trusted business insights.
I came into data engineering through IT infrastructure, networking, and cloud engineering — which means I think about data pipelines the way I think about networks: what happens when a link drops, where the bottleneck actually is, and who gets paged.
Today, I build cloud data platforms on Azure. Data Factory for orchestration, Databricks and PySpark for transformation, Data Lake for zoned storage, SQL for the modelling layer. The goal is always the same: turn raw, inconsistent source data into datasets people are willing to make decisions on.
The infrastructure background still does work. I have built and run three-tier applications on AWS, designed VPC networking with private endpoints, and handled snapshot and AMI disaster recovery. Most data engineers can write the transformation; fewer can tell you why the subnet routing broke it.
I am also working on where AI genuinely helps a data workflow — quality monitoring, documentation, anomaly explanation — with a firm line: models explain and draft; tests decide.