← Back to Paul Ian Lim

How I build data pipelines

Most of the pipelines I build follow the same shape. Data comes in from a few places, gets cleaned up and modelled in one spot, then goes back out to the people and tools that use it.

Sources: Platform feeds, APIs, File loads

  1. Ingestion. Everything lands in one warehouse, so there is a single place to work from.
  2. Transformation. I clean the raw data, make it consistent and join it together so it is ready to analyse.
  3. Modelling. I build features from the clean data and run machine-learning models over it on a schedule.
  4. Semantic layer. The results and key numbers are organised into tidy datasets that are easy to query.
  5. Apps and reporting. Internal apps and reports put those results in front of the people who need them.
  6. Activation. The useful results go back out to the platforms that act on them.

Outputs: Internal apps, Reports, Downstream platforms

Across every step: Scheduled jobs run each step in the right order. When something breaks, the logs help me find and fix it.

Orchestration and monitoring cover every step, so they sit underneath the rest.