Data / Business Intelligence

Enterprise BI Platform

A business intelligence dashboard backed by a real data pipeline: CSV, then validation, then cleaning, then dedup, then transform, then a KPI dashboard — not a mocked JSON fixture.

Next.jsshadcn/uiPythonRecharts
View on GitHub

01Problem

Manual reporting means numbers get pulled from five different systems and reconciled by hand — slow, error-prone, and with no visibility into duplicate or malformed source data before it reaches a dashboard.

02Why it matters

A dashboard is only as trustworthy as the pipeline feeding it. Centralizing information with a validated pipeline improves decision-making by making sure the numbers on screen are the same numbers that were actually checked.

03Solution

A Python pipeline takes raw CSV exports through validation, cleaning, duplicate detection, and transformation into a single JSON read model. A Next.js dashboard reads that model to render an Executive Overview and a deeper Analytics view.

04Architecture

data/raw/*.csv (departments, employees, revenue, expenses,
                 projects, suppliers, procurement, contracts)
  -> pipeline/run_pipeline.py
       - validation   (required fields, date format)
       - cleaning     (malformed fields nulled + logged)
       - dedup        (duplicate keys removed + counted)
       - transform    (joins, aggregation, KPI computation)
  -> data/processed/bi_dataset.json
  -> Next.js server components read it directly
  -> Executive Overview + Analytics dashboards

05Demo

The synthetic data generator deliberately injects 3 real data-quality issues (a duplicate row, a blank required field, a malformed date). Running the pipeline reports exactly what it found and fixed:

validation errors found: 2
  - employees: missing name for E0008
  - employees: malformed hire_date 'not-a-date' for E0013, cleaned to null
fields cleaned: 1
duplicate rows removed: 1
employees after cleaning: 139

06Technical decisions

JSON over a real database. A fixed, synthetic demo dataset doesn't need a hosted Postgres instance — the pipeline scripts are written so swapping the JSON write for a database write is a small, isolated change if this became a live tool.

Custom theme, not the default look. Built on a production-grade open-source dashboard starter, but shipped with a custom "Meridian" enterprise color system (deep slate-indigo + amber) rather than the starter's default palette.

07Security

No secrets or API keys required to run — auth runs in a free keyless mode. No real personal or financial data anywhere; every dataset is generated by the included pipeline script.

08Limitations

09Future improvements