A business intelligence dashboard backed by a real data pipeline: CSV, then validation, then cleaning, then dedup, then transform, then a KPI dashboard — not a mocked JSON fixture.
Manual reporting means numbers get pulled from five different systems and reconciled by hand — slow, error-prone, and with no visibility into duplicate or malformed source data before it reaches a dashboard.
A dashboard is only as trustworthy as the pipeline feeding it. Centralizing information with a validated pipeline improves decision-making by making sure the numbers on screen are the same numbers that were actually checked.
A Python pipeline takes raw CSV exports through validation, cleaning, duplicate detection, and transformation into a single JSON read model. A Next.js dashboard reads that model to render an Executive Overview and a deeper Analytics view.
data/raw/*.csv (departments, employees, revenue, expenses,
projects, suppliers, procurement, contracts)
-> pipeline/run_pipeline.py
- validation (required fields, date format)
- cleaning (malformed fields nulled + logged)
- dedup (duplicate keys removed + counted)
- transform (joins, aggregation, KPI computation)
-> data/processed/bi_dataset.json
-> Next.js server components read it directly
-> Executive Overview + Analytics dashboards
The synthetic data generator deliberately injects 3 real data-quality issues (a duplicate row, a blank required field, a malformed date). Running the pipeline reports exactly what it found and fixed:
validation errors found: 2 - employees: missing name for E0008 - employees: malformed hire_date 'not-a-date' for E0013, cleaned to null fields cleaned: 1 duplicate rows removed: 1 employees after cleaning: 139
JSON over a real database. A fixed, synthetic demo dataset doesn't need a hosted Postgres instance — the pipeline scripts are written so swapping the JSON write for a database write is a small, isolated change if this became a live tool.
Custom theme, not the default look. Built on a production-grade open-source dashboard starter, but shipped with a custom "Meridian" enterprise color system (deep slate-indigo + amber) rather than the starter's default palette.
No secrets or API keys required to run — auth runs in a free keyless mode. No real personal or financial data anywhere; every dataset is generated by the included pipeline script.