A deterministic security detection engine (brute force, impossible travel, port scan, request flood) over synthetic logs, with local-LLM incident summaries and a real-time dashboard.
A SOC analyst reviewing raw logs by hand can't keep up with volume, and a black-box ML anomaly score doesn't explain why something was flagged.
Security teams need to spot brute-force attempts, impossible-travel logins, and port scans quickly, with a reason an analyst can actually verify — not just a score.
Seven deterministic, threshold-based detection rules, each producing an incident with a human-readable reason and the exact evidence that triggered it, plus an optional local-LLM layer that turns any incident into a plain-language summary without overstating what the evidence shows.
data/raw/*.csv (authentication, firewall, web_access, endpoint, ip_reputation)
-> pipeline/run_detection.py
brute_force, post_brute_force_success, unusual_login_time,
impossible_travel, known_bad_ip, request_flood, port_scan
-> data/processed/security_dataset.json (KPIs + full incident list)
+--> Next.js dashboard (Security Overview page)
+--> pipeline/ai_incident_summary.py -> llama3.1:8b
plain-language summary, hedged unless confirmed-compromise
Verified live: the AI summary correctly distinguishes a confirmed compromise from a merely-suspicious pattern.
[Critical] post_brute_force_success AI summary: "...a brute-force attack was successful, allowing unauthorized access to the account 'karim.khoury'. Treat this as a likely account compromise, rotate the credentials immediately..." [Medium] unusual_login_time AI summary: "...this pattern may warrant investigation as it could indicate unauthorized access. Verify with the account owner..."
Rule-based over ML. Every flagged incident has a traceable, explainable reason a SOC analyst could verify by hand — the same principle used in the Contract Intelligence and AI Assistant projects: keep deterministic logic separate from anything generated.
A heavy OSS base was abandoned mid-way, deliberately. The first candidate (a mature, ML-heavy log analyzer) was too tightly coupled to Postgres/Kafka/gRPC to safely strip down — rather than fight it, the base was swapped for a lighter, provably working foundation.
All log data is synthetic; suspicious-looking IPs are drawn from RFC 5737 reserved documentation ranges, never real infrastructure. AI summaries never claim an attack succeeded unless the incident type is an explicit confirmed-compromise type.