Cybersecurity

AI Security Log Analyzer

A deterministic security detection engine (brute force, impossible travel, port scan, request flood) over synthetic logs, with local-LLM incident summaries and a real-time dashboard.

Next.jsPythonOllama12/12 tests
View on GitHub

01Problem

A SOC analyst reviewing raw logs by hand can't keep up with volume, and a black-box ML anomaly score doesn't explain why something was flagged.

02Why it matters

Security teams need to spot brute-force attempts, impossible-travel logins, and port scans quickly, with a reason an analyst can actually verify — not just a score.

03Solution

Seven deterministic, threshold-based detection rules, each producing an incident with a human-readable reason and the exact evidence that triggered it, plus an optional local-LLM layer that turns any incident into a plain-language summary without overstating what the evidence shows.

04Architecture

data/raw/*.csv (authentication, firewall, web_access, endpoint, ip_reputation)
  -> pipeline/run_detection.py
       brute_force, post_brute_force_success, unusual_login_time,
       impossible_travel, known_bad_ip, request_flood, port_scan
  -> data/processed/security_dataset.json (KPIs + full incident list)
       +--> Next.js dashboard (Security Overview page)
       +--> pipeline/ai_incident_summary.py -> llama3.1:8b
             plain-language summary, hedged unless confirmed-compromise

05Demo

Verified live: the AI summary correctly distinguishes a confirmed compromise from a merely-suspicious pattern.

[Critical] post_brute_force_success
AI summary: "...a brute-force attack was successful, allowing unauthorized
access to the account 'karim.khoury'. Treat this as a likely account
compromise, rotate the credentials immediately..."

[Medium] unusual_login_time
AI summary: "...this pattern may warrant investigation as it could
indicate unauthorized access. Verify with the account owner..."

06Technical decisions

Rule-based over ML. Every flagged incident has a traceable, explainable reason a SOC analyst could verify by hand — the same principle used in the Contract Intelligence and AI Assistant projects: keep deterministic logic separate from anything generated.

A heavy OSS base was abandoned mid-way, deliberately. The first candidate (a mature, ML-heavy log analyzer) was too tightly coupled to Postgres/Kafka/gRPC to safely strip down — rather than fight it, the base was swapped for a lighter, provably working foundation.

07Security

All log data is synthetic; suspicious-looking IPs are drawn from RFC 5737 reserved documentation ranges, never real infrastructure. AI summaries never claim an attack succeeded unless the incident type is an explicit confirmed-compromise type.

08Limitations

A real bug, found and fixed: the first version of the synthetic log generator produced 315 spurious "impossible travel" incidents from randomly-assigned login countries. Fixed by giving synthetic users a consistent home country — down to exactly 10 incidents, matching the deliberately injected patterns with zero noise.

09Future improvements