AI / Document Intelligence

AI Contract Intelligence

Local, Ollama-powered contract extraction, risk flagging, and natural-language search over PDF contracts. No paid API keys, no data leaves the machine.

OllamaFastAPIStreamlitChromaDB
View on GitHub

01Problem

Reviewing a stack of contracts by hand to answer "which of these expire in 90 days" or "which have automatic renewal" is slow, and it's easy to miss a renewal deadline buried in page 12 of a PDF.

02Why it matters

Contract review errors are expensive — a missed auto-renewal clause or an unflagged high-value contract can cost real money. A tool that reports what it couldn't find, instead of guessing, is more useful than one that always sounds confident.

03Solution

A local LLM extracts a fixed schema of contract fields from each PDF (never inferring beyond what the text states), deterministic rules flag real risk patterns, and natural-language search answers questions across the whole set with source citations.

04Architecture

PDF upload -> pdfplumber (text extraction, no OCR/poppler dependency)
   +--> RAG path: chunk -> nomic-embed-text embeddings -> Chroma vector store
   |       -> MultiQueryRetriever + llama3.1:8b -> answer + source citations
   +--> Extraction path: full text -> structured-JSON prompt -> llama3.1:8b
             -> extracted_facts (explicit-text-only) + missing_fields
             -> flag_contract_risks()  (deterministic: auto-renewal, high value)
             -> generate_summary()      (plain-language, from extracted_facts ONLY)

05Demo

Tested on a deliberately incomplete ("sparse") synthetic contract, the extractor honestly reports what it couldn't find rather than inventing values:

{
  "contract_number": "NL-2024-1003",
  "contract_value": null,
  "start_date": null,
  "missing_fields": ["contract_value", "start_date", "end_date",
                      "renewal_terms", "payment_terms", "responsible_department"]
}
risk flags: ["6 field(s) could not be extracted... Human review required."]

06Technical decisions

Facts vs. summary, structurally separate. The extraction prompt only returns explicit-text facts; the summary is generated strictly from those already-extracted facts, never from the raw document — bounding what the summary can claim.

pdfplumber over OCR pipelines. Pure-Python text extraction avoids a poppler system dependency (and a real licensing wrinkle found in the upstream base project's own issue tracker) at the cost of not handling scanned image-only PDFs.

07Security

Every extraction response carries a fixed AI-safety notice: facts are extracted from explicit text, the summary is model-generated, verify before making business or legal decisions — not legal advice. 100% local, no data sent to a third party.

08Limitations

09Future improvements