Local, Ollama-powered contract extraction, risk flagging, and natural-language search over PDF contracts. No paid API keys, no data leaves the machine.
Reviewing a stack of contracts by hand to answer "which of these expire in 90 days" or "which have automatic renewal" is slow, and it's easy to miss a renewal deadline buried in page 12 of a PDF.
Contract review errors are expensive — a missed auto-renewal clause or an unflagged high-value contract can cost real money. A tool that reports what it couldn't find, instead of guessing, is more useful than one that always sounds confident.
A local LLM extracts a fixed schema of contract fields from each PDF (never inferring beyond what the text states), deterministic rules flag real risk patterns, and natural-language search answers questions across the whole set with source citations.
PDF upload -> pdfplumber (text extraction, no OCR/poppler dependency)
+--> RAG path: chunk -> nomic-embed-text embeddings -> Chroma vector store
| -> MultiQueryRetriever + llama3.1:8b -> answer + source citations
+--> Extraction path: full text -> structured-JSON prompt -> llama3.1:8b
-> extracted_facts (explicit-text-only) + missing_fields
-> flag_contract_risks() (deterministic: auto-renewal, high value)
-> generate_summary() (plain-language, from extracted_facts ONLY)
Tested on a deliberately incomplete ("sparse") synthetic contract, the extractor honestly reports what it couldn't find rather than inventing values:
{
"contract_number": "NL-2024-1003",
"contract_value": null,
"start_date": null,
"missing_fields": ["contract_value", "start_date", "end_date",
"renewal_terms", "payment_terms", "responsible_department"]
}
risk flags: ["6 field(s) could not be extracted... Human review required."]
Facts vs. summary, structurally separate. The extraction prompt only returns explicit-text facts; the summary is generated strictly from those already-extracted facts, never from the raw document — bounding what the summary can claim.
pdfplumber over OCR pipelines. Pure-Python text extraction avoids a poppler system dependency (and a real licensing wrinkle found in the upstream base project's own issue tracker) at the cost of not handling scanned image-only PDFs.
Every extraction response carries a fixed AI-safety notice: facts are extracted from explicit text, the summary is model-generated, verify before making business or legal decisions — not legal advice. 100% local, no data sent to a third party.