feat: invoice-to-csv 0.4.0 (code-only extractor, validation, grounding, optional second reading) #6

Open
jonah wants to merge 3 commits from feat/invoice-to-csv into main
Owner

What

invoice-to-csv, first slice: PDF invoices in, validated CSV out. No model, no network, no download by default; about 1 ms per invoice.

  • Code-only extractor (rules/): label vocabularies in nl/en/de, a line-table reader keyed on its header words, locale-aware amounts and dates (decimal comma vs point detected per document). A layout it does not recognise, a missing value or an ambiguity goes to needs_review/ with the reason. Nothing is guessed.
  • Four layers keep wrong values out of the CSV: arithmetic (incl. quantity x unit price), identifier checksums, grounding (every extracted value must be printed on the page), and customer-clutter handling (customer-labelled lines ignored, competing candidates mean unreadable).
  • Optional second reading (--verify-model, extra [model]): a small GLiNER2.5 model with its vocabulary trimmed to nl/de/en (243 MB) reads supplier, invoice number and dates independently; a disagreement sends the invoice to review. tools/build_model.py builds it.
  • Synthetic invoice generator (4 print styles, injected defects, --distractors), committed development / held-out / clutter sets, an eval with a silent-error metric, CI floor tests.
  • Replaces the first design (an 8B LLM behind a model server: 60 min for 50 invoices, amounts right ~47%). The LLM client and config are removed.

Measured, and what it does not show

  • Every clean invoice accepted, every injected defect sent to review, 0 silent errors on the 50 development, 20 held-out-style and 40 clutter invoices, and on 800 invoices from 10 unseen seeds.
  • This is not a real-world failure rate. All layouts were written by the same author as the rules. The clutter set found a real silent error (a customer's valid VAT number replacing a corrupted supplier one); it was fixed, so that set is no longer unseen.
  • Second reading, with planted errors that pass arithmetic and grounding: costs 14% false reviews (12/88 clean), catches 32/32 order-number-as-invoice-number errors but only 14/32 customer-as-supplier errors. It is a partial safety net; supplier names are the gap and the case for fine-tuning.
  • 129 noisy German OCR texts (StefanStefan/german-invoice): 0 accepted (all refused for layout). It refuses what it cannot read.
  • Dynamic int8 quantization of the model was tried and rejected (slower, less accurate).

Not done

  • A published review rate on real invoices (a real set must be measured locally; none is uploaded anywhere).
  • OCR, --watch, an accounting export format, a Docker image, routing through redact-proxy.

Checked

320 hermetic tests, ruff, import-linter (4 contracts), pre-commit. The real model was run manually on all sets.

🤖 Generated with Claude Code

## What `invoice-to-csv`, first slice: PDF invoices in, validated CSV out. No model, no network, no download by default; about 1 ms per invoice. - **Code-only extractor** (`rules/`): label vocabularies in nl/en/de, a line-table reader keyed on its header words, locale-aware amounts and dates (decimal comma vs point detected per document). A layout it does not recognise, a missing value or an ambiguity goes to `needs_review/` with the reason. Nothing is guessed. - **Four layers keep wrong values out of the CSV:** arithmetic (incl. quantity x unit price), identifier checksums, **grounding** (every extracted value must be printed on the page), and customer-clutter handling (customer-labelled lines ignored, competing candidates mean unreadable). - **Optional second reading** (`--verify-model`, extra `[model]`): a small GLiNER2.5 model with its vocabulary trimmed to nl/de/en (243 MB) reads supplier, invoice number and dates independently; a disagreement sends the invoice to review. `tools/build_model.py` builds it. - Synthetic invoice generator (4 print styles, injected defects, `--distractors`), committed development / held-out / clutter sets, an eval with a **silent-error** metric, CI floor tests. - Replaces the first design (an 8B LLM behind a model server: 60 min for 50 invoices, amounts right ~47%). The LLM client and config are removed. ## Measured, and what it does not show - Every clean invoice accepted, every injected defect sent to review, **0 silent errors** on the 50 development, 20 held-out-style and 40 clutter invoices, and on 800 invoices from 10 unseen seeds. - **This is not a real-world failure rate.** All layouts were written by the same author as the rules. The clutter set found a real silent error (a customer's valid VAT number replacing a corrupted supplier one); it was fixed, so that set is no longer unseen. - Second reading, with planted errors that pass arithmetic and grounding: costs 14% false reviews (12/88 clean), catches 32/32 order-number-as-invoice-number errors but only 14/32 customer-as-supplier errors. It is a partial safety net; supplier names are the gap and the case for fine-tuning. - 129 noisy German OCR texts (`StefanStefan/german-invoice`): 0 accepted (all refused for layout). It refuses what it cannot read. - Dynamic int8 quantization of the model was tried and rejected (slower, less accurate). ## Not done - A published review rate on real invoices (a real set must be measured locally; none is uploaded anywhere). - OCR, `--watch`, an accounting export format, a Docker image, routing through redact-proxy. ## Checked 320 hermetic tests, ruff, import-linter (4 contracts), pre-commit. The real model was run manually on all sets. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Schema, validation (including quantity x unit price), grounding (every extracted value must be
printed in the PDF), outcome/pipeline/output, CLI, an eval with a silent-error metric, and the
synthetic invoice generator with committed development and held-out sets.

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Replace the model-based extractor (60 min per 50 invoices, ~47% amount accuracy) with
plain code: label vocabularies, a header-keyed line-table reader, locale-aware numbers and
dates. A layout it does not recognise goes to review; nothing is guessed.
- rules/: numbers, header, table, RulesExtractor
- data/invoices/generate.py --distractors (customer block, order/delivery data) and a
  committed 40-invoice adversarial set
- fixes found by that set: a corrupted supplier VAT number was replaced by the customer's
  valid one (and the supplier by the customer's postcode line); ids are now chosen by shape,
  customer lines are ignored, ambiguity gives None
- remove extract.py, config.py and the JSON-schema helpers; CLI and eval run the rules
- README and CHANGELOG describe the real design

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
feat: optional second reading, build tool, verifier and CLI flags
All checks were successful
CI / checks (pull_request) Successful in 1m10s
CI / policy (pull_request) Successful in 2s
CI / python (pull_request) Successful in 40s
96ff8e5133
A small GLiNER2.5 model (vocabulary trimmed to nl/de/en, 243 MB) independently reads the
supplier, invoice number and dates; a disagreement sends the invoice to review.
- verify.py (pure comparison), gliner.py (optional reader), cli.py, --verify-model/--strict
- tools/build_model.py: reproduces the trimmed model from the public base model
- number and date parsing moved to a shared module
- measured: best candidate at confidence 0.8 costs 14% false reviews, catches 32/32 planted
  invoice-number errors and 14/32 planted supplier errors; dynamic int8 rejected (slower, worse)
- extra invoice-to-csv[model]; README, CHANGELOG describe the numbers and the limits

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Author
Owner

🤖 Auto-review status · jonah/safe-ai-tools PR #6

📋 Review posted — fixes required, awaiting human action.

Last updated: 2026-10-06T19:42:21.992Z

<!-- opencode-status --> **🤖 Auto-review status** · `jonah/safe-ai-tools` PR #6 📋 **Review posted** — fixes required, awaiting human action. *Last updated: 2026-10-06T19:42:21.992Z*
jonah left a comment

[STATUS: FIXES_REQUIRED]

Verified first

Ran locally (had to pip install pdfplumber fpdf2 + editable-install the package into .venv; uv is not on this box):

  • ruff check . + ruff format --check . → clean
  • lint-imports → 4/4 contracts kept (including the two new invoice_to_csv ones)
  • pytest -m "not integration" → 320 passed (217 new for this package), 1 skipped (the model integration test), 2 deselected
  • No noqa/type: ignore/bare except/TODO anywhere in the diff; CHANGELOG updated alongside the VERSION bump (satisfies the CI policy job)

Must fix

  • ingest.py:24 — corrupt-but-openable PDFs crash the entire run, losing all output. except (PdfminerException, OSError) misses pdfplumber's MalformedPDFException (a sibling class, not a subclass — confirmed via MRO) and internal TypeErrors. Fuzzed 70 corruptions of a real invoice PDF (truncations + byte flips): 7 escaped IngestError (MalformedPDFException, TypeError: 'NoneType' object is not iterable). pipeline.process (pipeline.py:31) only catches (IngestError, ExtractError), so process_all aborts and run() raises before write_results — zero CSVs from the batch, on untrusted input the tool explicitly claims to handle ("unreadable PDF → review"). Root-cause fix is one place: read_text is the single seam — wrap broadly there (except Exception), keeping the existing message that already carries type(exc).__name__, so the failure stays diagnosable via reason.txt instead of destroying the run.

  • output.py:84 — _file_for_review assumes the rejected path is a readable regular file; it isn't always. Reproduced two ways, both aborting run() after processing with no output written:

    • dangling symlink ghost.pdf in the input folder → FileNotFoundError (ingest correctly converts this to IngestError → Rejected, then the copy reintroduces the crash)
    • a directory named weird.pdf → IsADirectoryError (folder.glob("*.pdf") at __main__.py:35 matches directories)
      Lazy fix: copy only if rejected.path.is_file(), always write the .reason.txt. One line.

Should fix (low)

  • output.py:89-93 — _clear_stale_review only prunes review copies of now-accepted invoices. PDFs deleted from the input folder (or a different folder run into the same --out) linger in needs_review/ forever, so a human works a stale queue that contradicts the docstring's "rewrite … from outcomes". Rebuild the folder from the current rejected set.
  • eval/score.py:82 — empty gold.jsonl → ZeroDivisionError traceback (/ total with total == 0) instead of an actionable message (§1.XVII). to_markdown already guards with _pct; accuracy doesn't.

Info (deliberate, no action needed)

  • output.py:34-40 _cell prefixes ' to any value starting with -, so a genuine negative amount (credit note — "creditnota" is in _TITLES) lands in the CSV as text, not a number. Standard OWASP CSV-injection trade-off; just know it if anything downstream expects numerics.
  • Env caveat: the .venv here predated this branch (invoice_to_csv, pdfplumber, fpdf were missing); CI's uv sync --all-packages --dev covers it. The GLiNER path is untested beyond the stub test (integration test skips without INVOICE_MODEL_DIR).
[STATUS: FIXES_REQUIRED] ## Verified first Ran locally (had to `pip install pdfplumber fpdf2` + editable-install the package into `.venv`; `uv` is not on this box): - `ruff check .` + `ruff format --check .` → clean - `lint-imports` → 4/4 contracts kept (including the two new `invoice_to_csv` ones) - `pytest -m "not integration"` → **320 passed** (217 new for this package), 1 skipped (the model integration test), 2 deselected - No `noqa`/`type: ignore`/bare `except`/TODO anywhere in the diff; CHANGELOG updated alongside the `VERSION` bump (satisfies the CI policy job) ## Must fix - **`ingest.py:24` — corrupt-but-openable PDFs crash the entire run, losing all output.** `except (PdfminerException, OSError)` misses pdfplumber's `MalformedPDFException` (a sibling class, not a subclass — confirmed via MRO) and internal `TypeError`s. Fuzzed 70 corruptions of a real invoice PDF (truncations + byte flips): **7 escaped `IngestError`** (`MalformedPDFException`, `TypeError: 'NoneType' object is not iterable`). `pipeline.process` (`pipeline.py:31`) only catches `(IngestError, ExtractError)`, so `process_all` aborts and `run()` raises **before `write_results`** — zero CSVs from the batch, on untrusted input the tool explicitly claims to handle ("unreadable PDF → review"). Root-cause fix is one place: `read_text` is the single seam — wrap broadly there (`except Exception`), keeping the existing message that already carries `type(exc).__name__`, so the failure stays diagnosable via `reason.txt` instead of destroying the run. - **`output.py:84` — `_file_for_review` assumes the rejected path is a readable regular file; it isn't always.** Reproduced two ways, both aborting `run()` after processing with no output written: - dangling symlink `ghost.pdf` in the input folder → `FileNotFoundError` (ingest correctly converts this to `IngestError` → `Rejected`, then the copy reintroduces the crash) - a **directory** named `weird.pdf` → `IsADirectoryError` (`folder.glob("*.pdf")` at `__main__.py:35` matches directories) Lazy fix: copy only if `rejected.path.is_file()`, always write the `.reason.txt`. One line. ## Should fix (low) - **`output.py:89-93` — `_clear_stale_review` only prunes review copies of *now-accepted* invoices.** PDFs deleted from the input folder (or a different folder run into the same `--out`) linger in `needs_review/` forever, so a human works a stale queue that contradicts the docstring's "rewrite … from `outcomes`". Rebuild the folder from the current rejected set. - **`eval/score.py:82` — empty `gold.jsonl` → `ZeroDivisionError` traceback** (`/ total` with `total == 0`) instead of an actionable message (§1.XVII). `to_markdown` already guards with `_pct`; `accuracy` doesn't. ## Info (deliberate, no action needed) - `output.py:34-40` `_cell` prefixes `'` to any value starting with `-`, so a genuine negative amount (credit note — "creditnota" is in `_TITLES`) lands in the CSV as text, not a number. Standard OWASP CSV-injection trade-off; just know it if anything downstream expects numerics. - Env caveat: the `.venv` here predated this branch (`invoice_to_csv`, `pdfplumber`, `fpdf` were missing); CI's `uv sync --all-packages --dev` covers it. The GLiNER path is untested beyond the stub test (integration test skips without `INVOICE_MODEL_DIR`).
Author
Owner

Review findings addressed in 9e35410:

  • ingest: any parser failure becomes an unreadable-PDF review entry (MalformedPDFException is not a PdfminerException; TypeErrors too).
  • output: reason file written even if the path is a dangling symlink or a directory.
  • output: needs_review/ rebuilt from the current run (stale PDF+reason pairs removed, foreign files kept).
  • eval: empty gold.jsonl gives a usage error.
    Checked: 326 hermetic tests, ruff, import-linter; 70 truncated/byte-flipped PDFs plus a folder and a dangling symlink run to completion (1 file survived intact and was accepted, 71 in review). The negative-amount CSV trade-off is noted in the CHANGELOG.

🤖 Generated with Claude Code

Review findings addressed in 9e35410: - ingest: any parser failure becomes an unreadable-PDF review entry (MalformedPDFException is not a PdfminerException; TypeErrors too). - output: reason file written even if the path is a dangling symlink or a directory. - output: needs_review/ rebuilt from the current run (stale PDF+reason pairs removed, foreign files kept). - eval: empty gold.jsonl gives a usage error. Checked: 326 hermetic tests, ruff, import-linter; 70 truncated/byte-flipped PDFs plus a folder and a dangling symlink run to completion (1 file survived intact and was accepted, 71 in review). The negative-amount CSV trade-off is noted in the CHANGELOG. 🤖 Generated with [Claude Code](https://claude.com/claude-code)
Author
Owner

🤖 Auto-review status · jonah/safe-ai-tools PR #6

📋 Review posted — fixes required, awaiting human action.

Last updated: 2026-10-06T20:18:13.390Z

<!-- opencode-status --> **🤖 Auto-review status** · `jonah/safe-ai-tools` PR #6 📋 **Review posted** — fixes required, awaiting human action. *Last updated: 2026-10-06T20:18:13.390Z*
jonah left a comment

[STATUS: FIXES_REQUIRED]

Verified first

  • ruff check . + ruff format --check . → clean; lint-imports → 4/4 contracts kept; pytest -m "not integration" → 326 passed (6 new tests from the fix commit), 2 deselected.
  • No noqa/type: ignore/TODO in the diff; the only bare-ish catch is the deliberate, commented except Exception at ingest.py:23.
  • Previous review's must-fixes confirmed resolved: fuzzed 300 corrupt PDFs (truncations + byte flips) through pipeline.process → 0 escapes (was 7/70); dangling-symlink and directory x.pdf repros now produce only .reason.txt (no crash); needs_review/ rebuild and the empty-gold parser.error both land with tests. CHANGELOG documents each fix.

Must fix

  • pipeline.py:33 — validate/grounding/verifiers run outside the try; two proven crashes abort the whole batch before write_results, zero CSVs out. The ingest seam was hardened, but except (IngestError, ExtractError) at pipeline.py:31 only covers lines 29–30; _problems (line 33 → 19) runs unguarded. End-to-end repros on PDFs that extract successfully:
    • formats.py:32 — number_readings builds Decimal(token…) for tokens mixing , and . with ≥2 of the decimal kind: numbers_in raises decimal.InvalidOperation on "1,2.3.4", "1,234.567.89", "192,168.1.1". ground.py:88 runs over the entire PDF text for every invoice that reaches validation, so one version-string/IP/token anywhere in the document crashes run() (verified: no output dir created).
    • validate.py:80 / :103 — .quantize(_CENT, ROUND_HALF_UP) raises InvalidOperation when an amount exceeds the 28-digit Decimal context; a table row with a 30-digit figure reproduced the same batch-killing traceback through run().
    • Same class covers --verify-model: gliner.py:39/42 index ["entities"]/["text"] directly, and any model runtime error on one document escapes the same unguarded call (untested — integration test skips without INVOICE_MODEL_DIR).
    • Fix at the single seam: in process, catch unexpected exceptions around _problems and return Rejected(path, (f"{type(exc).__name__} …",)) — same pattern as ingest, stays diagnosable via reason.txt, covers verifiers too. For the grounding case also fix number_readings semantically: a token it cannot convert is not a number — skip it, so such invoices are accepted rather than all landing in review.

Should fix (low)

  • output.py:86 — shutil.copy2 still re-raises after the is_file() guard. Reproduced: a stat-able but unreadable PDF (mode 000, e.g. dropped into a shared inbox by another user) → read_text correctly rejects it, then copy2 raises PermissionError → the run dies mid-review-write after the CSVs are written, remaining .reason.txt entries lost, raw traceback. Make the copy best-effort (try/except OSError → note the copy failure in the reason file).
  • eval/__main__.py:34-37 — only FileNotFoundError is handled; a malformed gold.jsonl (blank line, bad JSON, invalid invoice record) still surfaces a raw json/pydantic traceback instead of parser.error — same §1.XVII class as the empty-gold fix added right below it. Add the parse/validation errors to the existing except.

Info

  • __main__.py:35 glob("*.pdf") is case-sensitive: a folder containing only INVOICE.PDF reports "no PDF files" (arguably fine, noting for completeness).
  • Previous review's info items are settled: the CSV-injection/negative-amount trade-off is now documented in CHANGELOG, and score() documents its non-empty precondition with the CLI guarding it.
[STATUS: FIXES_REQUIRED] ## Verified first - `ruff check .` + `ruff format --check .` → clean; `lint-imports` → 4/4 contracts kept; `pytest -m "not integration"` → **326 passed** (6 new tests from the fix commit), 2 deselected. - No `noqa`/`type: ignore`/TODO in the diff; the only bare-ish catch is the deliberate, commented `except Exception` at `ingest.py:23`. - **Previous review's must-fixes confirmed resolved**: fuzzed **300** corrupt PDFs (truncations + byte flips) through `pipeline.process` → **0 escapes** (was 7/70); dangling-symlink and directory `x.pdf` repros now produce only `.reason.txt` (no crash); `needs_review/` rebuild and the empty-gold `parser.error` both land with tests. CHANGELOG documents each fix. ## Must fix - **`pipeline.py:33` — `validate`/grounding/verifiers run *outside* the `try`; two proven crashes abort the whole batch before `write_results`, zero CSVs out.** The ingest seam was hardened, but `except (IngestError, ExtractError)` at `pipeline.py:31` only covers lines 29–30; `_problems` (line 33 → 19) runs unguarded. End-to-end repros on PDFs that extract *successfully*: - **`formats.py:32`** — `number_readings` builds `Decimal(token…)` for tokens mixing `,` and `.` with ≥2 of the decimal kind: `numbers_in` raises `decimal.InvalidOperation` on `"1,2.3.4"`, `"1,234.567.89"`, `"192,168.1.1"`. `ground.py:88` runs over the **entire** PDF text for every invoice that reaches validation, so one version-string/IP/token anywhere in the document crashes `run()` (verified: no output dir created). - **`validate.py:80` / `:103`** — `.quantize(_CENT, ROUND_HALF_UP)` raises `InvalidOperation` when an amount exceeds the 28-digit `Decimal` context; a table row with a 30-digit figure reproduced the same batch-killing traceback through `run()`. - Same class covers `--verify-model`: `gliner.py:39`/`42` index `["entities"]`/`["text"]` directly, and any model runtime error on one document escapes the same unguarded call (untested — integration test skips without `INVOICE_MODEL_DIR`). - Fix at the single seam: in `process`, catch unexpected exceptions around `_problems` and return `Rejected(path, (f"{type(exc).__name__} …",))` — same pattern as ingest, stays diagnosable via `reason.txt`, covers verifiers too. For the grounding case also fix `number_readings` semantically: a token it cannot convert is not a number — skip it, so such invoices are *accepted* rather than all landing in review. ## Should fix (low) - **`output.py:86` — `shutil.copy2` still re-raises after the `is_file()` guard.** Reproduced: a stat-able but unreadable PDF (mode `000`, e.g. dropped into a shared inbox by another user) → `read_text` correctly rejects it, then `copy2` raises `PermissionError` → the run dies mid-review-write *after* the CSVs are written, remaining `.reason.txt` entries lost, raw traceback. Make the copy best-effort (`try/except OSError` → note the copy failure in the reason file). - **`eval/__main__.py:34-37` — only `FileNotFoundError` is handled; a malformed `gold.jsonl` (blank line, bad JSON, invalid invoice record) still surfaces a raw `json`/pydantic traceback** instead of `parser.error` — same §1.XVII class as the empty-gold fix added right below it. Add the parse/validation errors to the existing `except`. ## Info - `__main__.py:35` `glob("*.pdf")` is case-sensitive: a folder containing only `INVOICE.PDF` reports "no PDF files" (arguably fine, noting for completeness). - Previous review's info items are settled: the CSV-injection/negative-amount trade-off is now documented in CHANGELOG, and `score()` documents its non-empty precondition with the CLI guarding it.
All checks were successful
CI / checks (pull_request) Successful in 1m10s
CI / policy (pull_request) Successful in 2s
CI / python (pull_request) Successful in 40s
This pull request is broken due to missing fork information.
View command line instructions

Checkout

From your project repository, check out a new branch and test the changes.
git fetch -u origin feat/invoice-to-csv:feat/invoice-to-csv
git switch feat/invoice-to-csv

Merge

Merge the changes and update on Forgejo.

Warning: The "Autodetect manual merge" setting is not enabled for this repository, you will have to mark this pull request as manually merged afterwards.

git switch main
git merge --no-ff feat/invoice-to-csv
git switch feat/invoice-to-csv
git rebase main
git switch main
git merge --ff-only feat/invoice-to-csv
git switch feat/invoice-to-csv
git rebase main
git switch main
git merge --no-ff feat/invoice-to-csv
git switch main
git merge --squash feat/invoice-to-csv
git switch main
git merge --ff-only feat/invoice-to-csv
git switch main
git merge feat/invoice-to-csv
git push origin main
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
jonah/safe-ai-tools!6
No description provided.