feat: accuracy harness + synthetic test set (stacked on #1) #2

Merged
jonah merged 3 commits from feat/eval-harness into feat/core-and-redact-proxy 2026-10-05 21:14:48 +02:00
Owner

Stacked on #1: base is feat/core-and-redact-proxy. Merge #1 first, then retarget this to main.

What

  • Seeded NL/DE/EN test-set generator with exact gold spans, incl. checksum-failing look-alikes as unlabelled negatives.
  • python -m redact_proxy.eval: span-level recall, precision, F1, leak rate per entity.
  • Hermetic CI test: 99% recall floor on checksum entities.
  • README accuracy table (spaCy md, 1,000 docs per language).
  • Behaviour change: image/audio/file content parts return 400 by default (they were forwarded unredacted). allow_non_text_parts: true opts out.
  • 0.2.0.

Measured (synthetic, so an upper bound)

Leak rate: NL 16.5%, DE 7.4%, EN 8.8%. Weak: bare postcode/PLZ (33-47%), credit card (63-76%), phone (78-88%).

Not done

gliner2-pii backend, --verifier, Gretel/ai4privacy corpora, lg-model comparison. Docker, Forgejo CI and live Ollama still unverified (see #1).

🤖 Generated with Claude Code

https://claude.ai/code/session_01JZCbAYmY516SAe7umoLKWt

Stacked on #1: base is `feat/core-and-redact-proxy`. Merge #1 first, then retarget this to `main`. ## What - Seeded NL/DE/EN test-set generator with exact gold spans, incl. checksum-failing look-alikes as unlabelled negatives. - `python -m redact_proxy.eval`: span-level recall, precision, F1, leak rate per entity. - Hermetic CI test: 99% recall floor on checksum entities. - README accuracy table (spaCy md, 1,000 docs per language). - **Behaviour change:** image/audio/file content parts return 400 by default (they were forwarded unredacted). `allow_non_text_parts: true` opts out. - 0.2.0. ## Measured (synthetic, so an upper bound) Leak rate: NL 16.5%, DE 7.4%, EN 8.8%. Weak: bare postcode/PLZ (33-47%), credit card (63-76%), phone (78-88%). ## Not done `gliner2-pii` backend, `--verifier`, Gretel/ai4privacy corpora, lg-model comparison. Docker, Forgejo CI and live Ollama still unverified (see #1). 🤖 Generated with [Claude Code](https://claude.com/claude-code) https://claude.ai/code/session_01JZCbAYmY516SAe7umoLKWt
- data/testsets/generate.py: seeded NL/DE/EN generator with exact gold spans
- redact_proxy.eval: span scorer + CLI (per-entity recall/precision, leak rate)
- CI recall floor for checksum entities; README accuracy table
- Behaviour change: image/audio/file content parts now get a 400 instead of
  being forwarded unredacted (opt out: allow_non_text_parts)
- Bump to 0.2.0

Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JZCbAYmY516SAe7umoLKWt
Co-Authored-By: Claude Sonnet 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01JZCbAYmY516SAe7umoLKWt
Author
Owner

🤖 Auto-review status · jonah/safe-ai-tools PR #2

📋 Review posted — fixes required, awaiting human action.

Last updated: 2026-10-05T15:46:24.395Z

<!-- opencode-status --> **🤖 Auto-review status** · `jonah/safe-ai-tools` PR #2 📋 **Review posted** — fixes required, awaiting human action. *Last updated: 2026-10-05T15:46:24.395Z*
jonah left a comment

[STATUS: FIXES_REQUIRED]

Verified working: hermetic suite 79 passed (pytest -m "not integration"), ruff check / ruff format --check / lint-imports all clean, eval CLI smoke-tested end-to-end (generate.py → python -m redact_proxy.eval --lang nl --backend patterns), version bump 0.2.0 + CHANGELOG consistent with the CI policy job, import-linter contract correctly updated for redact_proxy.eval.

Must fix

  • Security — fail-open content gate: apps/redact-proxy/src/redact_proxy/app.py:41-45 — _has_non_text_parts only rejects non-text content when it's a list. Content that is neither str nor list (e.g. "content": {"type":"text","text":"BSN 111222333"}, or a number) passes the gate and _redact_content (app.py:33-38) returns it unredacted, so with allow_non_text_parts: false the proxy still forwards PII it promised to block. Reproduced: gate rejects dict content? False / pass-through returns the original text. Mainstream upstreams reject dict-shaped content themselves, but an OpenAI-compatible shim may not — and §XIX says deny, not grant. Fix at the root (one place): treat any message content that isn't None, str, or an all-text-part list as non-redactable → 400 unless allow_non_text_parts. Add one test alongside test_non_text_parts_are_rejected_by_default in tests/test_app.py.

Should fix (low)

  • Threshold drift: apps/redact-proxy/src/redact_proxy/eval/__main__.py:14 — THRESHOLD = 0.4 duplicates Config.score_threshold (config.py:46). The spacy-* backends implicitly use the config default while patterns uses the local constant; if the default ever moves, backends silently diverge and the accuracy table compares different thresholds. Derive one value from Config(...) and pass it to both.
  • Raw traceback on missing data: eval/__main__.py:46 — generated synth_*.jsonl is gitignored (.gitignore:14), so a fresh clone runs python -m redact_proxy.eval into a bare FileNotFoundError. One-line friendly message pointing at data/testsets/generate.py (the README reproduce line) is enough.

Nits

  • tests/test_eval_floor.py:16-20 loads the generator out-of-tree via importlib.util.spec_from_file_location — works, but a plain package import would be less brittle than a path computed from parents[3].
  • README accuracy table (spaCy md numbers, 1000 docs) couldn't be reproduced here — no spaCy models installed, integration test deselected. Numbers are honest-looking and caveat-laden, but CI/integration run should spot-check at least one row.

Everything else in the diff is solid: the non-text-parts 400 defaults to fail-safe and is tested in both directions, score.py is pure and well-specified with tests for overlap/leak/wrong-type semantics, generate.py is deterministic with exact-by-construction gold spans, and the S311 ignore is narrowly scoped to the seeded generator.

[STATUS: FIXES_REQUIRED] Verified working: hermetic suite `79 passed` (`pytest -m "not integration"`), `ruff check` / `ruff format --check` / `lint-imports` all clean, eval CLI smoke-tested end-to-end (`generate.py` → `python -m redact_proxy.eval --lang nl --backend patterns`), version bump 0.2.0 + CHANGELOG consistent with the CI policy job, import-linter contract correctly updated for `redact_proxy.eval`. ## Must fix - **Security — fail-open content gate: `apps/redact-proxy/src/redact_proxy/app.py:41-45`** — `_has_non_text_parts` only rejects non-text content when it's a *list*. Content that is neither `str` nor `list` (e.g. `"content": {"type":"text","text":"BSN 111222333"}`, or a number) passes the gate and `_redact_content` (`app.py:33-38`) returns it **unredacted**, so with `allow_non_text_parts: false` the proxy still forwards PII it promised to block. Reproduced: `gate rejects dict content? False` / pass-through returns the original text. Mainstream upstreams reject dict-shaped content themselves, but an OpenAI-compatible shim may not — and §XIX says deny, not grant. Fix at the root (one place): treat any message content that isn't `None`, `str`, or an all-text-part list as non-redactable → 400 unless `allow_non_text_parts`. Add one test alongside `test_non_text_parts_are_rejected_by_default` in `tests/test_app.py`. ## Should fix (low) - **Threshold drift: `apps/redact-proxy/src/redact_proxy/eval/__main__.py:14`** — `THRESHOLD = 0.4` duplicates `Config.score_threshold` (`config.py:46`). The `spacy-*` backends implicitly use the config default while `patterns` uses the local constant; if the default ever moves, backends silently diverge and the accuracy table compares different thresholds. Derive one value from `Config(...)` and pass it to both. - **Raw traceback on missing data: `eval/__main__.py:46`** — generated `synth_*.jsonl` is gitignored (`.gitignore:14`), so a fresh clone runs `python -m redact_proxy.eval` into a bare `FileNotFoundError`. One-line friendly message pointing at `data/testsets/generate.py` (the README reproduce line) is enough. ## Nits - `tests/test_eval_floor.py:16-20` loads the generator out-of-tree via `importlib.util.spec_from_file_location` — works, but a plain package import would be less brittle than a path computed from `parents[3]`. - README accuracy table (spaCy md numbers, 1000 docs) couldn't be reproduced here — no spaCy models installed, integration test deselected. Numbers are honest-looking and caveat-laden, but CI/integration run should spot-check at least one row. Everything else in the diff is solid: the non-text-parts 400 defaults to fail-safe and is tested in both directions, `score.py` is pure and well-specified with tests for overlap/leak/wrong-type semantics, `generate.py` is deterministic with exact-by-construction gold spans, and the `S311` ignore is narrowly scoped to the seeded generator.
jonah force-pushed feat/eval-harness from 3ec5505e48 to d9a44b9d01 2026-10-05 20:38:10 +02:00 Compare
Author
Owner

🤖 Auto-review status · jonah/safe-ai-tools PR #2

📋 Review posted — fixes required, awaiting human action.

Last updated: 2026-10-05T18:48:59.337Z

<!-- opencode-status --> **🤖 Auto-review status** · `jonah/safe-ai-tools` PR #2 📋 **Review posted** — fixes required, awaiting human action. *Last updated: 2026-10-05T18:48:59.337Z*
jonah left a comment

[STATUS: FIXES_REQUIRED]

Verified on current HEAD (d9a44b9; the only commit since the previous review is an empty chore: trigger PR review, so no fixes were pushed): hermetic suite 81 passed, 1 deselected, ruff check / ruff format --check / lint-imports (2 contracts KEPT) all green.

Must fix (unaddressed high-severity)

  • Security — fail-open content gate still open: apps/redact-proxy/src/redact_proxy/app.py:41-45 (call site app.py:100) — _has_non_text_parts only inspects content when it is a list. Reproduced on this HEAD:
    • gate rejects dict content? False → _redact_content (app.py:33-38) returns {'type': 'text', 'text': 'BSN 111222333'} unredacted, so with allow_non_text_parts: false the proxy forwards PII it promised to 400.
    • gate rejects int content? False (any non-str, non-list content passes).
    • gate rejects image list? True — the list path works; the hole is exactly the shapes the list check doesn't cover.
      Treat any message content that isn't None, str, or an all-text-part list as non-redactable → 400 unless allow_non_text_parts (§XIX: deny, not grant), and add one test beside test_non_text_parts_are_rejected_by_default in tests/test_app.py.

New (low)

  • README reproduce line drops the edge cases: README.md:95 — generate.py --n 1000 writes 1006 lines (the 6 hand-written NL edge cases are appended after the n template docs), and --limit 1000 then slices them off: verified first 1000 include any edge? False. So the documented reproduction of the published accuracy table never scores the dotted BSN (111.222.333), spaced IBAN, or +factuur email — the cases the harness exists to catch. Fix by ordering edge cases first in generate(), or documenting --limit ≥ total lines.

Non-blocking note

Two low items from the previous review remain open (not re-flagged per scope): THRESHOLD = 0.4 duplication at eval/__main__.py:14 vs Config.score_threshold, and the bare FileNotFoundError at eval/__main__.py:46 on a fresh clone (generated synth_*.jsonl is gitignored).

[STATUS: FIXES_REQUIRED] Verified on current HEAD (`d9a44b9`; the only commit since the previous review is an empty `chore: trigger PR review`, so no fixes were pushed): hermetic suite **81 passed, 1 deselected**, `ruff check` / `ruff format --check` / `lint-imports` (2 contracts KEPT) all green. ## Must fix (unaddressed high-severity) - **Security — fail-open content gate still open: `apps/redact-proxy/src/redact_proxy/app.py:41-45` (call site `app.py:100`)** — `_has_non_text_parts` only inspects `content` when it is a `list`. Reproduced on this HEAD: - `gate rejects dict content? False` → `_redact_content` (`app.py:33-38`) returns `{'type': 'text', 'text': 'BSN 111222333'}` **unredacted**, so with `allow_non_text_parts: false` the proxy forwards PII it promised to 400. - `gate rejects int content? False` (any non-str, non-list content passes). - `gate rejects image list? True` — the list path works; the hole is exactly the shapes the list check doesn't cover. Treat any message `content` that isn't `None`, `str`, or an all-text-part list as non-redactable → 400 unless `allow_non_text_parts` (§XIX: deny, not grant), and add one test beside `test_non_text_parts_are_rejected_by_default` in `tests/test_app.py`. ## New (low) - **README reproduce line drops the edge cases: `README.md:95`** — `generate.py --n 1000` writes 1006 lines (the 6 hand-written NL edge cases are appended *after* the n template docs), and `--limit 1000` then slices them off: verified `first 1000 include any edge? False`. So the documented reproduction of the published accuracy table never scores the dotted BSN (`111.222.333`), spaced IBAN, or `+factuur` email — the cases the harness exists to catch. Fix by ordering edge cases first in `generate()`, or documenting `--limit` ≥ total lines. ## Non-blocking note Two low items from the previous review remain open (not re-flagged per scope): `THRESHOLD = 0.4` duplication at `eval/__main__.py:14` vs `Config.score_threshold`, and the bare `FileNotFoundError` at `eval/__main__.py:46` on a fresh clone (generated `synth_*.jsonl` is gitignored).
Author
Owner

🤖 Auto-review status · jonah/safe-ai-tools PR #2

✅ PR is clean — no fixes required.

Last updated: 2026-10-05T19:14:14.123Z

<!-- opencode-status --> **🤖 Auto-review status** · `jonah/safe-ai-tools` PR #2 ✅ **PR is clean** — no fixes required. *Last updated: 2026-10-05T19:14:14.123Z*
jonah left a comment

[STATUS: CLEAN]

Verified on HEAD 767fb11 (fix: close fail-open content gate, derive eval threshold from Config, edge cases first — the only substantive commit since the previous review).

Checks run

  • Hermetic suite: pytest -m "not integration" → 84 passed, 1 deselected (81+1 before; +3 new tests); full pytest → 85 passed.
  • ruff check . clean, ruff format --check . → 30 files formatted, lint-imports → 2 contracts KEPT.

Previous must-fix items — all addressed

  • Fail-open content gate (high) — apps/redact-proxy/src/redact_proxy/app.py:41-48 now whitelists (None | str | all-text-part list) instead of sniffing for lists; call site app.py:103. Reproduced on this HEAD: dict content blocked=True, int content blocked=True, list of plain strings blocked=True, text part with non-str text blocked=True, mixed list blocked=True; str/None/missing/empty-list still pass. Test added beside the existing gate test (tests/test_app.py:69-75, parametrized over dict and int).
  • README edge cases dropped by --limit (low) — data/testsets/generate.py:160 returns edge + docs. Verified: generate.py --n 1000 → 1006 lines, all six nl-edge-* docs (including dotted BSN 111.222.333) present in the first 1000, output byte-identical across runs. New test tests/test_eval_floor.py:24-25; README README.md:95 note added.

Previous non-blocking notes — also resolved

  • THRESHOLD = 0.4 duplication removed; _backend (eval/__main__.py:21-29) takes config.score_threshold (confirmed 0.4).
  • Missing test set now exits via parser.error with the generate hint (verified: exit 2, actionable message) instead of a traceback.

No new issues introduced by the fix commit.

[STATUS: CLEAN] Verified on HEAD `767fb11` (`fix: close fail-open content gate, derive eval threshold from Config, edge cases first` — the only substantive commit since the previous review). **Checks run** - Hermetic suite: `pytest -m "not integration"` → **84 passed, 1 deselected** (81+1 before; +3 new tests); full `pytest` → 85 passed. - `ruff check .` clean, `ruff format --check .` → 30 files formatted, `lint-imports` → 2 contracts KEPT. **Previous must-fix items — all addressed** - **Fail-open content gate (high)** — `apps/redact-proxy/src/redact_proxy/app.py:41-48` now whitelists (`None` | `str` | all-text-part list) instead of sniffing for lists; call site `app.py:103`. Reproduced on this HEAD: `dict content blocked=True`, `int content blocked=True`, `list of plain strings blocked=True`, `text part with non-str text blocked=True`, `mixed list blocked=True`; `str`/`None`/missing/empty-list still pass. Test added beside the existing gate test (`tests/test_app.py:69-75`, parametrized over dict and int). - **README edge cases dropped by `--limit` (low)** — `data/testsets/generate.py:160` returns `edge + docs`. Verified: `generate.py --n 1000` → 1006 lines, all six `nl-edge-*` docs (including dotted BSN `111.222.333`) present in the first 1000, output byte-identical across runs. New test `tests/test_eval_floor.py:24-25`; README `README.md:95` note added. **Previous non-blocking notes — also resolved** - `THRESHOLD = 0.4` duplication removed; `_backend` (`eval/__main__.py:21-29`) takes `config.score_threshold` (confirmed 0.4). - Missing test set now exits via `parser.error` with the generate hint (verified: exit 2, actionable message) instead of a traceback. No new issues introduced by the fix commit.
jonah merged commit b21263c7eb into feat/core-and-redact-proxy 2026-10-05 21:14:48 +02:00
jonah deleted branch feat/eval-harness 2026-10-05 21:14:48 +02:00
Sign in to join this conversation.
No reviewers
No labels
No milestone
No project
No assignees
1 participant
Notifications
Due date
The due date is invalid or out of range. Please use the format "yyyy-mm-dd".

No due date set.

Dependencies

No dependencies set.

Reference
jonah/safe-ai-tools!2
No description provided.