feat: accuracy harness + synthetic test set (stacked on #1) #2
Loading…
Add table
Add a link
Reference in a new issue
No description provided.
Delete branch "feat/eval-harness"
Deleting a branch is permanent. Although the deleted branch may continue to exist for a short time before it actually gets removed, it CANNOT be undone in most cases. Continue?
Stacked on #1: base is
feat/core-and-redact-proxy. Merge #1 first, then retarget this tomain.What
python -m redact_proxy.eval: span-level recall, precision, F1, leak rate per entity.allow_non_text_parts: trueopts out.Measured (synthetic, so an upper bound)
Leak rate: NL 16.5%, DE 7.4%, EN 8.8%. Weak: bare postcode/PLZ (33-47%), credit card (63-76%), phone (78-88%).
Not done
gliner2-piibackend,--verifier, Gretel/ai4privacy corpora, lg-model comparison. Docker, Forgejo CI and live Ollama still unverified (see #1).🤖 Generated with Claude Code
https://claude.ai/code/session_01JZCbAYmY516SAe7umoLKWt
🤖 Auto-review status ·
jonah/safe-ai-toolsPR #2📋 Review posted — fixes required, awaiting human action.
Last updated: 2026-10-05T15:46:24.395Z
[STATUS: FIXES_REQUIRED]
Verified working: hermetic suite
79 passed(pytest -m "not integration"),ruff check/ruff format --check/lint-importsall clean, eval CLI smoke-tested end-to-end (generate.py→python -m redact_proxy.eval --lang nl --backend patterns), version bump 0.2.0 + CHANGELOG consistent with the CI policy job, import-linter contract correctly updated forredact_proxy.eval.Must fix
apps/redact-proxy/src/redact_proxy/app.py:41-45—_has_non_text_partsonly rejects non-text content when it's a list. Content that is neitherstrnorlist(e.g."content": {"type":"text","text":"BSN 111222333"}, or a number) passes the gate and_redact_content(app.py:33-38) returns it unredacted, so withallow_non_text_parts: falsethe proxy still forwards PII it promised to block. Reproduced:gate rejects dict content? False/ pass-through returns the original text. Mainstream upstreams reject dict-shaped content themselves, but an OpenAI-compatible shim may not — and §XIX says deny, not grant. Fix at the root (one place): treat any message content that isn'tNone,str, or an all-text-part list as non-redactable → 400 unlessallow_non_text_parts. Add one test alongsidetest_non_text_parts_are_rejected_by_defaultintests/test_app.py.Should fix (low)
apps/redact-proxy/src/redact_proxy/eval/__main__.py:14—THRESHOLD = 0.4duplicatesConfig.score_threshold(config.py:46). Thespacy-*backends implicitly use the config default whilepatternsuses the local constant; if the default ever moves, backends silently diverge and the accuracy table compares different thresholds. Derive one value fromConfig(...)and pass it to both.eval/__main__.py:46— generatedsynth_*.jsonlis gitignored (.gitignore:14), so a fresh clone runspython -m redact_proxy.evalinto a bareFileNotFoundError. One-line friendly message pointing atdata/testsets/generate.py(the README reproduce line) is enough.Nits
tests/test_eval_floor.py:16-20loads the generator out-of-tree viaimportlib.util.spec_from_file_location— works, but a plain package import would be less brittle than a path computed fromparents[3].Everything else in the diff is solid: the non-text-parts 400 defaults to fail-safe and is tested in both directions,
score.pyis pure and well-specified with tests for overlap/leak/wrong-type semantics,generate.pyis deterministic with exact-by-construction gold spans, and theS311ignore is narrowly scoped to the seeded generator.3ec5505e48tod9a44b9d01🤖 Auto-review status ·
jonah/safe-ai-toolsPR #2📋 Review posted — fixes required, awaiting human action.
Last updated: 2026-10-05T18:48:59.337Z
[STATUS: FIXES_REQUIRED]
Verified on current HEAD (
d9a44b9; the only commit since the previous review is an emptychore: trigger PR review, so no fixes were pushed): hermetic suite 81 passed, 1 deselected,ruff check/ruff format --check/lint-imports(2 contracts KEPT) all green.Must fix (unaddressed high-severity)
apps/redact-proxy/src/redact_proxy/app.py:41-45(call siteapp.py:100) —_has_non_text_partsonly inspectscontentwhen it is alist. Reproduced on this HEAD:gate rejects dict content? False→_redact_content(app.py:33-38) returns{'type': 'text', 'text': 'BSN 111222333'}unredacted, so withallow_non_text_parts: falsethe proxy forwards PII it promised to 400.gate rejects int content? False(any non-str, non-list content passes).gate rejects image list? True— the list path works; the hole is exactly the shapes the list check doesn't cover.Treat any message
contentthat isn'tNone,str, or an all-text-part list as non-redactable → 400 unlessallow_non_text_parts(§XIX: deny, not grant), and add one test besidetest_non_text_parts_are_rejected_by_defaultintests/test_app.py.New (low)
README.md:95—generate.py --n 1000writes 1006 lines (the 6 hand-written NL edge cases are appended after the n template docs), and--limit 1000then slices them off: verifiedfirst 1000 include any edge? False. So the documented reproduction of the published accuracy table never scores the dotted BSN (111.222.333), spaced IBAN, or+factuuremail — the cases the harness exists to catch. Fix by ordering edge cases first ingenerate(), or documenting--limit≥ total lines.Non-blocking note
Two low items from the previous review remain open (not re-flagged per scope):
THRESHOLD = 0.4duplication ateval/__main__.py:14vsConfig.score_threshold, and the bareFileNotFoundErrorateval/__main__.py:46on a fresh clone (generatedsynth_*.jsonlis gitignored).🤖 Auto-review status ·
jonah/safe-ai-toolsPR #2✅ PR is clean — no fixes required.
Last updated: 2026-10-05T19:14:14.123Z
[STATUS: CLEAN]
Verified on HEAD
767fb11(fix: close fail-open content gate, derive eval threshold from Config, edge cases first— the only substantive commit since the previous review).Checks run
pytest -m "not integration"→ 84 passed, 1 deselected (81+1 before; +3 new tests); fullpytest→ 85 passed.ruff check .clean,ruff format --check .→ 30 files formatted,lint-imports→ 2 contracts KEPT.Previous must-fix items — all addressed
apps/redact-proxy/src/redact_proxy/app.py:41-48now whitelists (None|str| all-text-part list) instead of sniffing for lists; call siteapp.py:103. Reproduced on this HEAD:dict content blocked=True,int content blocked=True,list of plain strings blocked=True,text part with non-str text blocked=True,mixed list blocked=True;str/None/missing/empty-list still pass. Test added beside the existing gate test (tests/test_app.py:69-75, parametrized over dict and int).--limit(low) —data/testsets/generate.py:160returnsedge + docs. Verified:generate.py --n 1000→ 1006 lines, all sixnl-edge-*docs (including dotted BSN111.222.333) present in the first 1000, output byte-identical across runs. New testtests/test_eval_floor.py:24-25; READMEREADME.md:95note added.Previous non-blocking notes — also resolved
THRESHOLD = 0.4duplication removed;_backend(eval/__main__.py:21-29) takesconfig.score_threshold(confirmed 0.4).parser.errorwith the generate hint (verified: exit 2, actionable message) instead of a traceback.No new issues introduced by the fix commit.