Part of the CustodyFlow suite: defensible DFIR and eDiscovery workflow tools.
Defensible eDiscovery production tooling: load file QC, TAR validation statistics, and matter estimates, with a tamper-evident audit trail.
▶ Try it in your browser: no install, and files never leave your device
One click runs QC on a sample production with nine planted defects. You can also drop in your own production folder. The page's security policy blocks all network requests.
ProdFlow covers three questions that come up on almost every matter:
| Question | Command | Output |
|---|---|---|
| Is this production clean enough to serve, or to load? | prodflow qc |
QC report (Markdown + JSON), pass/fail exit code |
| Can we defend the review's recall? | prodflow elusion |
Exact confidence intervals + methodology memo |
| What will this matter cost and how long will review take? | prodflow estimate |
Low/high volume, hours and cost with every assumption visible |
It comes in two forms built from one specification:
| Browser | Command line | |
|---|---|---|
| Install | None: open the page | Python 3.9+, no other dependencies |
| Best for | Demos, spot checks, sharing results with attorneys | Full-size productions, scripting, chain-of-custody logging |
| Output | Same QC report, byte for byte | Same QC report, byte for byte |
CI runs both engines against the same fixtures (clean, defective and adversarial fixtures that trigger all 34 finding types) and fails if a single byte differs.
git clone https://github.com/ryandus/ProdFlow.git && cd ProdFlow
python samples/generate.py # builds clean + defective synthetic productions
python -m prodflow qc samples/defective_production --out ~/prodflow_outFAIL: 12 records, 9 errors, 5 warnings -> ~/prodflow_out
The synthetic sample production has nine planted defects. ProdFlow finds all of them, plus the downstream effects a real reviewer would hit. See the full report: examples/qc_report_defective.md.
Optional install: pip install . adds a prodflow command.
Validates a Concordance DAT against its OPT or LFP image load file, and both against the files actually on disk.
prodflow qc /path/to/PROD001 --out ./qc_PROD001 --operator "R. Hanks"
prodflow qc /path/to/PROD001 --dat DATA/PROD001.dat --lfp DATA/PROD001.lfp --field beg=ProdBegBates| Check | Code | Severity |
|---|---|---|
| Rows with the wrong number of fields | FIELD_COUNT |
error |
| Text qualifier never closed (corrupt export) | UNTERMINATED_QUALIFIER |
error |
| Duplicate, overlapping, reversed or malformed Bates | DUPLICATE_BATES BATES_OVERLAP BATES_REVERSED BATES_FORMAT BATES_PREFIX |
error |
| Gaps in the Bates sequence (withheld ranges needing a log) | BATES_GAP |
warning |
| Inconsistent zero-padding within a prefix | BATES_PADDING |
warning |
| Image count ≠ Bates span, DAT docs with no images, images with no DAT row | PAGE_MISMATCH NO_IMAGE_DOC ORPHAN_IMAGE_DOC |
error |
| Family ranges that don't contain the document or don't resolve to real docs | FAMILY_RANGE FAMILY_PARENT FAMILY_END |
error |
| Native / text / image paths that don't exist | FILE_MISSING |
error |
| Paths that are absolute or escape the production root | ABSOLUTE_PATH PATH_OUTSIDE_ROOT |
error |
| Paths whose case differs from disk (breaks on Linux platforms) | CASE_MISMATCH |
warning |
| Native doesn't match the SHA-256 in the load file | HASH_MISMATCH |
error |
| Unparseable date fields | DATE_FORMAT |
error |
| Empty extracted-text files | EMPTY_TEXT |
warning |
| Non-UTF-8 load file | ENCODING |
warning |
Field names are recognized across common vendor conventions (BegBates, BEGDOC, ProdBeg, BegAttach, NativeLink…). Anything else can be mapped with --field ROLE=NAME.
Exit codes: 0 pass · 1 findings with errors · 2 could not run (bad input, integrity failure).
prodflow convert DATA/PROD001.opt ../converted/PROD001.lfpPage counts are recomputed from document breaks. The OPT → LFP → OPT round trip is covered by tests.
After a technology-assisted review, the question opposing counsel asks is how much did you miss? ProdFlow answers it with an elusion test on a random sample of the null set:
prodflow elusion --null-set 250000 --sample 1500 --found 6 --responsive 42000 \
--matter "Synthetic v. Example" --out memo/elusion_memo.mdEstimated recall is 97.67% (95% CI: 95.08% – 99.13%). An estimated 1,000 responsive documents remain in the null set (95% CI: 367 – 2,171).
- Intervals are exact Clopper-Pearson binomial intervals, the conservative choice for contested methodology. The implementation is checked against
scipy.stats.betareference values in the test suite. - The memo states the method and its limitations in plain language, ready for a meet-and-confer or a declaration exhibit. See
examples/elusion_memo.md.
Plan the sample before you draw it:
prodflow sample-size --confidence 0.95 --margin 0.02 --population 250000 # -> 2379prodflow estimate --gb 120 --months 6 --out estimate.md| Measure | Low | High |
|---|---|---|
| Review set after culling | 117,600 | 288,000 |
| Review hours (incl. QC) | 2,352 | 7,680 |
| Total | $147,744 | $630,312 |
Every assumption (docs/GB, de-dupe rate, culling rate, review pace, rates, QC overhead, hosting) is printed with the result and can be overridden from a JSON file. The defaults are illustrative planning figures, not vendor quotes. Full output: examples/estimate_120gb.md.
These are guarantees the code enforces, not conventions:
| Guarantee | How |
|---|---|
| Source is read-only | Inputs are opened read-only. ProdFlow refuses to write output anywhere inside the production folder. |
| Integrity verified before and after | Every input load file is SHA-256 hashed before processing and re-hashed after. Any change aborts the run with exit code 2. |
| Tamper-evident audit log | Each run appends to audit.jsonl: UTC timestamp, tool version, operator, host, command, input hashes (before/after) and output hashes. Every entry includes the hash of the previous entry, so an edited, deleted or reordered entry breaks the chain. |
| Reproducible output | Reports are sorted and carry no timestamps. The same inputs always produce a byte-identical report. CI verifies this. |
| No silent failures | Malformed rows, undecodable files and unresolvable paths are reported as findings or stop the run. Nothing is skipped quietly. |
prodflow verify-log ./qc_PROD001/audit.jsonl
# Audit chain intact.docs/ browser version (GitHub Pages): engine.js is a line-for-line port of the Python checks
prodflow/ CLI and library (stdlib only)
qc.py production QC checks and report rendering
loadfile.py DAT / OPT / LFP parsing and conversion
tar.py Clopper-Pearson, elusion, recall, sample size, memo
estimate.py volume and cost model
integrity.py SHA-256, read-only guards, hash-chained audit log
samples/ generator for clean and defective synthetic productions
examples/ reports generated from the samples
tests/ unittest suite, incl. browser-vs-CLI parity (runs on Linux and Windows in CI)
Rebuild the samples at any time with python samples/generate.py. Output is byte-identical on every run.
python -m unittest discover -s tests -vThe suite covers: every planted defect is detected with the expected counts, a clean production passes with zero findings, the source tree hashes identically after a run, hostile paths (..\..\, C:\) are rejected, audit-log tampering is detected, and the statistics match reference values. The parity tests run the browser engine under Node.js and require byte-identical samples and reports, matching statistics and memos, and audit entries from the browser that verify with prodflow verify-log.
To run the browser version offline, for example on an air-gapped lab machine: python -m http.server -d docs, then open http://localhost:8000.
- The browser version reads load files and any natives it hashes into memory. For productions with multi-gigabyte load files, use the command line.
- Image files are checked for existence, not decoded. Blank or corrupt TIFFs aren't detected yet.
- Only SHA-256 hash fields are verified. MD5/SHA-1 fields are carried through but not checked.
- LFP image-type codes (
2TIFF,4JPEG,7PDF) follow common usage. Confirm them against your review platform's import spec. - The recall estimate treats the review's responsive count as exact and assumes a simple random sample. Both are stated in the generated memo.
All sample data is synthetic and generated by samples/generate.py. No case, client or law-enforcement data is included or was used to build this project.
Ryan C. Hanks — digital forensics and eDiscovery. Part of the CustodyFlow suite with TraceFlow and SyncFlow. Related work: EDRM Litigation Forensics Showcase · DFIR Investigation Framework
MIT License.