The archive had accumulated over years across multiple business lines — contracts, statements of work, attestations, vendor agreements, regulated correspondence. The standing protocol was a 1% quarterly sample: it produced an audit report, not a confident assessment of the corpus.
The pipeline ran every document through three layers: classification (type, business line, governing policy sections), rule evaluation against the client’s own policy taxonomy — their rules, not generic compliance heuristics — and pattern surfacing across the full corpus.
The third layer produced the most surprising findings: provisions systematically missing from one business line during one date range — process gaps, not one-off oversights. Patterns invisible at sample scale became obvious at full coverage.
The by-product mattered as much as the audit: a fully indexed, classified archive that any future audit re-runs against in days instead of quarters. The reading became machine work. Judgment on every flagged exception stayed with the compliance team.