Workflow Automation · BraivIQ AI Engineering Playbook
Building A PII Detection And Redaction Pipeline In Code: Rules, NER, LLM Classification And Auditable Masking Before Anything Leaves The Building
Every UK business sends documents, messages and datasets across boundaries where personal data should not travel - to a supplier, to a model, to a customer who is not the subject, to an archive - and under UK GDPR a single unredacted record in the wrong place is a reportable breach. The workflow that prevents it is a PII detection and redaction pipeline, and in 2026 it is one of the highest-value automations a business can build, because the pieces have matured: deterministic rules for structured identifiers, named-entity recognition for names and addresses, and language-model classification for the contextual personal data that neither catches. This playbook is a code-side guide to building one properly: a layered detector where each layer catches what the others miss, confidence scoring that routes uncertain spans to a human rather than guessing, redaction strategies from irreversible masking to reversible tokenisation with a vault, per-field audit logs that prove what was masked and why, evaluation with precision and recall on labelled documents where a miss is a breach, and one pipeline that runs in both batch and streaming modes so the same protection applies to a nightly export and a live chat.
· 13 min read · By BraivIQ Engineering
3 layers - Deterministic rules, named-entity recognition and LLM classification - each catching what the others miss · A miss is a breach - Under UK GDPR, evaluation is about recall: an unmasked record in the wrong place is reportable · Mask or tokenise - Irreversible masking for data that must never come back; reversible tokenisation with a vault where it must · Batch + streaming - One pipeline, two modes - the same protection for a nightly export and a live message
Personal data crosses boundaries inside a business constantly, and most of those crossings are the ordinary business of getting work done: a support transcript sent to a supplier for analysis, a dataset exported for a model to learn from, a customer letter that references a third party, a document archived where more people can search it, a prompt sent to an AI provider. Under UK GDPR, each crossing is a place where a single unredacted name, address, account number or health detail in the wrong hands is a reportable breach - with notification to the people affected, the Information Commissioner's attention and, increasingly, a public record. The workflow that prevents it is a PII detection and redaction pipeline: a system that finds personal data in whatever is about to cross a boundary and masks or tokenises it before it does. It used to be crude - a handful of regular expressions that caught card numbers and missed everything else - and in 2026 it has become one of the highest-value automations a business can build, because the three techniques that together catch personal data reliably have all matured: deterministic rules for structured identifiers, named-entity recognition for names, addresses and organisations, and language-model classification for the contextual personal data that neither of the others can see. As an AI Agency Developer London that builds automation platforms for regulated UK clients, we treat this pipeline as foundational infrastructure, and this playbook is how to build it properly.
The Layered Detector
The detector is the heart of the pipeline and it is best built as a sequence of layers that each annotate the same text with spans - start, end, type, confidence, and which layer found it - so that later stages work on a single merged set of findings rather than three disagreeing ones. The first layer is rules: regular expressions for each structured identifier type, each paired with a validator that checks the checksum or format constraints so that a sixteen-digit number is only flagged as a card if it passes the Luhn check, a UK postcode only if its structure is valid, an NI number only if its letters are permitted - validation is what keeps rule-based detection precise. The second layer is named-entity recognition: a model trained to tag persons, locations, organisations and similar entities in text, run over the document with the rule-detected spans already marked so it does not re-classify them, and tuned for your domain, because a general NER model trained on news performs poorly on medical notes or trading chat. The third layer is language-model classification: a small, cost-effective model given a chunk of text and asked, with a constrained structured output, to list any spans that constitute personal data not already marked, and why - the layer that catches contextual identification and sensitive facts, run last because it is the most expensive and the least precise. A merge step then reconciles overlapping spans by precedence and combines confidences, producing one annotated document. The design principle is precision where it is cheap - rules with validation - and recall where it is hard - NER and the model - with each layer's findings and confidence preserved for the decisions that follow.
- Rules with validators - regex per identifier type paired with checksum and format checks (Luhn for cards, structure for postcodes and NI numbers) for precision.
- Domain-tuned NER - persons, addresses, organisations, dates of birth; tune on your document types, because news-trained models miss clinical or financial text.
- LLM classification last - a small model with constrained structured output flags contextual personal data the other layers cannot see.
- Merge to one annotation - reconcile overlapping spans by precedence, keep every layer's confidence and provenance.
- Never rely on one layer - each has a blind spot exactly the shape of the others.
Confidence, Redaction Strategies, And Human Review
A finding is not a decision, and the pipeline's second half turns annotated spans into safe output through policy. Confidence drives routing: high-confidence findings are redacted automatically, low-confidence spans - a name that might be a place, a number that might be a reference rather than an identifier, a model flag with weak justification - are queued for a human to confirm or dismiss, with the document page and the proposed redaction shown together, so that a person decides the uncertain cases in seconds rather than reviewing whole documents. The redaction itself is a strategy chosen per data type and per destination, and the choice matters. Irreversible masking replaces the span with a placeholder or a category label and is right when the data must never come back - an export to a third party, a training set. Reversible tokenisation replaces the span with a token that maps, in a separately secured vault, back to the original, and is right when the process downstream must not see the data but the business must be able to re-identify later - a support case routed through an external model, a document that must be restored for the customer. Format-preserving replacement substitutes a realistic but fake value of the same shape, so systems that validate formats keep working on masked data. Consistency across a document matters too: the same person must map to the same token everywhere, or the redacted text becomes incoherent. And each decision - what was found, by which layer, at what confidence, what strategy was applied, who approved an uncertain span - is written to a per-field audit log, because the purpose of the pipeline is partly to prove to a regulator that protection was applied, and an unlogged redaction proves nothing.
Evaluation, And One Pipeline For Batch And Streaming
The pipeline must be measured before it is trusted, and the measurement has an asymmetry the team must internalise: a false positive is an over-redaction that a human can restore, while a false negative is personal data that left the building - a breach. Evaluation therefore centres on recall per data type on a labelled set of your own documents, with precision tracked so the human-review queue stays manageable, and the labelled set must reflect your real document population - the formats, the domains, the messiness - because a detector that scores well on clean samples and misses names in scanned letters is a liability with a good dashboard. Run the evaluation in continuous integration so that a change to a rule, a model or a threshold is measured before it ships, and feed every human correction from the review queue back into the labelled set, so recall improves over time rather than staying where it was on day one. Finally, build the pipeline once and run it in two modes: batch, for nightly exports, archive migrations and dataset preparation, where throughput matters and latency does not; and streaming, for live messages, chat and outbound communications, where each item must be processed in milliseconds before it is sent. The detector layers, the policy and the audit are identical in both; only the orchestration differs - which is exactly why the pipeline belongs in a workflow-automation platform rather than as a script bolted onto one export. A business that applies the same protection to a live chat and a nightly file has closed the boundary properly; one that protects only the export it remembered to script has not.
The Bottom Line
Personal data crosses boundaries inside every UK business constantly, and under UK GDPR each crossing is a potential reportable breach, which makes a PII detection and redaction pipeline one of the highest-value automations a business can build in 2026 - and one that the matured techniques now make buildable properly. The detector is layered because no single technique suffices: validated rules for structured identifiers where precision is cheap, domain-tuned named-entity recognition for names and addresses, and a small language model with constrained output for the contextual personal data neither can see, merged into one annotated document with every finding's confidence and provenance preserved. Policy then turns findings into safe output - automatic redaction of high-confidence spans, a human-review queue for the uncertain ones, redaction strategies chosen per data type and destination from irreversible masking through reversible tokenisation with a vault to format-preserving replacement, consistency across a document, and a per-field audit log that proves protection was applied. It is evaluated on recall against labelled documents from your own population, because a miss is a breach, measured in continuous integration and improved by every human correction. And it runs as one pipeline in batch and streaming modes, so the same protection reaches the nightly export and the live message. Built this way it is foundational infrastructure for a regulated business; built as a script on one export it is a false sense of safety - and building the former is exactly the workflow-automation engineering we do.
References & Further Reading
- ICO - guide to UK GDPR: personal data breaches (what must be reported, and when): https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/personal-data-breaches/
- ICO - guidance on anonymisation, pseudonymisation and privacy-enhancing technologies: https://ico.org.uk/for-organisations/uk-gdpr-guidance-and-resources/data-sharing/anonymisation/
- Microsoft Presidio - open-source PII detection and anonymisation (rules plus NER reference architecture): https://microsoft.github.io/presidio/
- AppScale - structured output engineering: reliable JSON from LLMs (constrained classification outputs): https://appscale.blog/en/blog/structured-output-engineering-reliable-json-from-llms-2026
- Solutions Review - top MarTech news from the week of September 25th (Precisely's PII Masking Insight and Compliance Risk Scoring agents): https://solutionsreview.com/crm/2026/09/25/top-martech-news-from-the-week-of-september-25th/