Teams adding a language model to a document workflow run into the same problem early. The documents worth summarising, classifying or extracting from are the ones full of personal data: loan files, claims, support tickets, contracts, medical letters. Sending them to a third-party model as they are is something most security reviews will not sign off on.
The obvious fix is to redact first. The obvious way to redact, black boxes, works badly here.
Why black boxes break the task
Imagine asking a model to check whether the account holder on a bank statement matches the applicant on a loan form. With both account numbers boxed out, the model sees two blanks and cannot say whether they match. Ask it to summarise a dispute involving two cards and it cannot tell them apart. Ask it to pull every identifier into a structured form and it returns a list of █████.
A model reasons over structure. It needs to know that something is there, what kind of thing it is, and whether it is the same thing that appears elsewhere. It does not need the value.
Tokens: keep the structure, drop the value
Token mode replaces each identifier with a label of its type and a number:
| Original | Sent to the model |
|---|---|
| Card 4111 1111 1111 1111 was charged twice. | Card [PAYMENT_CARD_1] was charged twice. |
| Refund to GB82WEST12345698765432. | Refund to [IBAN_1]. |
| Customer SSN 123-45-6789, spouse SSN 078-05-1120. | Customer SSN [US_SSN_1], spouse SSN [US_SSN_2]. |
| Confirm SSN 123-45-6789 on file. | Confirm SSN [US_SSN_1] on file. |
The same value always gets the same token within a document, so the model can see that the SSN in paragraph one and paragraph six are the same, and that the spouse's is different. It can answer "do these match?" correctly without ever seeing either number.
Tokens are consistent within one document and never across documents. [US_SSN_1] in one file and [US_SSN_1] in another are unrelated. If tokens were stable across files, they would become a pseudonymous ID that could link a person's records together, which is the re-identification risk redaction was supposed to remove.
Mapping the answer back
The redaction response never hands the real values back. Its replacements map pairs each token with a masked preview, such as [US_SSN_1] → *******6789. That is enough to audit what was replaced, and useless to anyone who intercepts it.
If the model's output needs real values, for example a filled form, you already have them: they are in the original document, which never left your system. The response lists every identifier with its type and character offsets in the original text, so your code can put the values back after the model answers. The model provider's logs hold only labels.
When to use which mode
| Mode | Output | Good for |
|---|---|---|
token | [IBAN_1] | LLM prompts, analytics, anything that reasons over the text |
mask | ************1111 | Support staff who need the last four digits to match a customer's card |
pseudonym | A reserved stand-in in the same format | Test environments where downstream code still has to parse the field |
box | Content removed and covered | Documents leaving the organisation for good |
Documents, not just text
Most of this applies to PDFs as well as strings. In token mode the identifier is removed from the PDF's content stream and the label is written in its place, so a model reading the extracted text gets the token. A drawn box with the text still underneath would leak the value straight into the prompt. Scanned pages are OCR'd first, which matters, because a scanned page passed to a vision model carries every identifier on it whatever its text layer says.
What redaction does not cover
Checksum-validated detection is very good at structured identifiers: cards, IBANs, SSNs, NHS numbers, passport MRZs. It does not find personal names or street addresses, which have no structure to check and need a statistical model with its own error rate. If your threat model includes names, add that layer deliberately and measure it. Do not assume an identifier redactor covers it.
Use mode=token on MaskAadhaar's global redaction endpoints, or send text straight to /api/v1/global/redact/text. Same API key and quota as the Aadhaar and PAN endpoints, and nothing is charged when nothing is found.