All notes

Build notes10 min read

A black rectangle is not a redaction

Drawing a box over text leaves the text in the file. What a PDF still carries after the box, why Redact PDF rebuilds a marked page as an image, and how it checks its own export before you send it.

In this note
  1. A PDF keeps the words in more places than the page
  2. There are two honest ways to remove text
  3. The file is rebuilt, not edited in place
  4. The export checks itself, and says what it cannot prove
  5. A detector offers candidates, and the reader decides
  6. It runs in the browser, so the file is never uploaded
  7. A redaction tool owes you six things
  8. Before you send a redacted PDF, test it yourself

In January 2019, lawyers for Paul Manafort filed a PDF in federal court with passages blacked out. The same day, reporters were reading the blacked-out passages, because the black only covered the text and the text was still in the file. Anyone who selected the page, copied it and pasted it into a text editor got everything.

A black rectangle drawn over text in a PDF hides the text and removes nothing: the words stay in the file, where copying, searching and any program that reads text will find them.

The American Bar Association’s Judges’ Journal wrote the case up as one of a series of embarrassing redaction failures. Earlier incidents have the same shape: a 2005 Pentagon report whose blacked-out paragraphs readers restored by copy and paste, and a 2006 AT&T court filing with the same flaw. Each document was edited to look redacted, not to be redacted.

The failure is easy to repeat: any viewer with markup tools can draw the rectangle, and the page then looks finished. Redact PDF (opens in a new tab), which we built, takes the other route.

A PDF keeps the words in more places than the page

A PDF page is a list of drawing instructions, a content stream, that a viewer runs in order to produce a picture. Text is placed by instructions that carry the characters as data; a filled rectangle is another instruction, placed later in the list. The viewer paints the text, then paints the rectangle over it. Copy, search, screen readers and indexers never look at the paint. They read the data.

That is the obvious case. Two more need explaining.

The hidden text on a scan. A scanned document that has been through OCR, optical character recognition, can carry two things: the image of the page, and hidden text laid over it so the scan can be searched. Blacking out part of the image leaves the hidden text exactly where it was. The page looks right and searches wrong.

Incremental updates. The format allows a file to be changed by appending to it rather than rewriting it. Section 7.5.6 of the specification says the original contents are left intact and the changes are added after them. A new trailer gives the location of the previous cross-reference section, the earlier index of where each object sits. It is a sensible design for saving quickly and for keeping a signature valid. It also means a “redacted” file saved this way can still contain the complete, unredacted page, and anyone who follows the trail back can recover it.

The full list is longer, and a rectangle on the page touches none of it.

Where the content survives What it is How it is read back
Under the box Text drawn before the rectangle Select, copy and paste, or search
On a scan Hidden OCR text laid over the image Search, or copy
In an earlier revision The file as it stood before an incremental save Follow the previous cross-reference section
In metadata Title, author, subject and keywords, in the information dictionary and in XMP, a block of XML Document properties
In separate objects Bookmarks, comments, form field values, attached files The reader’s own panels
In a thumbnail A miniature image of the page, stored as its own object Inspect the file’s objects

The specification itself separates marking from removing. Section 12.5.6.23 defines a redaction annotation, a marker that “identifies content that is intended to be removed from the document”, as a first step. The second step, removing the content and every trace of it, is the application’s job. A tool that stops at the annotation, or at a plain rectangle, has shipped half the job.

There are two honest ways to remove text

The first is to edit the content stream: find the instructions that place the characters inside the marked region and delete them. Then do the same for the OCR text, the annotations, the metadata and the earlier revisions. This keeps the rest of the page as live text, and it is hard to do completely. Text can be positioned character by character, and a run of characters can straddle the edge of a region. A transformation matrix can land text inside the region from coordinates that look outside it. A Form XObject, a reusable block of drawing instructions, can draw the same content from somewhere else in the file. Even with every character gone, the spacing left behind can give a name away: a 2022 study of eleven redaction tools, Adobe Acrobat among them, found that they leaked information about the text they had redacted.

The second is to stop treating the page as text. Render the marked page to an image, burn the bars into the pixels, and build a new page from that image. No content stream from the original page is copied. There is nothing underneath, because there is no underneath: the page is a picture, and the picture does not contain the words.

Redact PDF takes the second route for any page that carries a mark, and keeps the text of every other page exactly as it was. The tool’s page states the cost. Every word on a redacted page, not only the covered ones, stops being selectable and searchable; a screen reader can no longer read it; the file is larger. The rebuild also drops the document’s bookmarks, its tag structure and its named destinations. We think the trade is right: the choice is between a less convenient file and a file that leaks. A picture still shows how wide the bar is, so where a name’s length would give it away, draw the bar wider.

The file is rebuilt, not edited in place

“Build a new page from the image” hides a second decision: build the file from scratch, too.

The tempting implementation is to open the original, replace the marked pages, and save. That inherits every object that was not on those pages, and, saved as an incremental update, it inherits the original pages too. Redact PDF instead constructs a new document and copies into it only what it has decided to keep: the untouched pages, and the rendered images for the marked ones. Nothing is appended to the original, which is left untouched. (The export also flattens form fields, turning each into fixed page content. The site has separate tools for that and for removing metadata.)

We learned why this is needed the hard way. An earlier version of Redact PDF replaced the marked page in place. The original page’s content stayed in the file as objects nothing pointed to, a bookmark kept a whole unredacted scan alive, and a text check passed it three times. The tool’s own guide, How to properly redact a PDF (opens in a new tab), has the full account.

Rebuilding is the only way we found to make a claim we could stand behind. The claim is not “we removed the text you marked”. It is “the new file contains nothing from the marked pages but pixels”. The first depends on the implementation finding every occurrence. The second depends only on the construction.

The export checks itself, and says what it cannot prove

A redaction tool has a particular obligation that most software does not: its failure is silent. A badly compressed image looks bad. A redacted page that still contains the words looks perfect.

So the last step in Redact PDF is a report. It opens the exported bytes again from scratch, in the same browser tab, and checks four things. Each marked page has no text left in it. Nothing you searched for remains on the pages that kept their text. No Social Security, card or account number is left there. The file carries no hidden text, metadata, attachments, scripts or earlier versions.

If anything it removed can still be found, it says so before you send the file anywhere. It also says what it cannot prove: a page rebuilt as an image has no text to search, so the report shows you those pages and asks you to look. This is the copy-and-paste test the Judges’ Journal describes, run by the tool instead of left to the person.

A detector offers candidates, and the reader decides

Much of what gets redacted has a shape: card, account and routing numbers, Social Security and employer ID numbers, phone numbers, email addresses, dates. Redact PDF’s find-and-redact panel looks for those shapes on every page and turns each match into a bar you can see and adjust. A number on page eleven is not missed because the reader was tired by page nine.

The detectors are patterns with a check behind them: a card number has to pass the checksum the banks use, and a routing number the ABA check. A match is still a candidate, not a verdict. An order reference can pass the card checksum by chance; a date in a footer may be nothing anyone needs to hide. Nothing is applied until you say so, and the final look through the document stays with the reader. One limit is built in: Find reads the text layer, so it cannot see a scanned page that has none, and the panel says which pages those are.

It runs in the browser, so the file is never uploaded

All of this happens on the person’s own device. The file is opened with the browser’s File API, worked on in memory, rendered and rebuilt there, and the download is assembled there. The tool’s privacy page says there is no upload endpoint and no storage bucket, because the site is static files with nothing behind them. It invites you to verify the claim in the simplest way: load the page, turn off the network, and keep working.

For a tool that runs on the web, we think this is the only architecture that makes sense. A document that needs redacting contains something the person does not want to share, and the first thing a server-side redaction service asks for is the unredacted file. We did not want a tool whose first step contradicts its purpose.

Very large files are the honest weak spot. With no server’s memory to spare, a long scanned document is slow, and on a phone it can run the browser out of memory. That is the price of what our Web Applications page describes: tools that do their work inside the browser tab.

A redaction tool owes you six things

Our list, from building one:

  1. It removes the content, not the appearance of the content.
  2. It builds the output rather than editing the input, so earlier revisions, metadata and attachments do not come along uninvited.
  3. It helps you find every occurrence, and says which pages it could not read.
  4. It checks its own export and tells you what it found, before you send the file.
  5. It never sees your document. The file does not leave the device.
  6. It tells you the cost. A rebuilt page is not searchable and a screen reader cannot read it; the file is bigger; the bookmarks are gone. Hiding the trade-off would be a smaller version of the original problem.

None of that is novel. The failures above run from 2005 to 2019, and guidance has existed since the year of the first: the NSA’s Redacting with Confidence is dated December 2005. What we wanted, and built, is a tool that defaults to the slow, complete method and then audits itself.

Before you send a redacted PDF, test it yourself

Whatever tool made the file, ours included:

  1. Redact a copy and keep the original. A real redaction cannot be undone.
  2. Open the export, select all, copy, and paste into a plain text editor. Search it for each thing you removed. If a word appears, the file is not redacted.
  3. Treat a clean paste as the weaker proof. It reads the pages, and what a careless tool leaves behind can sit outside them.
  4. Look at every page that was rebuilt as an image. A search cannot read a picture, so your eyes are the check.
  5. Send the export, never the working file.

Redact PDF is one of the tools in Kordal Labs, the software we build and run ourselves. It is free, with no account and no upload; its privacy page says the plan is to pay for the site with advertising, which is not running yet. If you have a document workflow where the obvious implementation looks right and leaks, tell us about it.

Sources

  1. Embarrassing Redaction Failures, by Judge Herbert B. Dixon Jr. (The Judges’ Journal, 2019) (opens in a new tab)American Bar Association. Accessed 3 Oct 2026
  2. Failed redaction reveals Paul Manafort’s ‘lies to FBI’ (8 January 2019) (opens in a new tab)BBC News. Accessed 3 Oct 2026
  3. Readers ‘declassify’ US document (2 May 2005) (opens in a new tab)BBC News. Accessed 3 Oct 2026
  4. AT&T leaks sensitive info in NSA suit, by Declan McCullagh (26 May 2006) (opens in a new tab)CNET. Accessed 3 Oct 2026
  5. NSA: Redacting With Confidence, by Steven Aftergood (20 January 2006), on the NSA report of 13 December 2005 (opens in a new tab)Federation of American Scientists. Accessed 3 Oct 2026
  6. ISO 32000-1:2008, Document management — Portable document format — Part 1: PDF 1.7, sections 7.5.5 File Trailer, 7.5.6 Incremental Updates, 8.10 Form XObjects, 9.4 Text Objects, 12.3 Document-Level Navigation, 12.5.6.23 Redaction Annotations, 12.8 Digital Signatures and 14.3 Metadata (opens in a new tab)ISO (the copy Adobe made available under agreement with ISO as PDF 32000-1:2008, archived by the Internet Archive). Accessed 3 Oct 2026
  7. Story Beyond the Eye: Glyph Positions Break PDF Text Redaction, by Maxwell Bland, Anushya Iyer and Kirill Levchenko (2022) (opens in a new tab)arXiv. Accessed 3 Oct 2026
  8. Hidden Gems in Acrobat DC: How to Optimize Hidden OCR Text (10 March 2016) (opens in a new tab)Adobe. Accessed 3 Oct 2026
  9. Using files from web applications (the File API) (opens in a new tab)MDN Web Docs. Accessed 3 Oct 2026
  10. Zero-Trust PDF Studio (Redact PDF), the redaction tool, its privacy page (updated 29 September 2026), its About page and its flatten and metadata tools (opens in a new tab)Kordal Systems. Accessed 3 Oct 2026
  11. How to properly redact a PDF (reviewed 29 September 2026) (opens in a new tab)Kordal Systems. Accessed 3 Oct 2026
  12. Check these claims (reviewed 29 September 2026) (opens in a new tab)Kordal Systems. Accessed 3 Oct 2026

Amiya BeraFounder, Kordal Systems

Published . Build notes.

More notes

All notes
  1. Build notes10 min read

    When supplier quotes cannot honestly be compared

    Two quotations for the same material can differ in unit, tax basis, delivery point and version, and a ranking that ignores any one of them is a guess. The rules we gave Rateslip for refusing to rank, and what it does instead.

  2. Field notes9 min read

    AI-generated software is a prototype until someone owns the failure modes

    Generating a working application has become cheap. What still costs money is knowing how it fails, who is told when it does, and what happens next: the work between a convincing demo and a system a business can run on.