Here's a little experiment you can try right now, if you've got a redacted PDF lying around somewhere. Open it. Press Ctrl+A (Cmd+A on a Mac). Copy. Paste the whole thing into Notepad or any plain text editor.
Did the "hidden" name show up?
Quite often it does. Just sitting there in plain text, between two perfectly ordinary sentences, as if nothing had happened.
Nobody hacked anything to get it, either. There's no clever trick involved. A black rectangle drawn over a line of text in a PDF is exactly what it sounds like: a rectangle, painted on top of a line of text. The text underneath has no idea it's supposed to be a secret. It stays in the file, intact and selectable, and (this is the part that tends to ruin somebody's week) searchable.
Think overhead projector, not sheet of paper
If you're old enough to remember overhead projectors at school, the humming ones with the transparent sheets the teacher would stack on top of each other, you already understand PDFs better than most software manuals explain them.
A PDF page isn't a flat picture. It's a set of instructions. Put this letter, in this font, at these coordinates. Draw a line here. Place an image there. Your PDF viewer reads those instructions one after another and paints the result on your screen, and whatever gets painted later ends up on top of whatever came before.
So when you draw a black box over a sentence, you're adding one more instruction to the pile: paint a black rectangle here. The instruction that says "write the sentence" is still in the file. On screen, the rectangle wins, because it was painted last. But copy and paste, the search function, screen readers, text extraction tools, search engines that index PDFs... they don't look at the painted picture at all. They read the text instructions directly. For them, the black box barely exists.
With a real piece of paper and a thick marker, the ink actually covers the letters. On a transparency, you've just laid another sheet on top. Lift it, and there it all is.
It's happened to people with entire legal departments
It's tempting to file this under rookie mistakes. It isn't. Or rather, it's a rookie mistake that professionals keep making, because the result looks so convincingly finished.
In January 2019, lawyers for Paul Manafort, Donald Trump's former campaign chairman, filed a court document with several passages blacked out. Journalists noticed pretty quickly that the black bars could be bypassed by copying the text and pasting it somewhere else. The hidden passages turned out to contain some of the most newsworthy material in the whole filing: allegations that Manafort had shared campaign polling data with Konstantin Kilimnik, a consultant the FBI had assessed as having ties to Russian intelligence. Not exactly the paragraph you'd want to leak by accident.
Roughly ten years earlier, in December 2009, the US Transportation Security Administration posted a version of an airport screening procedures manual online. The sensitive parts (and in a document about airport security, there were quite a few) had been covered with black boxes. Same story. The text underneath was still in the file, and fully readable copies were soon making the rounds.
Two very different organizations, same mistake. Neither was careless in the usual sense. They did something that looked done. That's the whole trap, really.
The "redactions" that don't actually redact
Beyond the classic black rectangle, there's a surprisingly creative list of ways to hide text that leave it exactly where it was.
The black highlighter in Word
Someone opens the document in Word, selects the sensitive bits, picks black as the highlight color and exports to PDF. It looks perfect. It is also, functionally, just highlighted text. Highlighting changes the background behind the letters, not the letters themselves.
White (or black) text
A close cousin: changing the font color so the text blends into the background. White text on a white page, black text on a black bar. Drag your mouse across the area and it lights right back up.
Covering a scanned page that has been OCR'd
This one's sneaky, and it catches people who think they're being careful. When a scanned document goes through text recognition (OCR), many tools add an invisible text layer on top of the image, so you can search and copy from the scan. Now, if you draw a box over the image of a word, you've covered the picture of that word. The invisible text layer? Untouched. Ctrl+F finds it happily.
Cropping
In a lot of PDF tools, cropping a page doesn't cut anything off. It just tells the viewer to display a smaller portion of the page. Everything outside the crop area often stays in the file, and plenty of programs will show it again with a couple of clicks.
Blurring and pixelating
Mostly a screenshot problem, but those screenshots end up in PDFs all the time. Blur and pixelation feel destructive. They aren't; they scramble, they don't delete. Security researchers have shown more than once that pixelated text can be reconstructed, particularly when the font is known and the text is short. Like, say, an account number. If it matters, don't blur it. Remove it.
"I deleted the text and saved"
Better. Usually fine, even. But not always: some programs save changes by appending them to the end of the file instead of rewriting the whole thing (an "incremental save", in PDF jargon). Older versions of the content can then linger inside the file, invisible in a normal viewer but recoverable by anyone who goes looking. Whether this affects you depends entirely on the program you used, which is kind of the annoying part.
Then there's all the stuff you never see on the page
Let's say the visible text is genuinely gone. Great. A PDF can still carry a small suitcase of extra information around with it, and people tend to forget it's there at all.
Document properties. The author name, the title field, the software used, creation and modification dates. The title field in particular loves to hold the original Word file name. Something like "Miller_dismissal_draft3_DO_NOT_SEND.docx", which says rather a lot about a document in which you've carefully blacked out the name Miller.
Comments and annotations. Sticky notes, review comments, markup from three rounds of edits ago. Often tucked away in a side panel that the recipient's viewer may well open by default.
Form fields. Values typed into form fields can survive even when the field looks empty or has been covered up.
Bookmarks. The outline sidebar sometimes has entries like "Section 4 – Settlement amount for J. Novak." Oops.
Attachments and layers. PDFs can contain embedded files and hidden layers. Rare in everyday documents, sure. But when they are there, nobody remembers they're there.
And one more that's easy to underestimate: the size of the box itself. A redaction bar that's exactly as wide as a short surname, in a document where everyone involved knows the four people who could possibly be meant, isn't much of a secret. Academic researchers have gone a step further and shown that the precise spacing and positioning of the letters around a redaction can help narrow down what was removed. For a bank statement you're sending to a landlord, this won't keep you up at night. For a court filing or a whistleblower document? It absolutely should be on your radar.
Okay. So what actually works?
The core idea is simple, even if the tools aren't always: the information has to be removed from the file, not hidden inside it. There are a few honest ways to get there.
A proper redaction tool. Some professional PDF software comes with dedicated redaction functions that delete the underlying text and graphics in the marked area, and can strip hidden data like metadata and comments as well. Adobe Acrobat Pro is the best-known example. These work well when used correctly, but "correctly" usually means marking the areas and then explicitly applying the redactions, a second step people skip more often than you'd believe. They also generally aren't free.
Flattening the page into an image. This is the approach behind the mousePDF redaction tool, and admittedly it's the brute-force option. You cover whatever needs to go, and when you save, the page is flattened into an image. No text layer left to copy from, no hidden sentence under the box, because there's no "under" anymore. Just pixels, and the pixels where the box was are black. It's hard to get wrong, which is sort of the point.
There are trade-offs, and it's only fair to mention them. Text on a flattened page can't be selected or searched anymore. The file may get a bit bigger. Zoomed in very far, the page can look a hair softer than crisp original text. If you need the page searchable again, you can run text recognition over it afterwards. OCR will only ever see what's visible, and the black parts are, well, black.
The analog route. Print it. Black it out with a thick marker, ideally twice, because one pass with a cheap marker lets letters ghost through when the light hits it. Scan it back in. Clunky? Very. Your document will look like it's been through a fax machine from 1994. But the information is genuinely gone from the digital file.
A small detour about where you do your redacting
This bit is a little ironic, and we'd be lying if we said it wasn't part of the reason mousePDF exists in the first place.
A lot of free online PDF tools work by uploading your file to a server, processing it there and sending the result back. For plenty of tasks, maybe that's a fair trade. But think for a second about what redaction means. You have a document containing something so sensitive that you've decided nobody else should see it. And the very first thing you do is send the complete, unredacted version to a server run by a company you've probably never heard of, in a country you might not be able to name.
Kind of defeats the purpose, doesn't it?
Many of these services are perfectly legitimate and delete files after an hour or so. But you're taking their word for it. And trusting that their security is good, and that nobody misconfigured a storage bucket that morning.
mousePDF does all of its processing locally in your browser. The file is opened on your device, edited on your device and saved on your device, and it never gets sent to us. If you're curious about the technical side, we've written up how the editor works without a server. Whatever tool you end up using, though, ask where your file actually goes before you drop something sensitive into it.
Who actually needs to care about this?
More people than you'd guess. Most of them aren't lawyers.
You're renting a flat and the landlord wants to see bank statements, but not necessarily every pharmacy receipt and every transfer to your ex. You work in HR and you're answering a data access request under the GDPR, where the employee is entitled to their own data, not their colleagues' names in the same email thread. You're filing documents with a US federal court, where the rules generally require things like Social Security and financial account numbers to be cut down to the last four digits. You're a teacher sharing a great piece of student work with next year's class. Or you're just sending a contract to a friend for a second opinion and don't want the other party's details floating around.
None of that is dramatic. All of it goes wrong in exactly the same way when the black box is just a black box.
Before you hit send: the two-minute check
Whatever method you used, run through this list. It takes a couple of minutes, and they're honestly the best two minutes you'll spend on that document.
- Copy and paste everything. Select all, copy, paste into a plain text editor. Read through it. If anything you redacted shows up, stop right there.
- Search for it. Use Ctrl+F (or Cmd+F) and type the exact names, numbers or phrases you removed. Try partial searches too, like the last few digits of an account number.
- Try to select under the box. Drag your cursor across a redacted area. If text highlights, it's still in there.
- Open it in a second viewer. Your browser's built-in PDF viewer is fine for this. Programs display things slightly differently, and every now and then one reveals what another hides.
- Check the document properties. Title, author, subject. Clear out anything that gives the game away.
- Look at the side panels. Comments, bookmarks, attachments, layers. Anything in there you didn't mean to share?
- Zoom way in on the edges. At 400 percent or so, a box that's a few pixels too short can expose the tops of letters or the tail of a "g".
- Rename the file. "Termination_Novak_CONFIDENTIAL.pdf" is a questionable name for a document in which Novak has been redacted.
- Read it like a stranger would. If you've removed a name on page two but page five mentions "the head of the Leeds office", you haven't really removed the name.
One last thought
Bad redactions don't keep happening because people are sloppy. They happen because the screen lies a little. It shows you a black bar, and a black bar looks final. Our brains believe what they see, and the PDF format just happens to be built in a way where "what you see" and "what's actually in the file" are two separate questions.
So don't trust the bar. Test it. Paste the text somewhere, search for the word you removed, give it those two minutes. Worst case, you've lost two minutes of your day. Best case, you're not the next person whose "redacted" document becomes the most interesting thing in the news that week.

