Public Document Redaction Assistant

Is that blacked-out document really safe?
A black rectangle may not remove the text inside a PDF

A name that looks blacked out may still be readable by copying and pasting. The original article describes recurring Japanese disclosure incidents from 2022 onward. We explain how this can happen and how to reduce the risk, without specialist jargon.

Published: Updated:

“It is blacked out, so no one can read it.” That assumption can fail. The original article describes documents published in 2025 where names remained copyable beneath the black masking, with affected files left available for a substantial period. The problem lies in the file’s contents, not simply its appearance.

The essential distinction

Covering text and removing it are different operations. Putting a black rectangle over a page can leave the original text inside the file. A person may not see it on screen, but software may still be able to read it.

EXPERIMENT

Try it for yourself

The example below looks like a redacted document. Select “Select and read the text” to reveal the information available to software. One click demonstrates the same underlying failure mode.

Reference numberDisclosure request 2026-0142
Applicant nameTaro Yamada
Applicant address1-2-3 Shibaura, Minato-ku, Tokyo, Seavans S Building, top floor
Phone number03-0000-0000
DepartmentGeneral Affairs, Information Policy Division
The text is still readable. The black rectangles are only shapes placed on top; the underlying characters remain in the file. The same can happen in a real PDF when someone selects and copies text or runs a text-extraction tool.

This is a fictional illustration, unrelated to actual people or addresses.

CASES

Examples described in the original article

The original Japanese article includes the following anonymized incident examples. They illustrate recurring failure modes rather than establishing findings about identified organizations; the page does not supply individual source links for these cases. The issue can affect ordinary workflows, not only unusually careless teams.

2022
Example AThe original article describes minutes published as a PDF with customer information blacked out. Names and other personal information reportedly remained recoverable on multiple occasions.The reported issue was not recognized, and checks were inadequate
2023
Example BThe original article describes a document supplied externally whose covered text could still be extracted using PDF-editing features. It reports exposure affecting more than 200 staff and other individuals; the incident is anonymized here.A document was visually masked in an editor and then converted to PDF
2024
Example CThe original article describes a PDF released in response to an information request in which personal information intended to be withheld remained accessible.Masking in a document editor was followed by PDF conversion
2025
Example DThe original article describes published application-related documents in which personal information about multiple applicants could be recovered.Underlying text had not been removed
2025
Example EThe original article describes published material where supposedly blacked-out names could still be read by copying the text.Black coverage did not prevent copying the content
2026
Example FThe original article describes an Excel workbook published online with personal information in hidden areas that became accessible through search results.Hidden sheets or cells remained in the workbook
2026
Example GThe original article describes a published PDF in which masked addresses remained accessible.Data beneath the masking remained
The common pattern is using a tool designed to hide content visually as though it removed the content itself. The lesson concerns choosing and verifying the right operation, not merely telling staff to be more careful.
WHY

Why it happens: three misconceptions

Misconception 01

A PDF cannot be edited, so converting to PDF makes it safe

PDF is often less convenient to edit, but it does not automatically turn text into an image. When Word or Excel files are exported to PDF, their text usually remains text. That is why it can be searched and copied.

The correction: converting to PDF does not remove text. Preserving usable text is ordinarily one of PDF’s advantages.

Misconception 02

A black rectangle makes the information invisible, so it is safe

A filled shape, highlighting or a background color may simply sit on top of the content. The text underneath is not necessarily removed. Selecting and copying it, or using a text-extraction tool, can reveal it, as the demonstration above illustrates.

The correction: remove the sensitive content itself, rather than merely covering it.

Misconception 03

A second visual check will catch every problem

Visual review can identify areas someone forgot to cover, but cannot reliably detect text that remains beneath an apparently correct black rectangle. Adding more reviewers does not solve that limitation if they all inspect only the appearance.

The correction: inspect the exported file’s recoverable content as well as its appearance.

SELF CHECK

Self-check: are your workplace’s redactions safe?

Select the statements that apply. This self-check does not send your selections to a server.

Eight questions about public-document redaction

Selected: 0 / 8
Select the statements that apply
Check items to see an indicative risk level and suggested next steps.
HOW TO

Three steps for effective redaction

  • Use a dedicated redaction feature. Apply an operation that removes sensitive page content, rather than adding a shape over it. A PDF editor’s Redact function is distinct from its drawing tools. Apply and save the redactions, and use the tool’s sanitization features where appropriate.
  • If using image output, create a new PDF from the correctly redacted page images. This can avoid carrying the original text layer into the output. The concealed pixels must actually be replaced, and the new file must not retain the original pages, attachments or an OCR layer containing the sensitive text. Image conversion alone is not a universal guarantee.
  • Check the exported file for recoverable information. Before publication, use a text-extraction tool or copy all text into a text editor. If a supposedly removed name appears, the redaction has failed. A blank result is useful evidence, but not sufficient by itself: also inspect the rendered pages and hidden content such as metadata, layers and attachments. Make these checks part of the publication procedure.
The first two steps concern how the output is produced; the third concerns how it is verified. Verification of the actual exported file is easy to omit and crucial to preventing disclosure.
REALITY

Even a correct manual process has limits

Knowing the method does not solve the volume problem. The original article cites Japan’s FY2024 administrative information-disclosure report: 215,425 requests, up 9,765 from the previous year. It reports that 96.6% arrived at a counter or by post, reflecting workflows still heavily based on paper and PDFs. The cited Japanese statistical report is linked below.

215,425

FY2024 disclosure requests in Japan, as cited in the original article; up 9,765 year on year

96.6%

Submitted at a counter or by post; the cited online share is 3.4%

91.5%

Decisions within the statutory 30-day period; the cited report lists 548 cases taking over one year

2,381

Administrative review requests; the cited figure includes 1,041 challenges to non-disclosure decisions

Staff review every page, find names, addresses, contact details and dates, redact them, and check the result. An omission can expose personal information, making the task demanding and time-consuming. They also need to explain afterward why an item was withheld. At this scale, “be more careful” is not a complete solution.

SOLUTION

Why we support the work without sending documents outside

AI services can help identify redaction candidates, and many process documents in the cloud. Japan’s Digital Agency guidelines revised in June 2026 state that, in principle, information requiring confidentiality cannot be handled by cloud generative AI used merely by accepting standard terms. The original article also cites the Ministry of Internal Affairs and Communications’ local-government security guidance. Applicable procurement arrangements, classifications and organizational rules must be checked.

A redaction workflow necessarily sees the very information that must not be disclosed. Sending it outside for efficiency therefore requires a deliberate security and governance decision.

Public Document Redaction Assistant

Generative AI prepares a redaction draft for review

This redaction-support application runs on Sovereign GaiXer. From OCR and AI candidate assessment to PDF generation, the local workflow runs within the compact appliance. In its closed-network configuration, the document is not sent to an external AI service.

Japanese-language source screenshots comparing highlighted redaction candidates with the exported, blacked-out PDFThe Sovereign GaiXer appliance
0 charactersCharacters extracted from the demonstrated output PDF

The application rebuilds pages as images so the original text layer is not carried into the output

AI proposes candidates; staff make the final decision on what to remove
Each candidate shows a reason, such as a personal name, together with OCR confidence information
Staff can drag over the preview to add areas that AI did not identify
Operations are automatically recorded in annual CSV logs, including the user, time and action
Supports PDF, Word and Excel; Office files are converted to PDF internally for processing

This application is under development. A working demonstration exists, and remaining work before production use is being assessed. Contact us to discuss a demonstration and the intended deployment.

COMPARISON

Comparing three approaches

ItemManual workCloud AIAI within your appliance
Work timeVisual review of every page; effort grows with volumeCandidate suggestions can reduce effortCandidate suggestions can reduce effort
Document destinationRemains within the organizationSent to an external providerRemains on the organization’s appliance
Confidential informationSubject to internal controlsSubject to applicable guidelines and service arrangementsSubject to internal controls and configuration
Processing recordsMay require a separate manual processDepends on the serviceAutomatic operation logs
Cost modelStaff timeSubscription or usage chargesAppliance purchase; local usage has no metered charge, with operating and application terms to confirm
FAQ

Frequently asked questions

Can the AI miss personal information?

Yes. That is why this application does not leave the final deletion decision to AI. AI identifies candidates; a staff member decides on screen what to redact. The initial policy includes names and dates among the candidates and flags uncertain items to reduce missed information.

If AI reads the document, does the information leave our organization?

The demonstrated application performs OCR, AI assessment and PDF generation within the Sovereign GaiXer appliance. After required initial setup, this local workflow can run in a closed network without continuous internet access.

How can we check that the text has actually been removed?

Run a text-extraction tool against the exported PDF, or select all, copy and paste into a text editor. This can reveal hidden text. Our measured example yielded zero extracted characters. A blank result from one tool is not a complete security proof: also inspect attachments, metadata, layers and the visible image, and confirm that no OCR text was added back.

I am concerned about documents we have already published online.

Start by copying all text from the published file into a text editor and checking the output. If information intended to be withheld appears, promptly remove public access and follow your incident-response process. Assess search-result or cached-copy removal where available, retain evidence and determine any notification obligations. Replacing the file alone may not remove copies already obtained.

Where should we start if we are considering this application?

A demonstration on the actual device is a good starting point. Bring an appropriate sample resembling your real document format so you can inspect the suggested redactions and the exported PDF.

SUMMARY

Three points to remember

  • Covering text and removing it are different operations. Adding a shape can leave the underlying text in the file.
  • Check the actual exported content, as well as its appearance. Include copy-and-paste or text extraction in the pre-publication checks, alongside review of hidden content and rendered pages.
  • AI can help with volume, but it handles the information you intend to withhold. Where processing takes place deserves as much attention as the application’s features.