Che thông tin PDF
Xóa vĩnh viễn văn bản và ảnh, thay vì che chúng đi.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
Redaction here removes the content rather than covering it. Every page carrying a redaction is rebuilt from pixels: the boxes are drawn into the page, the page is rendered at the resolution you choose, and the render becomes the new page. The original content stream is discarded rather than edited, so there is nothing underneath to select, copy or search. Pages with no redactions are copied through untouched.
Whenever a document has to go out with something taken out of it: names in a court filing, salaries in a board pack, an account number on an invoice, patient details in a medical record, an internal comment in a contract going to the other side.
Cách hoạt động
- Drop the PDF onto this page and draw boxes over whatever has to go. Each box is one redaction.
- Add a phrase under find and redact to catch every occurrence of a name or a number wherever it appears, instead of hunting page by page.
- Set the resolution for the rebuilt pages. 200 dpi reads well on screen and prints acceptably; 300 is better for print and produces a larger file.
- Leave keep the document searchable on unless you have a reason not to, and leave remove metadata on - a document's properties often hold the very name you have just taken off the page.
- Press Redact, then open the result and try to select the redacted text. There will be nothing there.
Drawing a black rectangle over text is not redaction, and the mistake has leaked court records, medical files and government documents repeatedly for twenty years. The rectangle is a drawing instruction added on top; the text underneath is still in the content stream, where anyone can select it, copy it or pull it out with a two-line script. The same is true of cropping.
Rebuilding from pixels is the only approach a browser can take that is genuinely safe, and it is deliberately blunt. The boxes are drawn into the document first and the page is rendered afterwards rather than the other way round, so no rounding in raster space can leave a one-pixel sliver of the original visible along an edge. There is no underneath, because the underneath was thrown away.
Losing search would be a real cost, so the surviving text is put back. The text layer is extracted before anything is drawn, every run falling inside a redaction box is dropped, and the rest is drawn back invisibly over the image. The document stays searchable and the redacted words are genuinely absent - the combination people assume a black box gives them and it almost never does.
The trade-off is size and fidelity on the pages you touched. A vector page that becomes a 200 dpi JPEG is usually several times larger, and links, form fields and annotations on that page do not survive the rebuild - which is why only affected pages are rebuilt. Check the result rather than trusting it: search for the word you removed, then select across the box and paste.
Công cụ này không làm được gì
- Pages containing redactions are replaced by images at the resolution you choose, so they grow, stop scaling cleanly, and lose any links, form fields or annotations they carried.
- Find and redact matches text within a single run, so a phrase broken across a line or a change of font can be missed. Read the result before sending it.
- A scanned document has no text layer, so find and redact has nothing to search. Draw the boxes by hand, or run OCR PDF first.
- Redaction removes what is on the page. Content elsewhere in the file - attachments, bookmark titles, an embedded XMP record - is a separate matter; metadata removal is offered here and Sanitise PDF goes further.
Câu hỏi thường gặp
- Đây là che thông tin thật hay chỉ là một ô đen?
- Thật. Các trang bạn che được kết xuất thành ảnh với các ô đã tô đè lên, và những ảnh đó thay thế các trang, nên các đối tượng văn bản gốc hoàn toàn không có trong tệp xuất ra. Một hình chữ nhật đen vẽ đè lên văn bản để văn bản nằm nguyên bên dưới nó, đó là cách các vụ che thông tin thất bại lên báo.
- Làm sao tôi chắc chắn văn bản thực sự đã biến mất?
- Hãy kiểm tra, bằng hai cách chỉ mất vài giây. Mở tệp xuất ra và tìm một từ đã che: không có kết quả nào khớp cả. Rồi quét chọn qua vùng đã che, sao chép, và dán vào một trình biên tập văn bản: không có gì hiện ra.
- Nó có thể che mọi lần xuất hiện của một cái tên không?
- Có. Hãy gõ cụm từ vào ô tìm và che, và mọi đoạn văn bản khớp đều được khoanh ô, khắp tài liệu hoặc trong một phạm vi trang bạn đặt. Việc khớp không phân biệt chữ hoa chữ thường, và có sẵn biểu thức chính quy cho các mẫu như số tài khoản. Vì việc khớp hoạt động theo từng đoạn, một cụm từ bị chia qua một dấu ngắt dòng có thể bị bỏ sót.
- Vì sao tệp của tôi lớn hơn?
- Vì các trang đã che giờ là ảnh. Một trang văn bản vector thì gọn nhẹ; cùng trang đó dưới dạng JPEG 200 dpi thường lớn gấp vài lần. Chỉ những trang có chứa phần che mới được chuyển đổi, nên tác động tỷ lệ với lượng bạn đã che. Hạ xuống 150 dpi sẽ giúp giảm bớt.
- Nó có dọn cả thuộc tính tài liệu không?
- Có, theo mặc định. Siêu dữ liệu thường chứa tên tác giả, đường dẫn tệp gốc và phần mềm đã tạo ra nó - đôi khi chính cái tên bạn vừa gỡ khỏi trang. Có thể tắt công tắc này đi khi bạn cần giữ nguyên thuộc tính.
- Tệp có được tải lên để che thông tin không?
- Không, và với công cụ này thì đó là toàn bộ lập luận. Một tài liệu bạn đang che thông tin, theo định nghĩa, là tài liệu có thứ gì đó nhạy cảm bên trong, và gửi nó tới một máy chủ để loại bỏ phần nhạy cảm là ngược đời. Kết xuất, che và ghi lại đều diễn ra trong các Web Worker trong trình duyệt của bạn.