Loại bỏ siêu dữ liệu
Loại bỏ những chi tiết ẩn mà một PDF mang theo về người tạo ra nó.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
This clears the information a PDF carries about itself rather than anything on its pages. That is more than the author field people expect: the document properties, the XMP packet, the private data an editor left under /PieceInfo, per-page metadata, and the identifier in the trailer. The file is read first, so the result can tell you what was actually found and what was removed.
Use it before a document leaves your organisation: a tender response that should not name the person who drafted it, a CV that would otherwise carry your employer's file server path, or a template you are publishing that still credits the client you built it for.
Cách hoạt động
- Drop the PDF onto this page.
- Leave Keep the title off unless the title is deliberate and public - it is often a file path or a working name nobody meant to publish.
- Leave Keep the dates off if you are anonymising. The timestamps are a fingerprint, and clearing them clears the document identifier with them.
- Press Remove metadata.
- Read the report. It lists what was found in the file and what was taken out, so nothing to remove is a real and useful answer.
The gap between what people expect this to remove and what a PDF actually carries is the reason the tool exists. A file routinely records the application that produced it, the operating system account signed in at the time, the full path the source document lived at, every editing session's timestamps, and - in anything that passed through Illustrator or InDesign - a private data blob under /PieceInfo that no PDF reader will ever display to you.
The trailer identifier matters more than its obscurity suggests. It is a pair of hashes derived from the file and the moment it was written, and it survives every cleanup tool that only clears the document properties dialog. Two files carrying the same first identifier are demonstrably versions of one original, which is precisely what someone anonymising a document does not want to leave behind.
Removal here means removal, not blanking. Each entry is deleted from the file and anything only that entry pointed at goes with it, so the values are not left sitting in the bytes for anyone who opens the file in a text editor. When the properties dictionary ends up empty it is dropped entirely, because an empty one is itself a small signal about how the file was made.
This does not touch what is written on the pages. A name in a letterhead, a signature block, a footer with a file path - all of that is content, and content is what Redact PDF is for. Metadata removal makes the file anonymous; it does not make the document anonymous.
The reading and the rewriting both happen in your browser with pdf-lib. There is a pleasing consistency in that: a tool whose job is to stop a document telling strangers about you would be a strange thing to run by sending the document to a stranger.
Công cụ này không làm được gì
- Only the file's own metadata is removed. Names, addresses and paths written into the page content stay exactly where they are - use Redact PDF for those.
- Metadata stored inside embedded images, such as a photograph's EXIF block with its camera and GPS fields, is not read or stripped.
Câu hỏi thường gặp
- Chính xác thì những gì bị loại bỏ?
- Từ điển thông tin tài liệu - tiêu đề, tác giả, chủ đề, từ khóa, ứng dụng tạo, phần mềm xuất và dấu thời gian - cộng thêm gói siêu dữ liệu XMP, các khối /PieceInfo mà trình biên tập để lại ở cấp tài liệu và cấp trang, các luồng siêu dữ liệu theo từng trang, và mã định danh trong trailer. Kết quả liệt kê những mục nào thực sự có mặt trong tệp của bạn.
- Làm sao tôi kiểm chứng nó đã có hiệu quả?
- Hãy mở kết quả trong bất kỳ trình đọc nào và xem thuộc tính tài liệu: các trường đều trống. Để kiểm tra kỹ hơn, hãy chạy exiftool trên tệp, hoặc mở nó trong một trình biên tập văn bản và tìm tên tác giả - các mục bị xóa chứ không phải bị làm trống, nên chẳng còn gì để tìm thấy.
- Việc này có xóa tên tôi khỏi văn bản của tài liệu không?
- Không, và sự phân biệt này rất quan trọng. Việc này xóa những gì tệp nói về chính nó. Một cái tên in trên tiêu đề thư, gõ trong khối chữ ký hay nằm ở chân trang là nội dung trang, và xóa nó nghĩa là biên tập che trang - Che thông tin PDF xóa cả điểm ảnh lẫn văn bản bên dưới cùng lúc.
- Vì sao ngày tháng và mã định danh tài liệu bị xóa cùng nhau?
- Vì chúng là cùng một loại chứng cứ. Thời điểm tạo và sửa đổi cho biết tệp được ghi khi nào; mã định danh trong trailer được suy ra từ tệp và chính khoảnh khắc đó, và nó liên kết các bản sao của một tài liệu với nhau. Giữ lại ngày tháng là giữ lại danh tính của tệp, nên hai thứ này đi cùng nhau thay vì cho bạn cảm giác sai lầm rằng tệp đã sạch.
- Tệp đi về đâu?
- Không đi đâu cả. Nó được pdf-lib đọc và ghi lại trong một Web Worker trên thiết bị của bạn, và không có yêu cầu nào rời khỏi trang. Với một công cụ được dùng riêng để ngăn một tài liệu tiết lộ những điều về bạn, đó là thiết kế hợp lý duy nhất - và bạn có thể xác nhận điều đó trong tab mạng.