ข้ามไปยังเนื้อหา

ลบเมทาดาทา

ลบรายละเอียดที่ซ่อนอยู่ซึ่ง PDF พกติดไว้ว่าใครเป็นคนสร้าง

ประมวลผลบนเครื่องของคุณ

เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์

เครื่องมือนี้ทำอะไร

This clears the information a PDF carries about itself rather than anything on its pages. That is more than the author field people expect: the document properties, the XMP packet, the private data an editor left under /PieceInfo, per-page metadata, and the identifier in the trailer. The file is read first, so the result can tell you what was actually found and what was removed.

Use it before a document leaves your organisation: a tender response that should not name the person who drafted it, a CV that would otherwise carry your employer's file server path, or a template you are publishing that still credits the client you built it for.

วิธีทำงาน

  1. Drop the PDF onto this page.
  2. Leave Keep the title off unless the title is deliberate and public - it is often a file path or a working name nobody meant to publish.
  3. Leave Keep the dates off if you are anonymising. The timestamps are a fingerprint, and clearing them clears the document identifier with them.
  4. Press Remove metadata.
  5. Read the report. It lists what was found in the file and what was taken out, so nothing to remove is a real and useful answer.

The gap between what people expect this to remove and what a PDF actually carries is the reason the tool exists. A file routinely records the application that produced it, the operating system account signed in at the time, the full path the source document lived at, every editing session's timestamps, and - in anything that passed through Illustrator or InDesign - a private data blob under /PieceInfo that no PDF reader will ever display to you.

The trailer identifier matters more than its obscurity suggests. It is a pair of hashes derived from the file and the moment it was written, and it survives every cleanup tool that only clears the document properties dialog. Two files carrying the same first identifier are demonstrably versions of one original, which is precisely what someone anonymising a document does not want to leave behind.

Removal here means removal, not blanking. Each entry is deleted from the file and anything only that entry pointed at goes with it, so the values are not left sitting in the bytes for anyone who opens the file in a text editor. When the properties dictionary ends up empty it is dropped entirely, because an empty one is itself a small signal about how the file was made.

This does not touch what is written on the pages. A name in a letterhead, a signature block, a footer with a file path - all of that is content, and content is what Redact PDF is for. Metadata removal makes the file anonymous; it does not make the document anonymous.

The reading and the rewriting both happen in your browser with pdf-lib. There is a pleasing consistency in that: a tool whose job is to stop a document telling strangers about you would be a strange thing to run by sending the document to a stranger.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • Only the file's own metadata is removed. Names, addresses and paths written into the page content stay exactly where they are - use Redact PDF for those.
  • Metadata stored inside embedded images, such as a photograph's EXIF block with its camera and GPS fields, is not read or stripped.

คำถามที่คนมักถาม

มีอะไรถูกลบออกไปบ้าง
พจนานุกรมข้อมูลเอกสาร ได้แก่ ชื่อเรื่อง ผู้เขียน หัวเรื่อง คำสำคัญ ผู้สร้าง ผู้ผลิต และเวลาประทับ รวมถึงแพ็กเก็ตเมทาดาทา XMP ก้อนข้อมูล /PieceInfo ที่โปรแกรมแก้ไขทิ้งไว้ทั้งระดับเอกสารและระดับหน้า สตรีมเมทาดาทารายหน้า และตัวระบุในส่วนท้ายไฟล์ ผลลัพธ์จะแสดงรายการว่าสิ่งใดมีอยู่จริงในไฟล์ของคุณ
จะตรวจสอบได้อย่างไรว่าได้ผล
เปิดผลลัพธ์ในโปรแกรมอ่านใดก็ได้แล้วดูคุณสมบัติเอกสาร ช่องต่างๆ จะว่างเปล่า หากต้องการตรวจที่เข้มงวดกว่านั้น ให้รัน exiftool กับไฟล์ หรือเปิดไฟล์ในโปรแกรมแก้ไขข้อความแล้วค้นหาชื่อผู้เขียน รายการเหล่านั้นถูกลบทิ้งไม่ใช่แค่ทำให้ว่าง จึงไม่เหลืออะไรให้ค้นเจอ
การทำแบบนี้ลบชื่อฉันออกจากเนื้อหาของเอกสารด้วยหรือไม่
ไม่ และความต่างนี้สำคัญ เครื่องมือนี้ล้างสิ่งที่ไฟล์บอกเกี่ยวกับตัวมันเอง ชื่อที่พิมพ์อยู่ในหัวจดหมาย ในบล็อกลายเซ็น หรือในท้ายกระดาษ คือเนื้อหาของหน้า การลบสิ่งเหล่านั้นคือการปกปิดข้อมูลในหน้า ซึ่งเครื่องมือปกปิดข้อมูลใน PDF จะลบทั้งพิกเซลและข้อความที่อยู่ข้างใต้ไปพร้อมกัน
ทำไมวันที่กับตัวระบุเอกสารจึงถูกล้างไปพร้อมกัน
เพราะทั้งสองอย่างเป็นหลักฐานชนิดเดียวกัน เวลาที่สร้างและเวลาที่แก้ไขบอกว่าไฟล์ถูกเขียนเมื่อใด ส่วนตัวระบุในส่วนท้ายไฟล์ได้มาจากตัวไฟล์และช่วงเวลาเดียวกันนั้น และมันเชื่อมโยงสำเนาของเอกสารเดียวกันเข้าด้วยกัน การเก็บวันที่ไว้จึงเท่ากับเก็บตัวตนของไฟล์ไว้ ทั้งสองอย่างจึงถูกจัดการไปพร้อมกัน แทนที่จะทำให้คุณเข้าใจผิดว่าไฟล์สะอาดแล้ว
ไฟล์ถูกส่งไปที่ใด
ไม่ไปที่ใดเลย ไฟล์ถูกอ่านและเขียนใหม่ด้วย pdf-lib ใน Web Worker บนอุปกรณ์ของคุณ และไม่มีคำขอใดออกจากหน้านี้ สำหรับเครื่องมือที่ใช้เพื่อไม่ให้เอกสารเปิดเผยเรื่องของคุณโดยเฉพาะ นี่เป็นการออกแบบเดียวที่สมเหตุสมผล และคุณยืนยันได้ในแท็บเครือข่าย

เครื่องมือที่เกี่ยวข้อง