ข้ามไปยังเนื้อหา

ปกปิดข้อมูลใน PDF

ลบข้อความและภาพออกอย่างถาวร แทนที่จะแค่บังทับไว้

ประมวลผลบนเครื่องของคุณ

เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์

เครื่องมือนี้ทำอะไร

Redaction here removes the content rather than covering it. Every page carrying a redaction is rebuilt from pixels: the boxes are drawn into the page, the page is rendered at the resolution you choose, and the render becomes the new page. The original content stream is discarded rather than edited, so there is nothing underneath to select, copy or search. Pages with no redactions are copied through untouched.

Whenever a document has to go out with something taken out of it: names in a court filing, salaries in a board pack, an account number on an invoice, patient details in a medical record, an internal comment in a contract going to the other side.

วิธีทำงาน

  1. Drop the PDF onto this page and draw boxes over whatever has to go. Each box is one redaction.
  2. Add a phrase under find and redact to catch every occurrence of a name or a number wherever it appears, instead of hunting page by page.
  3. Set the resolution for the rebuilt pages. 200 dpi reads well on screen and prints acceptably; 300 is better for print and produces a larger file.
  4. Leave keep the document searchable on unless you have a reason not to, and leave remove metadata on - a document's properties often hold the very name you have just taken off the page.
  5. Press Redact, then open the result and try to select the redacted text. There will be nothing there.

Drawing a black rectangle over text is not redaction, and the mistake has leaked court records, medical files and government documents repeatedly for twenty years. The rectangle is a drawing instruction added on top; the text underneath is still in the content stream, where anyone can select it, copy it or pull it out with a two-line script. The same is true of cropping.

Rebuilding from pixels is the only approach a browser can take that is genuinely safe, and it is deliberately blunt. The boxes are drawn into the document first and the page is rendered afterwards rather than the other way round, so no rounding in raster space can leave a one-pixel sliver of the original visible along an edge. There is no underneath, because the underneath was thrown away.

Losing search would be a real cost, so the surviving text is put back. The text layer is extracted before anything is drawn, every run falling inside a redaction box is dropped, and the rest is drawn back invisibly over the image. The document stays searchable and the redacted words are genuinely absent - the combination people assume a black box gives them and it almost never does.

The trade-off is size and fidelity on the pages you touched. A vector page that becomes a 200 dpi JPEG is usually several times larger, and links, form fields and annotations on that page do not survive the rebuild - which is why only affected pages are rebuilt. Check the result rather than trusting it: search for the word you removed, then select across the box and paste.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • Pages containing redactions are replaced by images at the resolution you choose, so they grow, stop scaling cleanly, and lose any links, form fields or annotations they carried.
  • Find and redact matches text within a single run, so a phrase broken across a line or a change of font can be missed. Read the result before sending it.
  • A scanned document has no text layer, so find and redact has nothing to search. Draw the boxes by hand, or run OCR PDF first.
  • Redaction removes what is on the page. Content elsewhere in the file - attachments, bookmark titles, an embedded XMP record - is a separate matter; metadata removal is offered here and Sanitise PDF goes further.

คำถามที่คนมักถาม

นี่คือการปกปิดข้อมูลจริง หรือแค่แถบดำ
จริง หน้าที่คุณปกปิดจะถูกเรนเดอร์เป็นภาพพร้อมกล่องที่ระบายทับไว้ แล้วภาพเหล่านั้นเข้าแทนที่หน้าเดิม อ็อบเจกต์ข้อความต้นฉบับจึงไม่อยู่ในไฟล์ผลลัพธ์เลย ส่วนสี่เหลี่ยมสีดำที่วาดทับข้อความจะทิ้งข้อความไว้ข้างใต้ ซึ่งเป็นเหตุที่ความล้มเหลวในการปกปิดข้อมูลกลายเป็นข่าวอยู่เรื่อยไป
จะมั่นใจได้อย่างไรว่าข้อความหายไปจริง
ตรวจได้สองวิธีที่ใช้เวลาไม่กี่วินาที เปิดไฟล์ผลลัพธ์แล้วค้นหาคำที่ปกปิดไว้ ต้องไม่พบผลลัพธ์ใด จากนั้นลากเลือกทับบริเวณที่ปกปิด คัดลอก แล้ววางลงในโปรแกรมแก้ไขข้อความ ต้องไม่มีอะไรปรากฏขึ้นมา
ปกปิดชื่อได้ทุกครั้งที่ปรากฏหรือไม่
ได้ พิมพ์วลีลงในช่องค้นหาและปกปิด แล้วทุกช่วงข้อความที่ตรงกันจะถูกครอบด้วยกล่อง ทั้งเอกสารหรือภายในช่วงหน้าที่คุณกำหนด การจับคู่ไม่แยกตัวพิมพ์ใหญ่เล็ก และมีนิพจน์ปกติให้ใช้สำหรับรูปแบบอย่างเลขที่บัญชี เนื่องจากการจับคู่ทำทีละช่วงข้อความ วลีที่ถูกตัดข้ามบรรทัดจึงอาจหลุดไปได้
ทำไมไฟล์ของฉันใหญ่ขึ้น
เพราะหน้าที่ปกปิดข้อมูลกลายเป็นภาพไปแล้ว หน้าที่เป็นข้อความเวกเตอร์มีขนาดกะทัดรัด ส่วนหน้าเดียวกันในรูปแบบ JPEG ที่ 200 dpi มักใหญ่กว่าหลายเท่า เฉพาะหน้าที่มีการปกปิดเท่านั้นที่ถูกแปลง ผลจึงแปรผันตามปริมาณที่คุณปกปิด การลดลงมาที่ 150 dpi ช่วยได้
เครื่องมือนี้ล้างคุณสมบัติของเอกสารด้วยหรือไม่
ล้างโดยค่าเริ่มต้น เมทาดาทามักเก็บชื่อผู้เขียน เส้นทางไฟล์เดิม และซอฟต์แวร์ที่ผลิตไฟล์ไว้เป็นประจำ บางครั้งก็เป็นชื่อเดียวกับที่คุณเพิ่งลบออกจากหน้าเอกสาร คุณปิดสวิตช์นี้ได้เมื่อต้องการให้คุณสมบัติยังอยู่ครบ
ไฟล์ถูกอัปโหลดเพื่อปกปิดข้อมูลหรือไม่
ไม่ และสำหรับเครื่องมือนี้ นั่นคือเหตุผลทั้งหมด เอกสารที่คุณกำลังปกปิดข้อมูลย่อมมีบางอย่างที่อ่อนไหวอยู่ข้างในตามนิยาม การส่งไปยังเซิร์ฟเวอร์เพื่อให้ลบส่วนที่อ่อนไหวออกจึงเป็นเรื่องกลับหัวกลับหาง การเรนเดอร์ การบัง และการเขียนใหม่ ทั้งหมดเกิดขึ้นใน Web Worker ในเบราว์เซอร์ของคุณ

เครื่องมือที่เกี่ยวข้อง