เปรียบเทียบไฟล์ PDF
ดูว่าอะไรเปลี่ยนไประหว่างสองเวอร์ชัน ทั้งในระดับคำและระดับพิกเซล
เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์
เครื่องมือนี้ทำอะไร
This compares two PDFs twice and reports the results separately: a word-level diff of what the documents say, and a page-by-page pixel comparison of what they look like. What you get is an HTML report naming the changes, not a blended similarity score. If the two files turn out to be identical, the report says so in its title rather than leaving you to interpret an empty list.
When a document has come back and you need to know what moved: a contract returned by the other side, a specification revised by a supplier, a design proof against the version you approved, or two copies of a report you are no longer certain are the same file.
วิธีทำงาน
- Drop both PDFs onto this page. The first is treated as the original and the second as the revision, which is what the labels in the report mean.
- Choose text, appearance or both. Both is the default and costs one render of each document.
- For the appearance pass, set the resolution and the sensitivity - sensitivity is how large a pixel difference has to be before it counts as a change rather than as rendering noise.
- Leave ignore whitespace on so that a reflowed paragraph does not report as a change on every line.
- Press Compare and open the report. It lists the added and removed passages and the percentage of pixels that differ on each page.
A lawyer wants the words and a designer wants the pixels, and one blended percentage would answer neither. The text pass diffs the documents word by word, so a changed sentence reports the words that changed instead of a scatter of letters. The appearance pass renders both documents and compares page one with page one, counting the pixels that differ beyond the tolerance you set.
The text comparison treats each document as a single stream of words, so it tells you what changed rather than which page it changed on. That is deliberate - a paragraph inserted on page two would otherwise make every later page report as different - but it means the appearance pass is where you look when you need location.
The appearance pass compares by page number, which is right for a revision and misleading for an insertion: add a page at the front and everything after it differs from its counterpart. Pages present in only one document are flagged as such, and two documents on different paper sizes are compared over the region they share, with the rest counted as different.
Two renders of genuinely identical pages are not bit-identical, because anti-aliasing is not deterministic to the last channel value, so a comparison with no tolerance reports that noise as change. The default sensitivity of 20 per cent ignores small differences. The report also lists at most 200 individual changes and then says how many it has not listed, rather than truncating quietly.
สิ่งที่เครื่องมือนี้ทำไม่ได้
- The report lists at most 200 individual text changes, and says how many more there were. Wholesale rewrites are better compared by exporting the text and diffing it elsewhere.
- Scanned documents have no text layer, so the text pass finds nothing to compare. Run OCR PDF on both first, or compare by appearance only.
- The appearance pass compares page 1 with page 1 and so on, so an inserted or deleted page makes everything after it report as different.
- The report gives counts and percentages for each page, not a marked-up image showing where on the page the difference falls.
คำถามที่คนมักถาม
- การเทียบข้อความกับการเทียบลักษณะที่ปรากฏต่างกันอย่างไร
- การเทียบข้อความจะดึงคำออกจากทั้งสองเอกสารแล้วเทียบกัน ซึ่งจับตัวเลขที่เปลี่ยนไปหรือข้อสัญญาที่ถูกลบได้ ไม่ว่ามันจะย้ายไปอยู่ที่ใด ส่วนการเทียบลักษณะที่ปรากฏจะเรนเดอร์ทั้งสองไฟล์แล้วเทียบพิกเซล ซึ่งจับขอบกระดาษที่ขยับ โลโก้ที่ถูกสลับ หรือฟอนต์ที่ถูกแทนที่ ซึ่งลำพังตัวคำจะไม่มีวันเผยให้เห็น
- แถบเลื่อนความไวทำอะไร
- มันกำหนดว่าพิกเซลสองจุดต้องต่างกันแค่ไหนจึงจะนับเป็นความต่าง การเรนเดอร์ไม่ได้ให้ผลเหมือนเดิมทุกค่าช่องสัญญาณ ค่าที่ต่ำมากจึงรายงานการลบขอบหยักว่าเป็นการเปลี่ยนแปลง ค่าเริ่มต้นที่ 20 เปอร์เซ็นต์เป็นจุดตั้งต้นที่เหมาะสม ลดลงเพื่อจับความเปลี่ยนแปลงเล็กน้อย และเพิ่มขึ้นเมื่อรายงานเต็มไปด้วยความต่างที่ไม่มีความหมาย
- เปรียบเทียบเอกสารสแกนสองไฟล์ได้หรือไม่
- ในแง่ลักษณะที่ปรากฏได้ นั่นคือสิ่งที่รอบการเทียบพิกเซลมีไว้ แม้ว่าการสแกนกระดาษแผ่นเดียวกันสองครั้งจะต่างกันเล็กน้อยไปทั้งหน้าและต้องใช้ความไวที่สูงขึ้น ส่วนในแง่ข้อความนั้นไม่ได้ เว้นแต่ไฟล์สแกนผ่าน OCR มาแล้ว เพราะเมื่อไม่มีชั้นข้อความก็ไม่มีอะไรให้เทียบ
- สุดท้ายแล้วจะได้อะไร
- รายงาน HTML ที่คุณบันทึก สั่งพิมพ์ หรือส่งต่อได้ เริ่มด้วยบทสรุปว่าเหมือนกันหรือไม่ มีการเปลี่ยนข้อความกี่จุด และมีกี่หน้าที่เปลี่ยนไป จากนั้นจึงแสดงรายการข้อความที่เพิ่มและที่ถูกลบ แล้วให้เปอร์เซ็นต์พิกเซลรายหน้า ใครก็เปิดรายงานนี้ได้โดยไม่ต้องใช้เว็บนี้
- เอกสารทั้งสองถูกอัปโหลดหรือไม่
- ไม่ ทั้งสองไฟล์ถูกอ่าน เรนเดอร์ และเปรียบเทียบภายใน Web Worker ในเบราว์เซอร์ของคุณ ซึ่งก็เป็นเหตุผลที่การเทียบไฟล์คู่ละห้าสิบหน้าไม่ขึ้นกับการเชื่อมต่อของคุณ รายงานก็ถูกสร้างขึ้นในเครื่องเช่นกัน