Chuyển đến nội dung

So sánh PDF

Xem những gì đã thay đổi giữa hai phiên bản, bằng chữ và bằng điểm ảnh.

Xử lý trên thiết bị của bạn

Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.

Công cụ này làm gì

This compares two PDFs twice and reports the results separately: a word-level diff of what the documents say, and a page-by-page pixel comparison of what they look like. What you get is an HTML report naming the changes, not a blended similarity score. If the two files turn out to be identical, the report says so in its title rather than leaving you to interpret an empty list.

When a document has come back and you need to know what moved: a contract returned by the other side, a specification revised by a supplier, a design proof against the version you approved, or two copies of a report you are no longer certain are the same file.

Cách hoạt động

  1. Drop both PDFs onto this page. The first is treated as the original and the second as the revision, which is what the labels in the report mean.
  2. Choose text, appearance or both. Both is the default and costs one render of each document.
  3. For the appearance pass, set the resolution and the sensitivity - sensitivity is how large a pixel difference has to be before it counts as a change rather than as rendering noise.
  4. Leave ignore whitespace on so that a reflowed paragraph does not report as a change on every line.
  5. Press Compare and open the report. It lists the added and removed passages and the percentage of pixels that differ on each page.

A lawyer wants the words and a designer wants the pixels, and one blended percentage would answer neither. The text pass diffs the documents word by word, so a changed sentence reports the words that changed instead of a scatter of letters. The appearance pass renders both documents and compares page one with page one, counting the pixels that differ beyond the tolerance you set.

The text comparison treats each document as a single stream of words, so it tells you what changed rather than which page it changed on. That is deliberate - a paragraph inserted on page two would otherwise make every later page report as different - but it means the appearance pass is where you look when you need location.

The appearance pass compares by page number, which is right for a revision and misleading for an insertion: add a page at the front and everything after it differs from its counterpart. Pages present in only one document are flagged as such, and two documents on different paper sizes are compared over the region they share, with the rest counted as different.

Two renders of genuinely identical pages are not bit-identical, because anti-aliasing is not deterministic to the last channel value, so a comparison with no tolerance reports that noise as change. The default sensitivity of 20 per cent ignores small differences. The report also lists at most 200 individual changes and then says how many it has not listed, rather than truncating quietly.

Công cụ này không làm được gì

  • The report lists at most 200 individual text changes, and says how many more there were. Wholesale rewrites are better compared by exporting the text and diffing it elsewhere.
  • Scanned documents have no text layer, so the text pass finds nothing to compare. Run OCR PDF on both first, or compare by appearance only.
  • The appearance pass compares page 1 with page 1 and so on, so an inserted or deleted page makes everything after it report as different.
  • The report gives counts and percentages for each page, not a marked-up image showing where on the page the difference falls.

Câu hỏi thường gặp

Khác biệt giữa văn bản và hình thức là gì?
Văn bản trích các từ từ cả hai tài liệu và so sánh chúng, cách này bắt được một con số bị đổi hay một điều khoản bị xóa dù nó chuyển đi đâu. Hình thức kết xuất cả hai và so sánh các điểm ảnh, cách này bắt được một lề bị dịch, một logo bị tráo hay một phông chữ bị thay thế mà chỉ riêng các từ sẽ không bao giờ để lộ.
Thanh trượt độ nhạy làm gì?
Nó đặt hai điểm ảnh phải khác nhau đến mức nào thì khác biệt mới được tính. Việc kết xuất không mang tính tất định đến giá trị kênh cuối cùng, nên một thiết lập rất thấp sẽ báo cả khử răng cưa là thay đổi. Mặc định 20 phần trăm là một điểm khởi đầu hợp lý: hạ nó xuống để bắt các dịch chuyển tinh tế, nâng nó lên khi báo cáo ngập trong những khác biệt chẳng có ý nghĩa gì.
Tôi có thể so sánh hai tài liệu quét không?
Bằng hình thức thì được - đó là mục đích của lần so sánh điểm ảnh, dù hai lần quét riêng biệt của cùng một tờ giấy sẽ khác nhau đôi chút ở khắp nơi và cần một độ nhạy cao hơn. Bằng văn bản thì không, trừ khi các bản quét đã qua OCR: không có lớp văn bản thì chẳng có gì để so sánh.
Cuối cùng tôi nhận được gì?
Một báo cáo HTML bạn có thể lưu, in hoặc gửi đi. Nó mở đầu bằng một bản tóm tắt - giống nhau hay không, bao nhiêu thay đổi văn bản, bao nhiêu trang thay đổi - rồi liệt kê các đoạn được thêm và bị xóa, rồi đưa ra phần trăm điểm ảnh theo từng trang. Bất kỳ ai cũng có thể mở nó mà không cần trang web này.
Cả hai tài liệu có được tải lên không?
Không. Cả hai đều được đọc, kết xuất và so sánh bên trong một Web Worker trong trình duyệt của bạn, đó cũng là lý do việc so sánh một cặp năm mươi trang không phụ thuộc vào kết nối của bạn. Báo cáo cũng được tạo ra ngay tại chỗ.

Công cụ liên quan