Chuyển đến nội dung

Kiểm tra PDF

Một báo cáo về cấu trúc, bảo mật, quyền riêng tư và khả năng đọc của một tài liệu.

Xử lý trên thiết bị của bạn

Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.

Công cụ này làm gì

This reads a document and reports on it rather than changing it. Four passes answer four different questions: is the file structurally sound, is it locked, what does it give away about whoever made it, and can a screen reader read it. Every finding that has a fix names the tool that applies it, so the report is something you can act on instead of a list of things to worry about.

Before sending a document you did not create, or one you did: to see what metadata is riding along, whether there are attachments you had forgotten, whether it is a scan that nobody's screen reader can read, or why another program refuses to open it.

Cách hoạt động

  1. Drop the PDF onto this page. Nothing is written back to it - inspection is read-only.
  2. Leave all four checks on for a full report, or switch off the ones you do not need.
  3. Press Inspect. Three engines are used: qpdf for encryption and structure, pdf-lib for metadata, attachments and links, PDF.js for the text layer.
  4. Read the report. Each finding is marked ok, information, warning or danger, and the ones that can be fixed carry a link to the tool that fixes them.
  5. Keep the report as a web page, or switch the format to JSON under advanced options if a machine is going to read it.

The structural pass runs qpdf's own check, which walks the cross-reference table and the object streams and reports whatever does not add up. A warning here is usually why some other program refuses to open the file, and Repair PDF is usually the answer. A clean check is a stronger statement than it opened on my machine.

The security pass separates the two kinds of protection people confuse. A document may be encrypted so that it cannot be opened without a password, or encrypted with an empty password so that it opens for everyone and merely declares restrictions. The report says which it is, how many bits the handler uses, and lists the permission lines exactly as qpdf reports them. It does not verify signatures - that is Verify signature's job.

The privacy pass is the one that surprises people. A PDF routinely records who created it, which program produced it, and sometimes the full path of the file on that person's disk; it can carry attachments nobody remembers adding, and links to hosts you might not want to be seen visiting. The report names the identifying fields, lists the attachments, and shows only the hosts of external links.

The accessibility pass is the one almost nobody else runs. It checks whether the document has a text layer at all - five pages sampled evenly across the whole document, because a cover sheet is often the only page with real text in an otherwise scanned file - and whether the document has a title and an outline. Without a text layer the file is a picture of a document.

Công cụ này không làm được gì

  • The text-layer check samples five pages spread across the document, so a file with real text on only a page or two can be reported either way.
  • The signature line reports whether form fields exist; it verifies nothing. Use Verify signature for an actual answer.
  • The accessibility pass looks for a text layer, a title and an outline. It is not a PDF/UA conformance test - tags, reading order and alternative text are not examined.
  • A document that needs a password to open can only be reported on from the outside; its metadata, attachments and text layer cannot be read until it has been through Unlock PDF.

Câu hỏi thường gặp

Báo cáo cho tôi biết gì?
Số trang, phiên bản PDF và kích thước tệp; cấu trúc có qua được bài kiểm tra của qpdf không; tệp có được mã hóa không và với quyền hạn nào; nó mang theo siêu dữ liệu định danh, tệp đính kèm và liên kết ngoài nào; và nó có lớp văn bản, tiêu đề và dấu trang không. Mỗi phát hiện được xếp hạng, và những phát hiện có cách khắc phục sẽ liên kết thẳng tới công cụ áp dụng nó.
Một PDF có thể chứa thông tin cá nhân nào?
Nhiều hơn hầu hết mọi người nghĩ. Tên tác giả, chương trình đã tạo ra nó và đôi khi cả đường dẫn tệp đầy đủ trên máy đã tạo nó được lưu dưới dạng thuộc tính tài liệu thông thường và vẫn còn sau khi được gửi email vòng vòng. Các tệp cũng có thể mang theo những tệp đính kèm và liên kết mà chưa ai xem tới.
Làm sao tôi biết một PDF có thể truy cập được không?
Câu hỏi đầu tiên là nó có lớp văn bản không, và báo cáo này trả lời bằng cách lấy mẫu năm trang rải khắp tài liệu. Không có lớp văn bản thì nó là một hình ảnh của một tài liệu: không trình đọc màn hình nào đọc được nó và không tìm kiếm nào tìm thấy gì trong nó, điều mà OCR PDF khắc phục. Báo cáo cũng đánh dấu khi thiếu tiêu đề, và thiếu mục lục trên các tài liệu hơn hai mươi trang.
Nó tìm thấy vấn đề. Giờ tôi làm gì?
Hãy đi theo các liên kết. Một cảnh báo cấu trúc dẫn tới Sửa PDF, thiếu lớp văn bản tới OCR PDF, siêu dữ liệu định danh tới Loại bỏ siêu dữ liệu, tệp đính kèm tới Làm sạch PDF, thiếu tiêu đề tới Chỉnh sửa siêu dữ liệu. Một phát hiện để mặc bạn tự tìm ra cần tìm kiếm gì thì chẳng mấy hữu ích.
Tệp có được tải lên để kiểm tra không?
Không. Việc kiểm tra là ba engine chạy dưới dạng WebAssembly và JavaScript trong các Web Worker trong trình duyệt của bạn, và báo cáo cũng được tạo ra ở đó. Cũng không có gì được ghi ngược lại vào tệp. Sau khi trang đã tải xong một lần, nó hoạt động khi kết nối bị tắt.

Công cụ liên quan