Thêm và chỉnh sửa liên kết
Làm cho các URL trong tài liệu nhấp được, hoặc gỡ các liên kết ra ngoài đi.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
A URL printed in a PDF is usually not a link. It is text that looks like one, because whatever produced the document wrote the characters and never created the annotation that makes them clickable. This tool finds those and makes them real, lets you draw links by hand over anything else, and can strip existing links out. Links are annotations laid over a rectangle, so nothing about the page's appearance changes either way.
The common case is a document exported from a design tool or scanned from print where every web address is dead text. The other direction matters too: stripping outbound links from a document before it goes somewhere untrusted, or before it is published, so nobody follows a tracking address that was fine internally.
Cách hoạt động
- Drop the PDF onto this page.
- Leave automatic detection on for a first pass. Every URL and e-mail address found in the text is listed with the page it sits on.
- Draw a rectangle over anything else that should be clickable and give it a web address, or a page number to jump to inside the document.
- Turn on replacing existing links if the document already carries some and you would rather start clean.
- Press Save links and download the result.
Detection is harder than a regular expression over the page text, which is why most tools that offer it miss half the links. A PDF regularly splits one URL across several text runs, so a naive scan finds "https://exa" and "mple.com/report" and links neither. The runs on each line are joined before matching, which finds the whole address. Trailing punctuation is then trimmed - the full stop that ended the sentence rather than the one inside the domain - and a closing bracket is only dropped when it is unbalanced, because plenty of real addresses end in one and truncating those produces a link that leads nowhere. What is not linked is as deliberate as what is. A bare host with no dot in it is rejected, because in running prose "e.g" and "vs" match far more often than "http://localhost" appears, and a wrong link is worse than a missing one. An address beginning www. gets https:// in front of it, since a bare host is not a URI and readers show the annotation as broken. E-mail addresses become mailto: links, and an address inside a URL is not linked twice.
Links are written without a border and with the print flag set. The first is what stops readers drawing the black box round every link that made documents from the 1990s so recognisable; the second is what keeps the link alive when someone prints the document back to PDF. Internal jumps are written as explicit destinations that preserve the reader's current zoom and land at the top of the target page, rather than as named destinations, which break as soon as anyone splits or merges the file.
Running detection twice on the same document adds a second annotation over the same words unless you turn on replacing existing links first. Two stacked links behave unpredictably - readers pick whichever they find first - so turn that option on for a second pass, and leave it off when the document already carries hand-made links you want to keep.
Stripping outbound links is a separate pass and it ignores everything else on the page: every link with a web address is removed and the internal page-to-page jumps are kept, so a document keeps its own navigation and loses its ability to send a reader anywhere. Sanitise PDF is the wider version of this, removing JavaScript, launch actions and embedded files as well.
Detection needs the text of the document, which is read in your browser by PDF.js, and the annotations are written by pdf-lib in a Web Worker. No address in your document is looked up, resolved or checked against anything - which also means a link to a page that no longer exists is written exactly as printed.
Công cụ này không làm được gì
- Detection can only find what is in the text layer. A scan has no text until it has been through OCR, and an address broken across two lines is read as two fragments, so neither is linked.
Câu hỏi thường gặp
- Vì sao các URL trong PDF của tôi chưa nhấp được sẵn?
- Vì một liên kết trong một PDF là một chú thích - một hình chữ nhật kèm một hành động - chứ không phải một thuộc tính của văn bản. Word và hầu hết các công cụ thiết kế tạo ra chúng khi bạn dán một URL, nhưng các lối xuất làm mất chúng, và một bản quét thì chưa từng có: nó có một ảnh của một địa chỉ web. Phát hiện tự động đọc văn bản, tìm các địa chỉ và ghi các chú thích còn thiếu lên chúng.
- Tôi có thể liên kết từ một trang của tài liệu tới trang khác không?
- Có. Vẽ một hình chữ nhật và cho nó một số trang thay vì một địa chỉ web. Bước nhảy được ghi thành một đích rõ ràng giữ nguyên mức thu phóng hiện tại của trình đọc và đáp xuống đầu trang đích. Với cả một mục lục, Chỉnh sửa dấu trang thường là công cụ tốt hơn - nó cho trình đọc một bảng bên thay vì văn bản nhấp được.
- Làm sao để gỡ các liên kết khỏi một PDF?
- Hai cách, và chúng làm những việc khác nhau. Bật thay thế các liên kết hiện có sẽ xóa mọi liên kết trước khi ghi bất cứ gì bạn đã thêm, nên nó xóa hết chúng khi bạn không thêm gì. Gỡ các liên kết ra ngoài chỉ lấy đi những liên kết trỏ ra web, để nguyên các bước nhảy giữa các trang - đó là điều bạn muốn cho một tài liệu đi ra ngoài tổ chức.
- Một số URL của tôi chỉ được phát hiện một phần. Chuyện gì đang xảy ra?
- Hầu như luôn là một dấu ngắt dòng. Bộ phát hiện nối các đoạn văn bản trong một dòng trước khi khớp, nên một địa chỉ tách thành nhiều đoạn trên một dòng được tìm thấy nguyên vẹn - nhưng một địa chỉ xuống dòng tiếp theo đến dưới dạng hai mảnh và không mảnh nào trông giống một URL khi đứng riêng. Hãy vẽ những cái đó bằng tay: một hình chữ nhật trên mỗi dòng, cả hai mang cùng một địa chỉ.
- Các liên kết có được kiểm tra hay gửi đi đâu không?
- Không cái nào. Không địa chỉ nào được phân giải, tải về hay xác thực, và tài liệu không bao giờ được tải lên - việc trích xuất văn bản và ghi chú thích đều diễn ra trong trình duyệt của bạn. Điều đó có nghĩa là một lỗi gõ trong một địa chỉ in ra sẽ trở thành một liên kết với cùng lỗi gõ đó, nên đáng để phát hiện những cái rõ ràng hỏng trong danh sách trước khi lưu.