Chỉnh sửa dấu trang
Dựng mục lục mà một tài liệu dài lẽ ra phải có.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
Bookmarks are the panel down the side of a reader that lists a document's sections and jumps to them. A long PDF without one is a scroll bar and nothing else. This tool writes that outline: you build the tree by hand, or generate it from the headings the document already has, and each entry is written as a jump to a page in the file itself.
Use it on the documents that need an outline most and rarely have one: a scanned report, a bundle merged from a dozen sources, a manual exported by a tool that did not bother, a court filing where someone has to find exhibit 14 quickly. Renaming an existing entry that reads "Heading 2.1.4" is the other common reason people arrive.
Cách hoạt động
- Drop the PDF onto this page. Any outline it already has is read and shown as a tree.
- Add entries, retitle them, and drag them to change their level - a bookmark nested under another becomes a sub-entry in the reader's panel.
- Set the page each entry points at.
- Or turn on generation from headings and let the document's own heading structure become the tree, then tidy what it produced.
- Press Save bookmarks and the download starts.
Generating from headings uses the same size-and-position analysis that PDF to Markdown uses. There is no heading tag in a normal PDF - a heading is simply text that is bigger, bolder or more isolated than the text around it - so the structure is inferred from the page rather than read from it. A document that converts to Markdown with sensible headings gets sensible bookmarks from the same reading, and one that does not will need tidying afterwards.
Each entry is written as an explicit destination with the view position left unset, which tells a reader to jump to the page and keep the zoom the reader is already at. Naming a page beyond the end of the document is clamped to the last page rather than written as given: readers handle an out-of-range destination inconsistently and a few refuse to open the file at all.
Nesting is the part that makes an outline useful rather than merely long. A branch with children is written open, so a reader shows the chapters expanded on first opening rather than a row of collapsed triangles. Titles are stored as Unicode, so a heading with an accent, a dash or a non-Latin script comes through as written.
The tree replaces the document's outline rather than merging into it, so what you see in the editor before you save is exactly what the file will have. There is no way to save an empty outline here - the tool needs at least one entry - so this is not the way to strip bookmarks out of a document.
The reading, the analysis and the writing all happen in your browser. Generating an outline for a 400-page scan takes a while because every page's text has to be examined, and the progress bar counts pages rather than guessing; none of those pages leaves the device.
Công cụ này không làm được gì
- Generated bookmarks are inferred from type size and position, not read from tags. A document with decorative large text gets entries that are not chapters, one with a flat visual hierarchy gets a flat tree, and a scan with no text layer produces nothing at all - run OCR first, then come back.
Câu hỏi thường gặp
- Dấu trang PDF chính xác là gì?
- Chúng là mục lục của chính tài liệu, lưu dưới dạng một cây các mục mà mỗi mục trỏ tới một trang. Các trình đọc hiển thị chúng trong một bảng bên - Acrobat gọi là Bookmarks, Preview gọi là Table of Contents, Chrome hiển thị nó dưới dạng dàn bài tài liệu. Chúng tách biệt với một trang mục lục in ra, vốn chỉ là văn bản, và tách biệt với các dấu trang mà trình duyệt của bạn lưu.
- Nó có dựng mục lục cho tôi được không?
- Có, từ các tiêu đề của tài liệu. Văn bản của mỗi trang được xem xét và những đoạn trông giống tiêu đề - cỡ chữ lớn hơn, đậm hơn, nhiều khoảng trống quanh chúng - trở thành các mục lồng theo cấp bề ngoài của chúng. Nó hoạt động tốt trên các báo cáo và sổ tay tạo ra từ một trình xử lý văn bản, và kém hơn trên các tài liệu có thứ bậc trực quan phẳng. Hãy coi kết quả là một bản nháp đầu tiên và dọn dẹp nó.
- Làm sao để tạo dấu trang con?
- Lồng một mục dưới một mục khác trong cây và nó trở thành một mục con. Không có giới hạn độ sâu nào đáng lo - các trình đọc xử lý vài cấp một cách thoải mái, và ba cấp là khoảng mà sự hữu ích đạt đỉnh. Các nhánh có con được ghi ở trạng thái mở, nên trình đọc hiển thị chúng đã mở rộng khi tài liệu được mở lần đầu.
- PDF gộp của tôi mất dấu trang. Đây có phải cách khắc phục không?
- Đó là một cách, và cách tốt hơn là gộp bằng Gộp PDF, vốn dựng lại mục lục của mỗi nguồn và trỏ lại mỗi mục về vị trí trang mới của nó. Nếu thiệt hại đã xảy ra rồi, hoặc nếu các nguồn chưa từng có mục lục, hãy dựng một cái ở đây - tạo từ các tiêu đề thường nhanh hơn gõ tay các chương của một tập tài liệu.
- Tài liệu có được tải lên để phân tích không?
- Không. Việc phân tích tiêu đề chạy trên máy của bạn, trong một Web Worker, đó là lý do một tài liệu dài mất thời gian thấy được thay vì tức thì - công việc đang diễn ra ở đây chứ không phải trên một máy chủ có nhiều lõi hơn. Không có gì về tài liệu, kể cả các tiêu đề chương bạn gõ, được truyền đi.