PDF sang Markdown
Biến PDF thành Markdown, khôi phục tiêu đề, danh sách và bảng.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
This reads a PDF's text and writes it out as Markdown, working out which lines are headings, where paragraphs begin and end, and which blocks are lists or tables. A PDF stores none of that - only characters at coordinates - so the structure is inferred from geometry, except where the document has a bookmark outline, which is used instead because it is the author's own account of the structure. The result is one .md file that re-flows; it does not reproduce the original pagination or layout.
Reach for it when a document has to become editable text again: a report going into a wiki, documentation that arrived as a PDF and now needs to live in a repository beside the code, or notes you would rather search and rewrite than scroll through.
Cách hoạt động
- Drop the PDF onto this page. The first five pages are checked for a text layer, so a scan is caught before you download an empty file.
- Leave Use the bookmark outline on when the document has one. A real outline is more reliable than any guess made from type size.
- Choose what else to detect: headings where no bookmark covers them, lists, and tables written as Markdown pipe tables.
- Under advanced options, add a front matter block if the file is headed for a static site, and set a page range if you only want part of the document.
- Press Convert to Markdown. The .md file downloads on its own.
A line becomes a heading when it is set larger than the body text around it, is short, and is followed by prose - three conditions rather than one, because any of them alone promotes running heads and figure captions by mistake. Bookmarks override the lot, which is why a document with a proper outline converts noticeably better than one without.
Paragraphs are rebuilt by asking why each line ended. A line that runs to the right edge of its column was broken by the typesetter, so the next line is joined onto it; a line that stops short ended on purpose, so a new paragraph starts. Hyphens at a line end are rejoined only when the continuation begins with a lowercase letter, which is what stops Coca- and Cola being welded into a word that was never in the document.
Tables are the least certain part of the conversion. Column boundaries are found by clustering the horizontal positions of text runs, so a real table converts cleanly while a two-column page layout can occasionally be read as one.
Extraction runs through PDF.js in a Web Worker inside your browser. What would be uploaded here is the document's text, which is usually the confidential part. Nothing about the text is sent, and the one request a run makes asks whether you have tasks left - a tool name and a hash of the bytes.
Công cụ này không làm được gì
- Images are not extracted. The output is text and structure only; if you want the pages as pictures, PDF to image is the tool for that.
- Layout is not preserved. Columns, text boxes and absolute positioning collapse into a single reading order, which is what Markdown is for and what makes it re-flow.
Câu hỏi thường gặp
- Công cụ quyết định đâu là tiêu đề bằng cách nào?
- Nếu PDF có mục lục dấu trang, các dấu trang đó trở thành tiêu đề và độ sâu của chúng trở thành cấp tiêu đề, nên không có gì phải đoán. Ở nơi không có mục lục, một dòng được nâng lên khi nó lớn hơn văn bản thân bài xung quanh, ngắn, và theo sau là văn xuôi thông thường. Cả hai nguồn đều có công tắc riêng.
- Bảng có được chuyển đổi không?
- Có, dưới dạng bảng pipe của Markdown, khi bộ phát hiện đủ chắc chắn về các cột. Vị trí được gom cụm từ cả hai mép của văn bản, nên các cột số căn phải cũng được tìm thấy như các cột căn trái. Nếu các con số thuộc về một bảng tính, PDF sang CSV là lối đi tốt hơn.
- Các hình ảnh trong tài liệu của tôi thì sao?
- Chúng bị bỏ lại. Markdown tham chiếu tới hình ảnh chứ không chứa chúng, nên mang chúng theo sẽ đồng nghĩa với việc ghi ra một thư mục tệp bên cạnh tệp .md, và thay vào đó bản chuyển đổi tạo ra một tệp văn bản duy nhất, khép kín.
- PDF của tôi là bản quét và không có gì hiện ra. Tại sao?
- Một trang quét là ảnh chụp của văn bản: không có ký tự nào để bộ chuyển đổi đọc, nên công cụ dừng lại kèm lời giải thích thay vì ghi ra một tài liệu trống. Hãy chạy OCR PDF trước, rồi mới chuyển đổi.
- Tài liệu của tôi có được tải lên để chuyển đổi không?
- Không. Văn bản được trích xuất trong trình duyệt của bạn và Markdown cũng được ghi ở đó, không có yêu cầu gửi ra ngoài tại bất kỳ thời điểm nào. Bạn có thể kiểm chứng thay vì phải tin: theo dõi tab mạng trong khi chuyển đổi, hoặc ngắt kết nối hoàn toàn rồi vẫn chuyển đổi.