PDF sang Word
Dựng lại một tệp .docx chỉnh sửa được từ phần chữ trong PDF.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
This rebuilds a Word document from what a PDF actually contains, which is glyphs at coordinates rather than paragraphs. Headings, lists and tables are inferred from the geometry of the page and written into a real .docx. It is a reconstruction, not the original recovered: the text re-flows in Word, so line and page breaks will not always fall where the PDF put them.
Use it when you have to change words in a document nobody sent you the source for: a contract that needs amending, a template letter worth re-using, a report whose figures have moved on, or a page of text you would rather edit than retype.
Cách hoạt động
- Drop the PDF onto this page. One file at a time, up to 200 MB.
- Choose the layout mode. Flowing gives paragraphs that re-wrap as you edit; preserve keeps the original line breaks and looks closer to the PDF.
- Leave heading, table and list detection on unless the document is plain prose you want kept plain.
- Under advanced options, give a page range to convert one chapter rather than the whole file.
- Press Convert to Word.
- Open the .docx and check the tables first - they are where reconstruction is hardest.
A PDF does not contain the Word document it was made from. When the PDF was written, the paragraph, the heading style and the table were thrown away, and what survived was ink at coordinates. Turning that back into structure means inferring it: a line in a larger or bolder face becomes a heading, lines sharing a hanging indent become a list, text in consistent vertical gutters becomes a table. The same analysis feeds PDF to Markdown, so there is one detector to improve rather than two that would slowly disagree.
The two modes answer different questions. Flowing is for editing: sentences become paragraphs that re-wrap when you change a word, which is what makes the file useful and also why it will not look identical to the PDF. Preserve is for looking at: each line becomes its own paragraph, so the page keeps its shape and edits badly, since typing into one line pushes nothing along to the next.
A scan is checked for before any work starts. The first five pages are sampled for a text layer, and if there is none the conversion stops and points at OCR, because a scan is a picture of words with nothing in it to convert. That is a deliberate choice over handing you an empty document and letting you conclude the tool is broken.
The .docx is assembled in your browser - the OOXML written by hand and zipped locally - so no part of a confidential contract is uploaded to be converted. The only thing that crosses the network is the allowance check, which is worth knowing when the document is the sort you would not put through a web service in the first place.
Công cụ này không làm được gì
- The document is inferred from page geometry rather than recovered from an original, so expect to correct some headings, spacing and table edges by hand.
- Pictures are not carried into the .docx. The conversion works from the text layer; use PDF to image if you need the artwork from a page.
- Multi-column layouts, sidebars and footnotes are flattened into one stream of text in the order the analysis reads them.
- A scanned PDF has no text to convert. It is refused with a link to OCR rather than producing an empty document.
Câu hỏi thường gặp
- Tệp Word có giống hệt PDF không?
- Không, và không bộ chuyển đổi nào có thể thành thật hứa điều đó. Word dàn lại chữ theo cách ngắt dòng và số đo phông của riêng nó, nên chỗ ngắt trang xê dịch và khoảng cách khác đi. Chế độ giữ nguyên giữ các chỗ ngắt dòng gốc và trông gần giống hơn về mặt hình thức; chế độ trôi tự do trông ít giống hơn nhưng lại là chế độ bạn thực sự chỉnh sửa được. Với một lá thư hay một báo cáo thì khác biệt đó chỉ là bề ngoài.
- PDF của bạn là bản quét và không chuyển được. Vì sao?
- Vì một bản quét chứa ảnh chụp của chữ, chứ không phải chữ. Năm trang đầu được lấy mẫu để tìm lớp chữ, và khi không có lớp nào thì việc chuyển đổi dừng lại và hướng bạn sang OCR. Hãy chạy OCR PDF trước, rồi quay lại và chuyển kết quả.
- Bảng có được giữ lại không?
- Thường là có, khi bảng được bố trí thành các cột đều đặn. Việc nhận diện dựa trên hình học - nó tìm những khe dọc còn nguyên qua mọi hàng - nên một bảng tài chính gọn gàng chuyển tốt, còn một bảng chỉ dựa vào đường kẻ và khoảng cách không đều có thể ra thành các đoạn văn thông thường. Ô đã gộp không được dựng lại.
- Hình ảnh trong PDF của bạn thì sao?
- Chúng không được đưa vào tài liệu Word. Việc chuyển đổi làm từ lớp chữ, nên ảnh chụp, logo và sơ đồ bị bỏ lại. Nếu cần chúng, hãy chạy PDF sang ảnh trên các trang đó rồi tự chèn ảnh vào tệp .docx.
- Bạn nên dùng chế độ bố cục nào?
- Trôi tự do nếu bạn sắp chỉnh sửa chữ, vì các đoạn văn hành xử đúng như đoạn văn. Giữ nguyên nếu bạn chủ yếu cần một thứ trông giống PDF và chỉ để đọc, vì mỗi dòng trở thành một đoạn riêng.
- Tài liệu của bạn có bị tải lên để chuyển đổi không?
- Không. PDF được đọc và tệp .docx được tạo hoàn toàn trong trình duyệt của bạn, với tệp Word được ghép cục bộ chứ không phải bởi một máy chủ. Không có gì được truyền đi, điều bạn có thể tự xác nhận trong thẻ mạng của công cụ dành cho nhà phát triển, và công cụ vẫn tiếp tục hoạt động khi tắt kết nối.