ข้ามไปยังเนื้อหา

PDF เป็น Word

สร้างไฟล์ .docx ที่แก้ไขได้ขึ้นใหม่จากข้อความใน PDF

ประมวลผลบนเครื่องของคุณ

เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์

เครื่องมือนี้ทำอะไร

This rebuilds a Word document from what a PDF actually contains, which is glyphs at coordinates rather than paragraphs. Headings, lists and tables are inferred from the geometry of the page and written into a real .docx. It is a reconstruction, not the original recovered: the text re-flows in Word, so line and page breaks will not always fall where the PDF put them.

Use it when you have to change words in a document nobody sent you the source for: a contract that needs amending, a template letter worth re-using, a report whose figures have moved on, or a page of text you would rather edit than retype.

วิธีทำงาน

  1. Drop the PDF onto this page. One file at a time, up to 200 MB.
  2. Choose the layout mode. Flowing gives paragraphs that re-wrap as you edit; preserve keeps the original line breaks and looks closer to the PDF.
  3. Leave heading, table and list detection on unless the document is plain prose you want kept plain.
  4. Under advanced options, give a page range to convert one chapter rather than the whole file.
  5. Press Convert to Word.
  6. Open the .docx and check the tables first - they are where reconstruction is hardest.

A PDF does not contain the Word document it was made from. When the PDF was written, the paragraph, the heading style and the table were thrown away, and what survived was ink at coordinates. Turning that back into structure means inferring it: a line in a larger or bolder face becomes a heading, lines sharing a hanging indent become a list, text in consistent vertical gutters becomes a table. The same analysis feeds PDF to Markdown, so there is one detector to improve rather than two that would slowly disagree.

The two modes answer different questions. Flowing is for editing: sentences become paragraphs that re-wrap when you change a word, which is what makes the file useful and also why it will not look identical to the PDF. Preserve is for looking at: each line becomes its own paragraph, so the page keeps its shape and edits badly, since typing into one line pushes nothing along to the next.

A scan is checked for before any work starts. The first five pages are sampled for a text layer, and if there is none the conversion stops and points at OCR, because a scan is a picture of words with nothing in it to convert. That is a deliberate choice over handing you an empty document and letting you conclude the tool is broken.

The .docx is assembled in your browser - the OOXML written by hand and zipped locally - so no part of a confidential contract is uploaded to be converted. The only thing that crosses the network is the allowance check, which is worth knowing when the document is the sort you would not put through a web service in the first place.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • The document is inferred from page geometry rather than recovered from an original, so expect to correct some headings, spacing and table edges by hand.
  • Pictures are not carried into the .docx. The conversion works from the text layer; use PDF to image if you need the artwork from a page.
  • Multi-column layouts, sidebars and footnotes are flattened into one stream of text in the order the analysis reads them.
  • A scanned PDF has no text to convert. It is refused with a link to OCR rather than producing an empty document.

คำถามที่คนมักถาม

ไฟล์ Word จะหน้าตาเหมือน PDF ทุกอย่างไหม
ไม่ และไม่มีตัวแปลงไหนสัญญาแบบนั้นได้อย่างซื่อสัตย์ Word จัดเรียงข้อความใหม่ด้วยกฎการขึ้นบรรทัดและมาตรวัดฟอนต์ของตัวเอง จุดขึ้นหน้าใหม่จึงขยับและระยะห่างต่างออกไป โหมดคงบรรทัดจะเก็บการขึ้นบรรทัดเดิมไว้และหน้าตาใกล้เคียงกว่า ส่วนโหมดไหลต่อเนื่องดูเหมือนน้อยกว่าแต่เป็นแบบที่คุณแก้ไขได้จริง สำหรับจดหมายหรือรายงาน ความต่างนี้เป็นแค่เรื่องหน้าตา
PDF ของฉันเป็นงานสแกนและแปลงไม่ได้ เพราะอะไร
เพราะงานสแกนบรรจุภาพถ่ายของคำ ไม่ใช่ตัวคำ ระบบจะสุ่มตรวจชั้นข้อความจากห้าหน้าแรก และเมื่อไม่มีชั้นข้อความ การแปลงจะหยุดแล้วส่งคุณไปที่ OCR แทน ให้ใช้เครื่องมือ OCR PDF ก่อน แล้วค่อยกลับมาแปลงผลลัพธ์
ตารางถูกนำมาด้วยไหม
ตามปกติได้ เมื่อตารางถูกจัดเป็นคอลัมน์ที่สม่ำเสมอ การตรวจจับใช้เรขาคณิต คือมองหาช่องว่างแนวตั้งที่คงอยู่ทุกแถว ตารางการเงินที่สะอาดจึงแปลงได้ดี ส่วนตารางที่ยึดกันด้วยเส้นตีและระยะห่างที่ไม่สม่ำเสมออาจออกมาเป็นย่อหน้าธรรมดา เซลล์ที่ผสานกันไว้ไม่ถูกสร้างขึ้นใหม่
รูปภาพใน PDF ของฉันจะเป็นอย่างไร
รูปภาพไม่ถูกนำเข้าไปในเอกสาร Word การแปลงทำงานจากชั้นข้อความ ภาพถ่าย โลโก้ และแผนภาพจึงถูกทิ้งไว้ ถ้าคุณต้องการรูปเหล่านั้น ให้ใช้เครื่องมือ PDF เป็นรูปภาพกับหน้าที่ต้องการ แล้วนำรูปไปแทรกในไฟล์ .docx ด้วยตัวเอง
ฉันควรใช้โหมดการจัดหน้าแบบไหน
แบบไหลต่อเนื่องถ้าคุณจะแก้ไขข้อความ เพราะย่อหน้าจะทำตัวเป็นย่อหน้าจริง ๆ แบบคงบรรทัดถ้าคุณต้องการอะไรที่หน้าตาเหมือน PDF เป็นหลักและจะใช้อ่านอย่างเดียว เพราะทุกบรรทัดจะกลายเป็นย่อหน้าของตัวเอง
เอกสารของฉันถูกอัปโหลดเพื่อแปลงไหม
ไม่ ไฟล์ PDF ถูกอ่านและไฟล์ .docx ถูกเขียนขึ้นในเบราว์เซอร์ของคุณทั้งหมด โดยไฟล์ Word ถูกประกอบขึ้นในเครื่อง ไม่ใช่โดยเซิร์ฟเวอร์ ไม่มีอะไรถูกส่งออกไป ซึ่งคุณยืนยันได้ในแท็บเครือข่ายของเครื่องมือสำหรับนักพัฒนา และเครื่องมือนี้ยังทำงานได้แม้ปิดการเชื่อมต่อ

เครื่องมือที่เกี่ยวข้อง