PDF เป็นข้อความ
ดึงตัวอักษรออกมาเป็นไฟล์ .txt ธรรมดา
เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์
เครื่องมือนี้ทำอะไร
This extracts the words from a PDF and writes them into a plain .txt file. A PDF does not store text as sentences - it stores glyphs at coordinates - so reading order, word boundaries and paragraph breaks are all reconstructed from the geometry of the page. Every extractor guesses that slightly differently, which is why the same document gives different output in different tools, and why there are two modes here rather than one.
Lifting a quotation out of a report without retyping it, feeding a document to a script or a language model, searching a book you only have as a PDF, or checking what a page actually contains when a table detector insists it found nothing.
วิธีทำงาน
- Drop the PDF onto this page. One file at a time, up to 300 MB.
- Choose reading order for prose you intend to paste somewhere, or layout to keep columns and tables lined up with spaces.
- Decide whether pages are separated by a marker line and whether page numbers are printed into the file.
- Leave rejoin hyphenated words on so a word broken across a line break comes back together.
- Under advanced options, set a page range, or switch the encoding to UTF-8 with a BOM if Excel or Notepad will open the file.
- Press Convert to text and the .txt downloads on its own.
The two modes solve opposite problems. Reading order joins the lines of a paragraph back into flowing prose, which is what you want in an email, a script or a quotation. Layout mode keeps every line where it was on the page, padding with spaces so columns stay aligned - right for a table, a code listing or a form, and wrong for prose, where the padding becomes ragged whitespace you then strip out. If the output looks strange, the mode is usually why.
Rejoining hyphenated words is more careful than a search-and-replace. A hyphen at the end of a line is joined only when the character before it and the first character of the next line are both lowercase, so a word broken as inter- and national comes back as international, while a compound such as self-service, or a hyphen followed by a capitalised name, is left as written. The rule errs towards leaving a real hyphen alone rather than towards editing your text.
A scanned document is detected before any work happens. The first five pages are sampled for a text layer, and when there is none the tool stops and points at OCR, since a scan is a picture of words with nothing in it to extract. Reporting no text found would be technically true and practically useless.
The encoding option is small and saves a familiar afternoon. UTF-8 without a byte order mark is what a script, a git repository or a text editor wants. UTF-8 with a BOM is what makes Excel and Notepad recognise the file as UTF-8 instead of guessing a legacy code page and mangling every accented character. If names come out as strings of symbols, the BOM is what was missing.
Extraction runs in a Web Worker and the text file is written there too, so a confidential document is never sent anywhere to have its words read. The result reports the page count, the character and word counts and how many pages came out empty - often the quickest sign that part of a document is scanned while the rest is not.
สิ่งที่เครื่องมือนี้ทำไม่ได้
- Columns, sidebars, headers and footnotes come out in the order the reconstruction decides on, which is not always the order a person reads them in. Layout mode is often clearer for those pages.
- A scanned page contains no text to extract. That is detected and reported with a link to OCR rather than writing an empty file.
คำถามที่คนมักถาม
- ลำดับการอ่านกับการคงเค้าโครงหน้าต่างกันอย่างไร
- ลำดับการอ่านจะประกอบย่อหน้าขึ้นใหม่ โดยเชื่อมบรรทัดของแต่ละย่อหน้าเข้าเป็นข้อความที่ไหลต่อกันและนำไปวางที่ไหนก็ได้ ส่วนการคงเค้าโครงหน้าจะรักษารูปทรงของหน้าไว้ด้วยการเติมช่องว่าง คอลัมน์ ตาราง และรายการโค้ดจึงยังเรียงตรงกันเมื่อดูด้วยฟอนต์ความกว้างคงที่ ใช้ลำดับการอ่านกับร้อยแก้ว และใช้การคงเค้าโครงเมื่อตำแหน่งแนวนอนของข้อความมีความหมาย
- ไม่มีอะไรออกมาจาก PDF ของฉันเลย เพราะอะไร
- เกือบแน่นอนว่าเพราะมันเป็นงานสแกน หน้าที่สแกนมาบรรจุภาพถ่ายของคำ ไม่ใช่ตัวคำ จึงไม่มีชั้นข้อความให้ดึงออกมา ระบบจะสุ่มตรวจห้าหน้าแรกก่อนเริ่มทำงาน และถ้าไม่มีชั้นข้อความ เครื่องมือจะหยุดแล้วส่งคุณไปที่ OCR PDF ซึ่งเปลี่ยนภาพให้เป็นข้อความจริงที่คุณดึงออกมาได้ต่อไป
- ข้อความที่ได้สลับกันมั่ว แก้ได้ไหม
- ลองโหมดคงเค้าโครงหน้าก่อน ผลลัพธ์ที่สลับกันมักหมายถึงหน้าหลายคอลัมน์ที่บรรทัดถูกเชื่อมข้ามคอลัมน์กันไป การคงเค้าโครงจะเก็บแต่ละคอลัมน์ไว้ที่เดิม ซึ่งอ่านง่ายกว่าและจัดการต่อง่ายกว่า เอกสารบางฉบับยังวางหัวกระดาษ ท้ายกระดาษ และหมายเหตุริมหน้าไว้ในลำดับที่ไม่เกี่ยวกับวิธีอ่าน และไม่มีตัวดึงข้อความไหนรู้ดีกว่านั้นได้อย่างน่าเชื่อถือ
- อักขระที่มีเครื่องหมายกำกับเสียงดูผิดเพี้ยนตอนฉันเปิดไฟล์ ต้องทำอย่างไร
- เปลี่ยนการเข้ารหัสเป็น UTF-8 พร้อม BOM แล้วแปลงอีกครั้ง เครื่องหมายลำดับไบต์จะบอก Excel และ Notepad ว่าไฟล์เป็น UTF-8 แทนที่จะปล่อยให้เดาโค้ดเพจรุ่นเก่า ซึ่งเป็นสาเหตุที่ตัวอักษรมีเครื่องหมายกำกับกลายเป็นสัญลักษณ์สองตัว ให้ใช้ UTF-8 ธรรมดาสำหรับสคริปต์ โปรแกรมแก้ไข และระบบควบคุมเวอร์ชัน
- เอกสารถูกอัปโหลดเพื่อดึงข้อความไหม
- ไม่ ไฟล์ PDF ถูกแยกวิเคราะห์และไฟล์ข้อความถูกเขียนขึ้นในเบราว์เซอร์ของคุณ และไม่มีอะไรถูกส่งออกไปในขั้นตอนใดเลย ซึ่งเป็นเรื่องที่ควรรู้เมื่อเอกสารเป็นสัญญาหรือเวชระเบียน คุณดูแท็บเครือข่ายของเครื่องมือสำหรับนักพัฒนาระหว่างการแปลงเพื่อยืนยันได้ และเครื่องมือนี้ทำงานแบบออฟไลน์ได้