ข้ามไปยังเนื้อหา

PDF เป็น CSV

ดึงตารางออกจาก PDF เป็นไฟล์คั่นด้วยจุลภาค

ประมวลผลบนเครื่องของคุณ

เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์

เครื่องมือนี้ทำอะไร

A PDF has no idea what a table is. There are no cells and no rows in the file, only text placed at coordinates that happen to line up, so recovering a table means finding the columns again. This tool clusters the horizontal positions of every text run, looks for a stretch of lines whose runs all fall into the same bands, and writes what it finds as CSV. Because it is inference rather than reading, every table comes with a confidence score, and the threshold is yours to set.

The usual case is numbers you would otherwise retype: a bank statement that only arrives as a PDF, invoice lines that have to be reconciled, a price list from a supplier, or an export from a system that offers PDF and nothing else.

วิธีทำงาน

  1. Drop the PDF onto this page. If it is a scan with no text layer, you are sent to OCR PDF rather than given an empty spreadsheet.
  2. Pick the delimiter. Comma is standard; semicolon is what Excel expects in most of Europe; tab suits pasting straight into a sheet.
  3. Set the confidence threshold. 55 per cent is the default; lower it to catch loosely aligned tables, raise it if prose is being mistaken for one.
  4. Choose one file per table, or a single file with the tables one after another.
  5. Press Extract tables. Each output shows its page and its confidence, so a doubtful table is visible as doubtful before you use it.

The detail that separates a table detector that works from one that does not is right alignment. Number columns are set flush right, so their runs share an end position rather than a start position, and a detector that clusters only on where text begins misses every financial table - which is most of what anybody converts. Both edges are clustered here.

Nothing is returned below the confidence threshold, and that is a deliberate refusal rather than a gap. A spreadsheet full of misaligned fragments looks like a result and is worse than no result, because someone will use it.

The threshold is a real dial. At 30 per cent almost any run of aligned lines is offered, which is useful for a loosely typeset table you already know is there. At 90 per cent only tables with clean, consistent columns survive, which is what you want when converting a hundred pages unattended. The default of 55 is where a typical report's tables come through and its paragraphs do not.

Detection and writing both happen in your browser. The numbers in a bank statement are about as sensitive as a document gets, and none of them are transmitted: the PDF is read and the CSV written without any part of either reaching a network.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • Ruling lines are ignored. Columns are found from where the text sits, so a heavily bordered table whose text does not line up is still missed, and a table with no borders at all converts fine as long as it is aligned.
  • Merged and spanning cells flatten. A cell spanning three columns is written into one of them and the rest are left empty, because CSV has no way to say that a cell covers more than one column.
  • A scanned table has no text to cluster, so nothing is found. Run OCR PDF first to add a text layer, then come back.

คำถามที่คนมักถาม

เครื่องมือหาตารางเจอได้อย่างไร
ด้วยเรขาคณิต ตารางปรากฏเป็นชุดบรรทัดที่ต่อเนื่องกันซึ่งข้อความตกอยู่ในแถบแนวตั้งเดียวกัน ขอบทั้งสองด้านของแต่ละชุดถูกจัดกลุ่ม เพราะคอลัมน์ตัวเลขชิดขวาและมิฉะนั้นตัวตรวจจับจะมองไม่เห็น ทุกตารางที่เข้าข่ายจะได้คะแนนความมั่นใจ และสิ่งที่ต่ำกว่าเกณฑ์ของคุณจะถูกทิ้งแทนที่จะเดา
มีตารางที่เห็นอยู่ชัด ๆ แต่ถูกมองข้าม ต้องทำอย่างไรต่อ
ลดเกณฑ์ความมั่นใจลง ค่า 40 เปอร์เซ็นต์จับตารางที่จัดแนวหลวม ๆ ได้เกือบทั้งหมดซึ่งค่าเริ่มต้นปฏิเสธไป ถ้ายังไม่ช่วย ข้อความในตารางนั้นน่าจะไม่ได้เรียงเป็นคอลัมน์เลย ซึ่งเกิดกับเซลล์ที่ตัดขึ้นบรรทัดใหม่มาก ๆ อย่างน้อย PDF เป็นข้อความจะให้เนื้อหามาให้คุณจัดการเองด้วยมือ
ควรเลือกตัวคั่นแบบไหน
จุลภาค ถ้าไฟล์จะไปใช้กับอะไรก็ตามที่ไม่ใช่ Excel แบบยุโรป อัฒภาค ถ้าตัวเลขในภาษาของคุณใช้จุลภาคเป็นตัวคั่นทศนิยม เพราะ Excel จะอ่านไฟล์ที่คั่นด้วยอัฒภาคและทำไฟล์ที่คั่นด้วยจุลภาคเพี้ยน ส่วนแท็บปลอดภัยที่สุดสำหรับการวางลงในสเปรดชีตที่เปิดอยู่โดยตรง
อ่านตารางจากหน้าที่สแกนมาได้ไหม
ไม่ได้ด้วยตัวเอง ไฟล์สแกนคือภาพ จึงไม่มีตำแหน่งข้อความให้จัดกลุ่ม ให้ใช้ OCR PDF ก่อน แล้วตารางจะถูกตรวจจับได้ แต่ความแม่นยำของ OCR กับตัวเลขก็ควรตรวจทีละเซลล์ก่อนที่คุณจะเชื่อถือตัวเลขเหล่านั้น
ไฟล์ถูกอัปโหลดไปที่ไหนไหม
ไม่ PDF ถูกอ่านและ CSV ถูกเขียนในเบราว์เซอร์ของคุณทั้งหมด รายการเดินบัญชีหรือตารางเงินเดือนจึงไม่เคยออกจากเครื่องที่คุณเปิดมัน ไม่มีบัญชีผู้ใช้และไม่มีขีดจำกัดรายวัน และเครื่องมือยังทำงานได้แม้ปิดการเชื่อมต่อ

เครื่องมือที่เกี่ยวข้อง