PDF เป็น CSV
ดึงตารางออกจาก PDF เป็นไฟล์คั่นด้วยจุลภาค
เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์
เครื่องมือนี้ทำอะไร
A PDF has no idea what a table is. There are no cells and no rows in the file, only text placed at coordinates that happen to line up, so recovering a table means finding the columns again. This tool clusters the horizontal positions of every text run, looks for a stretch of lines whose runs all fall into the same bands, and writes what it finds as CSV. Because it is inference rather than reading, every table comes with a confidence score, and the threshold is yours to set.
The usual case is numbers you would otherwise retype: a bank statement that only arrives as a PDF, invoice lines that have to be reconciled, a price list from a supplier, or an export from a system that offers PDF and nothing else.
วิธีทำงาน
- Drop the PDF onto this page. If it is a scan with no text layer, you are sent to OCR PDF rather than given an empty spreadsheet.
- Pick the delimiter. Comma is standard; semicolon is what Excel expects in most of Europe; tab suits pasting straight into a sheet.
- Set the confidence threshold. 55 per cent is the default; lower it to catch loosely aligned tables, raise it if prose is being mistaken for one.
- Choose one file per table, or a single file with the tables one after another.
- Press Extract tables. Each output shows its page and its confidence, so a doubtful table is visible as doubtful before you use it.
The detail that separates a table detector that works from one that does not is right alignment. Number columns are set flush right, so their runs share an end position rather than a start position, and a detector that clusters only on where text begins misses every financial table - which is most of what anybody converts. Both edges are clustered here.
Nothing is returned below the confidence threshold, and that is a deliberate refusal rather than a gap. A spreadsheet full of misaligned fragments looks like a result and is worse than no result, because someone will use it.
The threshold is a real dial. At 30 per cent almost any run of aligned lines is offered, which is useful for a loosely typeset table you already know is there. At 90 per cent only tables with clean, consistent columns survive, which is what you want when converting a hundred pages unattended. The default of 55 is where a typical report's tables come through and its paragraphs do not.
Detection and writing both happen in your browser. The numbers in a bank statement are about as sensitive as a document gets, and none of them are transmitted: the PDF is read and the CSV written without any part of either reaching a network.
สิ่งที่เครื่องมือนี้ทำไม่ได้
- Ruling lines are ignored. Columns are found from where the text sits, so a heavily bordered table whose text does not line up is still missed, and a table with no borders at all converts fine as long as it is aligned.
- Merged and spanning cells flatten. A cell spanning three columns is written into one of them and the rest are left empty, because CSV has no way to say that a cell covers more than one column.
- A scanned table has no text to cluster, so nothing is found. Run OCR PDF first to add a text layer, then come back.
คำถามที่คนมักถาม
- เครื่องมือหาตารางเจอได้อย่างไร
- ด้วยเรขาคณิต ตารางปรากฏเป็นชุดบรรทัดที่ต่อเนื่องกันซึ่งข้อความตกอยู่ในแถบแนวตั้งเดียวกัน ขอบทั้งสองด้านของแต่ละชุดถูกจัดกลุ่ม เพราะคอลัมน์ตัวเลขชิดขวาและมิฉะนั้นตัวตรวจจับจะมองไม่เห็น ทุกตารางที่เข้าข่ายจะได้คะแนนความมั่นใจ และสิ่งที่ต่ำกว่าเกณฑ์ของคุณจะถูกทิ้งแทนที่จะเดา
- มีตารางที่เห็นอยู่ชัด ๆ แต่ถูกมองข้าม ต้องทำอย่างไรต่อ
- ลดเกณฑ์ความมั่นใจลง ค่า 40 เปอร์เซ็นต์จับตารางที่จัดแนวหลวม ๆ ได้เกือบทั้งหมดซึ่งค่าเริ่มต้นปฏิเสธไป ถ้ายังไม่ช่วย ข้อความในตารางนั้นน่าจะไม่ได้เรียงเป็นคอลัมน์เลย ซึ่งเกิดกับเซลล์ที่ตัดขึ้นบรรทัดใหม่มาก ๆ อย่างน้อย PDF เป็นข้อความจะให้เนื้อหามาให้คุณจัดการเองด้วยมือ
- ควรเลือกตัวคั่นแบบไหน
- จุลภาค ถ้าไฟล์จะไปใช้กับอะไรก็ตามที่ไม่ใช่ Excel แบบยุโรป อัฒภาค ถ้าตัวเลขในภาษาของคุณใช้จุลภาคเป็นตัวคั่นทศนิยม เพราะ Excel จะอ่านไฟล์ที่คั่นด้วยอัฒภาคและทำไฟล์ที่คั่นด้วยจุลภาคเพี้ยน ส่วนแท็บปลอดภัยที่สุดสำหรับการวางลงในสเปรดชีตที่เปิดอยู่โดยตรง
- อ่านตารางจากหน้าที่สแกนมาได้ไหม
- ไม่ได้ด้วยตัวเอง ไฟล์สแกนคือภาพ จึงไม่มีตำแหน่งข้อความให้จัดกลุ่ม ให้ใช้ OCR PDF ก่อน แล้วตารางจะถูกตรวจจับได้ แต่ความแม่นยำของ OCR กับตัวเลขก็ควรตรวจทีละเซลล์ก่อนที่คุณจะเชื่อถือตัวเลขเหล่านั้น
- ไฟล์ถูกอัปโหลดไปที่ไหนไหม
- ไม่ PDF ถูกอ่านและ CSV ถูกเขียนในเบราว์เซอร์ของคุณทั้งหมด รายการเดินบัญชีหรือตารางเงินเดือนจึงไม่เคยออกจากเครื่องที่คุณเปิดมัน ไม่มีบัญชีผู้ใช้และไม่มีขีดจำกัดรายวัน และเครื่องมือยังทำงานได้แม้ปิดการเชื่อมต่อ