PDF sang CSV
Rút các bảng ra khỏi PDF dưới dạng tệp phân tách bằng dấu phẩy.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
A PDF has no idea what a table is. There are no cells and no rows in the file, only text placed at coordinates that happen to line up, so recovering a table means finding the columns again. This tool clusters the horizontal positions of every text run, looks for a stretch of lines whose runs all fall into the same bands, and writes what it finds as CSV. Because it is inference rather than reading, every table comes with a confidence score, and the threshold is yours to set.
The usual case is numbers you would otherwise retype: a bank statement that only arrives as a PDF, invoice lines that have to be reconciled, a price list from a supplier, or an export from a system that offers PDF and nothing else.
Cách hoạt động
- Drop the PDF onto this page. If it is a scan with no text layer, you are sent to OCR PDF rather than given an empty spreadsheet.
- Pick the delimiter. Comma is standard; semicolon is what Excel expects in most of Europe; tab suits pasting straight into a sheet.
- Set the confidence threshold. 55 per cent is the default; lower it to catch loosely aligned tables, raise it if prose is being mistaken for one.
- Choose one file per table, or a single file with the tables one after another.
- Press Extract tables. Each output shows its page and its confidence, so a doubtful table is visible as doubtful before you use it.
The detail that separates a table detector that works from one that does not is right alignment. Number columns are set flush right, so their runs share an end position rather than a start position, and a detector that clusters only on where text begins misses every financial table - which is most of what anybody converts. Both edges are clustered here.
Nothing is returned below the confidence threshold, and that is a deliberate refusal rather than a gap. A spreadsheet full of misaligned fragments looks like a result and is worse than no result, because someone will use it.
The threshold is a real dial. At 30 per cent almost any run of aligned lines is offered, which is useful for a loosely typeset table you already know is there. At 90 per cent only tables with clean, consistent columns survive, which is what you want when converting a hundred pages unattended. The default of 55 is where a typical report's tables come through and its paragraphs do not.
Detection and writing both happen in your browser. The numbers in a bank statement are about as sensitive as a document gets, and none of them are transmitted: the PDF is read and the CSV written without any part of either reaching a network.
Công cụ này không làm được gì
- Ruling lines are ignored. Columns are found from where the text sits, so a heavily bordered table whose text does not line up is still missed, and a table with no borders at all converts fine as long as it is aligned.
- Merged and spanning cells flatten. A cell spanning three columns is written into one of them and the rest are left empty, because CSV has no way to say that a cell covers more than one column.
- A scanned table has no text to cluster, so nothing is found. Run OCR PDF first to add a text layer, then come back.
Câu hỏi thường gặp
- Công cụ tìm các bảng bằng cách nào?
- Bằng hình học. Một bảng hiện ra thành một chuỗi các dòng liên tiếp mà văn bản rơi vào cùng những dải dọc. Cả hai mép của mỗi chuỗi đều được gom cụm, vì các cột số căn phải và nếu không sẽ vô hình với bộ phát hiện. Mỗi ứng viên nhận một điểm chắc chắn, và bất cứ gì dưới ngưỡng của bạn đều bị loại bỏ thay vì đoán.
- Nó bỏ sót một bảng rõ ràng đang có ở đó. Giờ làm sao?
- Hạ ngưỡng chắc chắn xuống - 40 phần trăm bắt được hầu hết các bảng căn lỏng lẻo mà mức mặc định từ chối. Nếu điều đó không giúp được, có lẽ văn bản của bảng không xếp thẳng hàng thành cột chút nào, điều thường xảy ra với các ô ngắt dòng nhiều; PDF sang văn bản ít nhất cũng cho bạn nội dung để xử lý bằng tay.
- Tôi nên chọn dấu phân tách nào?
- Dấu phẩy nếu tệp sẽ vào bất cứ thứ gì khác ngoài một Excel bản châu Âu. Dấu chấm phẩy nếu số liệu ở vùng của bạn dùng dấu phẩy làm dấu thập phân, vì khi đó Excel đọc tệp phân tách bằng dấu chấm phẩy và làm hỏng tệp phân tách bằng dấu phẩy. Tab là lựa chọn an toàn nhất để dán thẳng vào một bảng tính đang mở.
- Công cụ có đọc được bảng từ một trang quét không?
- Không tự nó làm được. Bản quét là hình ảnh, nên không có vị trí văn bản nào để gom cụm. Hãy chạy OCR PDF trước và bảng sẽ trở nên phát hiện được - dù độ chính xác của OCR trên các con số đáng để kiểm tra từng ô trước khi bạn tin vào chúng.
- Các tệp có được tải lên đâu đó không?
- Không. PDF được đọc và CSV được ghi hoàn toàn trong trình duyệt của bạn, nên một bản sao kê ngân hàng hay bảng lương không bao giờ rời khỏi máy bạn mở nó. Không có tài khoản, không có giới hạn hằng ngày, và công cụ vẫn hoạt động khi tắt kết nối.