PDF sang Excel
Rút các bảng từ một PDF ra thành một bảng tính.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
This finds the tables in a document and writes them into an .xlsx workbook. A PDF never records that something is a table, so the columns are located geometrically, from the vertical gutters that survive every row. Values that read cleanly as numbers are written as numbers, which is the difference between a column you can sum immediately and one Excel refuses to add up.
Bank statements, invoice line items, a supplier price list published only as a PDF, the results table in a report you need to chart - anywhere the figures are visible on the page and retyping them is the only other option.
Cách hoạt động
- Drop the PDF onto this page. One file at a time, up to 200 MB.
- Leave the confidence threshold at 55%. Lower it if a table was missed, raise it if something that is not a table came out as one.
- Leave joining across pages on for a table that continues overleaf, or switch on a sheet per page to keep them separate.
- Leave the header row option on so a heading repeated on a continuation page is not written in as data.
- Press Convert to Excel.
- Check the first and last rows of each sheet, where detection is least certain.
Two facts drive the detection. The first is that numeric columns align on their right edge, not their left, so a detector clustering only the start of each run reads 12.00 and 1,234.00 as two different columns - exactly on the financial statements that are most of what anyone converts. Alignment here is scored on the best of the start, end and centre consistency, so a left-aligned description column and a right-aligned money column both score well in the same table.
The second is that a wrong table is worse than no table. Two columns of prose have pin-sharp left edges and a wide gutter down the middle, and by alignment alone they look better than most real tables. Geometry cannot separate them, so there is a further gate on what the columns contain, and it rejects outright rather than quietly lowering the score. That is why a page of newspaper-style text does not arrive as a two-column spreadsheet.
Numbers keep the formatting they were printed with, and it has to come off before a cell can hold a real number: thousands separators, currency symbols, trailing percent signs, and negatives written in accountancy brackets, which become a minus. A comma is only removed when exactly three digits follow it, because it is a thousands separator in English and a decimal point in Indonesian. Anything still ambiguous stays text on purpose - a misplaced decimal point is worse than an unformatted cell.
When nothing reads as a table, the tool says so and points at PDF to text rather than handing you an empty workbook, which would imply the document had nothing in it. Lowering the confidence threshold is the first thing to try; reading the page with PDF to text is the second.
Detection and the writing of the workbook both happen on your device, the OOXML built and zipped in the browser, so a bank statement is never sent anywhere to be read. That matters more here than on most tools: the documents people convert to spreadsheets are usually the ones with account numbers on them.
Công cụ này không làm được gì
- Tables drawn with ruled lines are still found by their text alignment. The rules themselves live in the page graphics, which this pass does not read, so a table held together only by its lines and irregular spacing may be missed.
- Merged cells are not reconstructed. A heading spanning three columns lands in the first of them and the other two come out empty.
- Only values come across. Fonts, colours, borders, column widths and number formats are not carried into the workbook.
Câu hỏi thường gặp
- Công cụ tìm các bảng bằng cách nào?
- Bằng hình học. Lớp chữ cho mỗi ký tự một vị trí, và một khối các dòng mà phần chữ tách ra thành cùng những khe dọc hết hàng này đến hàng khác được coi là một bảng. Mỗi ứng viên được chấm điểm theo mức độ các cột thẳng hàng đều đặn, và bất cứ thứ gì dưới ngưỡng của bạn đều bị loại. Một bước kiểm tra nội dung riêng loại bỏ những cột văn xuôi, vốn nếu không sẽ được điểm cao hơn cả bảng thật.
- Các số có đến dưới dạng số mà bạn cộng được không?
- Có, khi chúng rõ ràng. Ký hiệu tiền tệ, dấu phân cách hàng nghìn, dấu phần trăm ở cuối và dấu ngoặc kiểu kế toán được loại bỏ, và các giá trị trong ngoặc trở thành số âm. Một dấu phẩy chỉ được coi là dấu phân cách hàng nghìn khi có đúng ba chữ số theo sau, vì trong tiếng Indonesia nó là dấu thập phân. Bất cứ thứ gì không phân tích được rõ ràng đều giữ nguyên là văn bản chứ không đoán mò.
- Nó bỏ sót một bảng, hoặc tìm ra một bảng không có thật. Giờ làm sao?
- Hãy chỉnh ngưỡng độ tin cậy. Nó chạy từ 30 đến 90 và bắt đầu ở 55; hạ nó xuống sẽ chấp nhận những ứng viên yếu hơn và tìm được nhiều bảng hơn, đổi lại có một số bảng sai. Nếu không tìm thấy gì ở bất kỳ mức nào, có thể trang không có cấu trúc cột thật sự, và PDF sang văn bản sẽ cho bạn thấy trang thực sự chứa những gì.
- Bảng của bạn kéo dài qua nhiều trang. Nó có bị chia nhỏ không?
- Theo mặc định thì không. Một bảng bắt đầu ở đầu một trang với cùng số cột như bảng kết thúc ở cuối trang trước được coi là phần tiếp nối và ghép vào một trang tính, còn hàng tiêu đề lặp lại của nó bị bỏ đi chứ không ghi vào như dữ liệu. Hãy bật "mỗi trang một trang tính" để tách chúng ra.
- Bản sao kê của bạn có bị tải lên đâu đó không?
- Không. Việc nhận diện và tạo sổ tính đều diễn ra trong trình duyệt của bạn, và không phần nào của tài liệu được truyền đi. Điều này kiểm tra được trong thẻ mạng của công cụ dành cho nhà phát triển, và công cụ vẫn tiếp tục hoạt động khi tắt kết nối.