ข้ามไปยังเนื้อหา

PDF เป็น Excel

ดึงตารางออกจาก PDF ไปไว้ในสเปรดชีต

ประมวลผลบนเครื่องของคุณ

เครื่องมือนี้ทำงานในเบราว์เซอร์ของคุณทั้งหมด ไฟล์ของคุณไม่ถูกอัปโหลดเลย และคุณตรวจสอบได้เองในแท็บเครือข่ายของเบราว์เซอร์ ตรวจสอบด้วยตัวเอง เปิดแท็บเครือข่ายของเบราว์เซอร์แล้วดู คุณจะเห็นคำขอเล็ก ๆ หนึ่งรายการที่ถามว่าคุณยังมีโควตางานเหลืออยู่หรือไม่ ซึ่งมีแค่ชื่อเครื่องมือกับค่าแฮช ไม่ใช่ตัวไฟล์

เครื่องมือนี้ทำอะไร

This finds the tables in a document and writes them into an .xlsx workbook. A PDF never records that something is a table, so the columns are located geometrically, from the vertical gutters that survive every row. Values that read cleanly as numbers are written as numbers, which is the difference between a column you can sum immediately and one Excel refuses to add up.

Bank statements, invoice line items, a supplier price list published only as a PDF, the results table in a report you need to chart - anywhere the figures are visible on the page and retyping them is the only other option.

วิธีทำงาน

  1. Drop the PDF onto this page. One file at a time, up to 200 MB.
  2. Leave the confidence threshold at 55%. Lower it if a table was missed, raise it if something that is not a table came out as one.
  3. Leave joining across pages on for a table that continues overleaf, or switch on a sheet per page to keep them separate.
  4. Leave the header row option on so a heading repeated on a continuation page is not written in as data.
  5. Press Convert to Excel.
  6. Check the first and last rows of each sheet, where detection is least certain.

Two facts drive the detection. The first is that numeric columns align on their right edge, not their left, so a detector clustering only the start of each run reads 12.00 and 1,234.00 as two different columns - exactly on the financial statements that are most of what anyone converts. Alignment here is scored on the best of the start, end and centre consistency, so a left-aligned description column and a right-aligned money column both score well in the same table.

The second is that a wrong table is worse than no table. Two columns of prose have pin-sharp left edges and a wide gutter down the middle, and by alignment alone they look better than most real tables. Geometry cannot separate them, so there is a further gate on what the columns contain, and it rejects outright rather than quietly lowering the score. That is why a page of newspaper-style text does not arrive as a two-column spreadsheet.

Numbers keep the formatting they were printed with, and it has to come off before a cell can hold a real number: thousands separators, currency symbols, trailing percent signs, and negatives written in accountancy brackets, which become a minus. A comma is only removed when exactly three digits follow it, because it is a thousands separator in English and a decimal point in Indonesian. Anything still ambiguous stays text on purpose - a misplaced decimal point is worse than an unformatted cell.

When nothing reads as a table, the tool says so and points at PDF to text rather than handing you an empty workbook, which would imply the document had nothing in it. Lowering the confidence threshold is the first thing to try; reading the page with PDF to text is the second.

Detection and the writing of the workbook both happen on your device, the OOXML built and zipped in the browser, so a bank statement is never sent anywhere to be read. That matters more here than on most tools: the documents people convert to spreadsheets are usually the ones with account numbers on them.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • Tables drawn with ruled lines are still found by their text alignment. The rules themselves live in the page graphics, which this pass does not read, so a table held together only by its lines and irregular spacing may be missed.
  • Merged cells are not reconstructed. A heading spanning three columns lands in the first of them and the other two come out empty.
  • Only values come across. Fonts, colours, borders, column widths and number formats are not carried into the workbook.

คำถามที่คนมักถาม

เครื่องมือนี้หาตารางเจอได้อย่างไร
ด้วยเรขาคณิต ชั้นข้อความบอกตำแหน่งของอักขระทุกตัว และกลุ่มบรรทัดที่หมึกแยกออกเป็นช่องว่างแนวตั้งชุดเดียวกันซ้ำแถวแล้วแถวเล่าจะถูกถือว่าเป็นตาราง ผู้เข้าชิงแต่ละชุดจะถูกให้คะแนนตามความสม่ำเสมอของการเรียงคอลัมน์ และอะไรที่ต่ำกว่าเกณฑ์ของคุณจะถูกทิ้ง นอกจากนี้ยังมีการตรวจเนื้อหาแยกอีกชั้นเพื่อคัดคอลัมน์ของร้อยแก้วออก ซึ่งไม่อย่างนั้นจะได้คะแนนสูงกว่าตารางจริง
ตัวเลขจะมาถึงในสภาพที่ฉันบวกได้ไหม
ได้ เมื่อตัวเลขไม่กำกวม สัญลักษณ์สกุลเงิน ตัวคั่นหลักพัน เครื่องหมายเปอร์เซ็นต์ท้ายค่า และวงเล็บแบบบัญชีจะถูกตัดออก และค่าที่อยู่ในวงเล็บจะกลายเป็นค่าติดลบ เครื่องหมายจุลภาคจะถือเป็นตัวคั่นหลักพันเฉพาะเมื่อมีตัวเลขตามมาสามหลักพอดี เพราะในภาษาอินโดนีเซียมันคือจุดทศนิยม อะไรที่แปลงค่าไม่ได้อย่างชัดเจนจะยังเป็นข้อความแทนที่จะถูกเดา
มันหาตารางไม่เจอ หรือเจอตารางที่ไม่มีอยู่ ต้องทำอย่างไรต่อ
ให้ขยับเกณฑ์ความมั่นใจ ค่านี้อยู่ระหว่าง 30 ถึง 90 และเริ่มต้นที่ 55 การลดลงจะรับผู้เข้าชิงที่อ่อนกว่าเข้ามาและเจอตารางมากขึ้น โดยแลกกับตารางลวงบางส่วน ถ้าไม่เจออะไรเลยไม่ว่าตั้งค่าไหน หน้านั้นอาจไม่มีโครงสร้างคอลัมน์จริง และเครื่องมือ PDF เป็นข้อความจะแสดงให้คุณเห็นว่าในหน้านั้นมีอะไรอยู่
ตารางของฉันยาวข้ามหลายหน้า มันจะถูกแยกออกจากกันไหม
ไม่ ตามค่าเริ่มต้น ตารางที่เริ่มบนสุดของหน้าและมีจำนวนคอลัมน์เท่ากับตารางที่จบอยู่ท้ายหน้าก่อนหน้า จะถูกถือว่าเป็นส่วนต่อและถูกเชื่อมเข้าเป็นชีตเดียว โดยแถวหัวตารางที่ซ้ำจะถูกตัดออกแทนที่จะถูกเขียนลงไปเป็นข้อมูล เปิดตัวเลือกหนึ่งชีตต่อหนึ่งหน้าเพื่อแยกออกจากกัน
รายการเดินบัญชีของฉันถูกอัปโหลดไปที่ไหนไหม
ไม่ ทั้งการตรวจจับและการเขียนเวิร์กบุ๊กเกิดขึ้นในเบราว์เซอร์ของคุณ และไม่มีส่วนใดของเอกสารถูกส่งออกไป เรื่องนี้ตรวจสอบได้ในแท็บเครือข่ายของเครื่องมือสำหรับนักพัฒนา และเครื่องมือนี้ยังทำงานได้แม้ปิดการเชื่อมต่อ

เครื่องมือที่เกี่ยวข้อง