ข้ามไปยังเนื้อหา

ดึงข้อมูล

ดึงข้อมูลช่องเดียวกันออกจากเอกสารทั้งกอง เป็น JSON หรือ CSV

ข้อความถูกส่งไปยังบริการ AI

ไฟล์ PDF ของคุณไม่ถูกอัปโหลดเลย ข้อความถูกดึงออกมาในเบราว์เซอร์ของคุณ แสดงให้คุณดู และส่งออกไปหลังจากที่คุณอนุมัติแล้วเท่านั้น

เครื่องมือนี้ทำอะไร

This reads a set of documents and returns one row each, carrying the fields you asked for. Text is extracted in your browser and sent to the AI provider you choose with your own key; the files themselves never leave your device. The output is a JSON array or a CSV, with a _file column so every row says which document it came from.

It exists for the batch: twenty supplier invoices that have to become twenty rows in a ledger, a month of receipts before an expense claim, bank statements you want opening and closing balances from, or a folder of ID documents whose numbers need typing into a system.

วิธีทำงาน

  1. Drop up to twenty PDFs onto this page, then choose your provider and paste your API key.
  2. Pick a preset - invoice, receipt, ID document or bank statement - or choose your own fields and list the names you want, one per line.
  3. Leave the missing-fields toggle on. It is what makes an absent value come back empty instead of invented.
  4. Choose JSON or CSV, then press Extract. Each document is one request, and progress shows which one is running.
  5. Download the file. Every row has the same columns, in the order you asked for them, whether or not each document had them.

Leaving missing fields empty is the most important setting on this page, and it is on by default. Asked for an invoice number, a language model will produce one whether or not the document has one, and a plausible invented invoice number is far more damaging than a blank cell - it is wrong in a way nobody notices until a payment goes to the wrong place. With the setting on, the instruction is explicit: where a value is not clearly present, return nothing, and never estimate. Turn it off only when you want values the document implies but does not state.

Every row comes back with exactly the keys you asked for, in the order you asked for them, present in the document or not. Missing values become null in JSON and an empty cell in CSV. That is what makes the CSV open cleanly in a spreadsheet: rows of differing shapes make columns drift, and a total in one row ends up under a date in the next. If a model's answer cannot be parsed at all, that row is filled with blanks rather than aborting a twenty-file batch at file eleven.

Temperature is fixed at zero, because extraction is not a creative task and the same invoice should give the same answer twice. Numbers are requested as bare numbers, without currency symbols or thousands separators, and dates as ISO 8601 strings, so the CSV sorts and sums without a cleaning pass. The CSV is quoted to RFC 4180, which stops a supplier name with a comma in it from splitting a column.

The text arrives in reading order with the layout flattened, and that is a real limit on what can be pulled out. A total in the bottom-right corner of an invoice reads fine; a multi-column table of line items arrives as a stream of cells with nothing marking which column each belonged to. It is why the presets ask for header-level fields - totals, dates, parties - rather than line items. For a table you genuinely need as a table, PDF to Excel reconstructs the grid instead.

สิ่งที่เครื่องมือนี้ทำไม่ได้

  • The documents' text is sent to the AI provider you choose. The files themselves are not.
  • Line-item tables do not extract reliably, because the text arrives in reading order without column boundaries. Use PDF to Excel for a table you need as a table.
  • Even with missing fields left empty, a value can be read from the wrong place - a delivery date taken for an invoice date. Spot-check rows against the documents before using them.
  • A scanned PDF has no text to read. Run OCR a PDF over it first.

คำถามที่คนมักถาม

การเว้นช่องที่ไม่พบให้ว่างทำอะไร
มันบอกแบบจำลองว่าค่าที่ไม่ปรากฏชัดเจนในเอกสารต้องกลับมาเป็นค่าว่าง และการแต่งค่าที่ดูสมเหตุสมผลขึ้นมาถือเป็นข้อบกพร่อง ไม่ใช่ความช่วยเหลือ มันเปิดไว้เป็นค่าเริ่มต้น เพราะช่องว่างคือปัญหาที่คุณเห็นได้ทันที ส่วนเลขที่ใบแจ้งหนี้ที่ถูกกุขึ้นคือปัญหาที่คุณจะรู้ตัวช้ากว่ามาก ปิดมันก็ต่อเมื่อคุณต้องการค่าที่เอกสารสื่อถึงแต่ไม่ได้ระบุไว้
ฉันเลือกช่องข้อมูลของตัวเองได้ไหม
ได้ ตั้งชุดสำเร็จรูปเป็นช่องข้อมูลของฉันเองแล้วไล่ชื่อลงไป คั่นด้วยจุลภาคหรือขึ้นบรรทัดใหม่ บรรทัดละชื่ออ่านย้อนกลับง่ายที่สุด ชื่อเหล่านั้นจะกลายเป็นคีย์ของ JSON และหัวคอลัมน์ของ CSV ตรงตามที่คุณพิมพ์ ชื่อแบบ snake_case อย่าง purchase_order_number จึงใช้ได้ดี ชุดสำเร็จรูปที่มีมาให้ก็เป็นรายการชื่อช่องธรรมดาเช่นกัน ชุดใบแจ้งหนี้ขอมาสิบเอ็ดช่อง ตั้งแต่ invoice_number ไปจนถึง payment_terms
ทำได้ครั้งละกี่เอกสาร
ได้สูงสุดยี่สิบไฟล์ แต่ละไฟล์เป็นคำขอของตัวเอง และให้ผลออกมาไฟล์ละหนึ่งแถว ทั้งหมดแบ่งขีดจำกัดตัวอักษรกัน โดยค่าเริ่มต้นคือ 100,000 หารด้วยจำนวนไฟล์ เอกสารที่ยาวกว่าส่วนแบ่งของตัวเองจึงถูกตัด ถ้าเอกสารใดให้คำตอบที่แจงค่าไม่ได้ แถวของมันจะถูกเติมด้วยค่าว่างและงานทั้งชุดจะเดินหน้าต่อ แทนที่จะล้มกลางคัน
ใบแจ้งหนี้ของฉันไปที่ไหน
ตัวไฟล์อยู่บนเครื่องของคุณ ส่วนข้อความในไฟล์ไปยังผู้ให้บริการที่คุณเลือก เอกสารละหนึ่งคำขอ ส่งโดยเบราว์เซอร์ของคุณพร้อมคีย์ของคุณเอง ไม่มีเซิร์ฟเวอร์ของ CekPDF อยู่ในเส้นทางนี้และไม่มีบัญชีผู้ใช้ของ CekPDF จึงไม่มีอะไรตรงนี้ที่เก็บสำเนาไว้ ถ้าเอกสารมีข้อมูลส่วนบุคคล อย่างบัตรประจำตัวและรายการเดินบัญชี ข้อกำหนดที่สำคัญคือข้อกำหนดของผู้ให้บริการของคุณ
ค่าที่ได้เชื่อถือได้แค่ไหน
เชื่อถือได้พอที่จะช่วยให้ไม่ต้องพิมพ์เอง แต่ไม่พอที่จะข้ามการตรวจทาน ช่องระดับหัวเอกสารในใบแจ้งหนี้ที่สะอาด ทั้งเลขที่ วันที่ และยอดรวม กลับมาถูกต้องเกือบทุกครั้ง จุดที่มันพลาดคือความกำกวม เช่น มีสองวันที่บนหน้าและมีแค่อันเดียวที่เป็นวันที่ของใบแจ้งหนี้ ยอดรวมย่อยที่ดูเหมือนยอดรวม หรือตัวเลขภาษีที่ถูกแยกเป็นสองบรรทัด ค่าอุณหภูมิเป็นศูนย์ คำตอบจึงอย่างน้อยก็คงเส้นคงวาระหว่างการรันแต่ละครั้ง

เครื่องมือที่เกี่ยวข้อง