Trích xuất dữ liệu
Rút cùng các trường ra khỏi một chồng tài liệu, dưới dạng JSON hoặc CSV.
PDF của bạn không bao giờ được tải lên. Văn bản được trích xuất trong trình duyệt, hiển thị cho bạn, và chỉ được gửi sau khi bạn chấp thuận.
Công cụ này làm gì
This reads a set of documents and returns one row each, carrying the fields you asked for. Text is extracted in your browser and sent to the AI provider you choose with your own key; the files themselves never leave your device. The output is a JSON array or a CSV, with a _file column so every row says which document it came from.
It exists for the batch: twenty supplier invoices that have to become twenty rows in a ledger, a month of receipts before an expense claim, bank statements you want opening and closing balances from, or a folder of ID documents whose numbers need typing into a system.
Cách hoạt động
- Drop up to twenty PDFs onto this page, then choose your provider and paste your API key.
- Pick a preset - invoice, receipt, ID document or bank statement - or choose your own fields and list the names you want, one per line.
- Leave the missing-fields toggle on. It is what makes an absent value come back empty instead of invented.
- Choose JSON or CSV, then press Extract. Each document is one request, and progress shows which one is running.
- Download the file. Every row has the same columns, in the order you asked for them, whether or not each document had them.
Leaving missing fields empty is the most important setting on this page, and it is on by default. Asked for an invoice number, a language model will produce one whether or not the document has one, and a plausible invented invoice number is far more damaging than a blank cell - it is wrong in a way nobody notices until a payment goes to the wrong place. With the setting on, the instruction is explicit: where a value is not clearly present, return nothing, and never estimate. Turn it off only when you want values the document implies but does not state.
Every row comes back with exactly the keys you asked for, in the order you asked for them, present in the document or not. Missing values become null in JSON and an empty cell in CSV. That is what makes the CSV open cleanly in a spreadsheet: rows of differing shapes make columns drift, and a total in one row ends up under a date in the next. If a model's answer cannot be parsed at all, that row is filled with blanks rather than aborting a twenty-file batch at file eleven.
Temperature is fixed at zero, because extraction is not a creative task and the same invoice should give the same answer twice. Numbers are requested as bare numbers, without currency symbols or thousands separators, and dates as ISO 8601 strings, so the CSV sorts and sums without a cleaning pass. The CSV is quoted to RFC 4180, which stops a supplier name with a comma in it from splitting a column.
The text arrives in reading order with the layout flattened, and that is a real limit on what can be pulled out. A total in the bottom-right corner of an invoice reads fine; a multi-column table of line items arrives as a stream of cells with nothing marking which column each belonged to. It is why the presets ask for header-level fields - totals, dates, parties - rather than line items. For a table you genuinely need as a table, PDF to Excel reconstructs the grid instead.
Công cụ này không làm được gì
- The documents' text is sent to the AI provider you choose. The files themselves are not.
- Line-item tables do not extract reliably, because the text arrives in reading order without column boundaries. Use PDF to Excel for a table you need as a table.
- Even with missing fields left empty, a value can be read from the wrong place - a delivery date taken for an invoice date. Spot-check rows against the documents before using them.
- A scanned PDF has no text to read. Run OCR a PDF over it first.
Câu hỏi thường gặp
- Để trống các trường thiếu làm gì?
- Nó bảo mô hình rằng một giá trị không hiện diện rõ ràng trong tài liệu phải trở về trống, và rằng bịa ra một giá trị trông có vẻ hợp lý được tính là một khiếm khuyết chứ không phải là hữu ích. Nó bật mặc định vì một ô trống là một vấn đề bạn thấy ngay còn một số hóa đơn bịa ra là thứ bạn phát hiện muộn hơn nhiều. Chỉ tắt nó khi bạn muốn các giá trị mà tài liệu ngụ ý nhưng không bao giờ nêu.
- Tôi có thể chọn các trường của riêng mình không?
- Có. Đặt bộ định sẵn thành các trường của riêng bạn và liệt kê các tên, ngăn cách bằng dấu phẩy hoặc dấu xuống dòng - mỗi tên một dòng là dễ đọc lại nhất. Những tên đó trở thành các khóa JSON và các tiêu đề CSV đúng như bạn đã gõ, nên các tên kiểu snake_case như purchase_order_number hoạt động tốt. Các bộ định sẵn có sẵn cũng là các danh sách trường thông thường; bộ hóa đơn hỏi mười một trường, từ invoice_number đến payment_terms.
- Tôi có thể làm bao nhiêu tài liệu cùng lúc?
- Tối đa hai mươi, mỗi tài liệu một yêu cầu riêng, tạo ra một hàng mỗi cái. Chúng dùng chung giới hạn ký tự với nhau - mặc định 100.000 chia cho số tệp - nên một tài liệu dài hơn phần của nó sẽ bị cắt. Nếu một tài liệu tạo ra một câu trả lời không phân tích được, hàng của nó được điền các ô trống và mẻ vẫn tiếp tục thay vì thất bại giữa chừng.
- Các hóa đơn của tôi đi đâu?
- Các tệp ở lại trên thiết bị của bạn. Văn bản của chúng đi tới nhà cung cấp bạn đã chọn, mỗi tài liệu một yêu cầu, được trình duyệt của bạn gửi kèm khóa của riêng bạn. Không có máy chủ CekPDF nào trong đường đi và không có tài khoản CekPDF, nên không có gì ở đây giữ một bản sao. Nếu các tài liệu mang dữ liệu cá nhân - thẻ căn cước và sao kê ngân hàng thì có - các điều khoản quan trọng là của nhà cung cấp của bạn.
- Các giá trị đáng tin cậy đến mức nào?
- Đủ tin cậy để đỡ phải gõ, không đủ tin cậy để bỏ qua việc kiểm tra. Các trường ở cấp tiêu đề trên một hóa đơn sạch - số, ngày, tổng - trở về đúng phần lớn thời gian. Sự mơ hồ là nơi nó trượt: hai ngày trên trang và chỉ một trong số đó là ngày hóa đơn, một khoản tạm tính trông giống một khoản tổng, một con số thuế chia trên hai dòng. Nhiệt độ bằng không, nên các câu trả lời ít nhất là nhất quán giữa các lần chạy.