استخراج البيانات
استخرج الحقول نفسها من كومة مستندات، بصيغة JSON أو CSV.
لا يُرفع ملف PDF أبدًا. يُستخرج النص داخل متصفحك، ويُعرض عليك، ولا يُرسَل إلا بعد موافقتك.
ماذا تفعل هذه الأداة
This reads a set of documents and returns one row each, carrying the fields you asked for. Text is extracted in your browser and sent to the AI provider you choose with your own key; the files themselves never leave your device. The output is a JSON array or a CSV, with a _file column so every row says which document it came from.
It exists for the batch: twenty supplier invoices that have to become twenty rows in a ledger, a month of receipts before an expense claim, bank statements you want opening and closing balances from, or a folder of ID documents whose numbers need typing into a system.
كيف تعمل
- Drop up to twenty PDFs onto this page, then choose your provider and paste your API key.
- Pick a preset - invoice, receipt, ID document or bank statement - or choose your own fields and list the names you want, one per line.
- Leave the missing-fields toggle on. It is what makes an absent value come back empty instead of invented.
- Choose JSON or CSV, then press Extract. Each document is one request, and progress shows which one is running.
- Download the file. Every row has the same columns, in the order you asked for them, whether or not each document had them.
Leaving missing fields empty is the most important setting on this page, and it is on by default. Asked for an invoice number, a language model will produce one whether or not the document has one, and a plausible invented invoice number is far more damaging than a blank cell - it is wrong in a way nobody notices until a payment goes to the wrong place. With the setting on, the instruction is explicit: where a value is not clearly present, return nothing, and never estimate. Turn it off only when you want values the document implies but does not state.
Every row comes back with exactly the keys you asked for, in the order you asked for them, present in the document or not. Missing values become null in JSON and an empty cell in CSV. That is what makes the CSV open cleanly in a spreadsheet: rows of differing shapes make columns drift, and a total in one row ends up under a date in the next. If a model's answer cannot be parsed at all, that row is filled with blanks rather than aborting a twenty-file batch at file eleven.
Temperature is fixed at zero, because extraction is not a creative task and the same invoice should give the same answer twice. Numbers are requested as bare numbers, without currency symbols or thousands separators, and dates as ISO 8601 strings, so the CSV sorts and sums without a cleaning pass. The CSV is quoted to RFC 4180, which stops a supplier name with a comma in it from splitting a column.
The text arrives in reading order with the layout flattened, and that is a real limit on what can be pulled out. A total in the bottom-right corner of an invoice reads fine; a multi-column table of line items arrives as a stream of cells with nothing marking which column each belonged to. It is why the presets ask for header-level fields - totals, dates, parties - rather than line items. For a table you genuinely need as a table, PDF to Excel reconstructs the grid instead.
ما لا تستطيع هذه الأداة فعله
- The documents' text is sent to the AI provider you choose. The files themselves are not.
- Line-item tables do not extract reliably, because the text arrives in reading order without column boundaries. Use PDF to Excel for a table you need as a table.
- Even with missing fields left empty, a value can be read from the wrong place - a delivery date taken for an invoice date. Spot-check rows against the documents before using them.
- A scanned PDF has no text to read. Run OCR a PDF over it first.
أسئلة يطرحها الناس
- ماذا يفعل خيار "ترك الحقول المفقودة فارغة"؟
- يخبر النموذج أن أي قيمة غير موجودة بوضوح في المستند يجب أن تعود فارغة، وأن اختراع قيمة تبدو معقولة يُحسب عيبًا لا مساعدة. وهو مفعّل افتراضيًا لأن الخلية الفارغة مشكلة تراها فورًا، أما رقم الفاتورة الملفَّق فمشكلة تكتشفها بعد وقت طويل. لا توقفه إلا حين تريد قيمًا يوحي بها المستند دون أن يذكرها.
- هل يمكنني اختيار حقولي الخاصة؟
- نعم. اضبط الإعداد المسبق على حقولك الخاصة واسرد الأسماء مفصولة بفواصل أو بأسطر جديدة - وسطر لكل اسم أسهل في القراءة. وتصبح تلك الأسماء مفاتيح JSON وعناوين CSV كما كتبتها تمامًا، لذا تعمل الأسماء بصيغة snake_case مثل purchase_order_number جيدًا. والإعدادات المسبقة المدمجة هي أيضًا قوائم حقول عادية؛ فإعداد الفاتورة يطلب أحد عشر حقلًا، من invoice_number إلى payment_terms.
- كم مستندًا يمكنني معالجته في وقت واحد؟
- حتى عشرين، لكل منها طلبه الخاص، وينتج كل منها صفًا واحدًا. وتتقاسم حدّ الأحرف بينها - 100000 مقسومًا على عدد الملفات افتراضيًا - فالمستند الأطول من حصته يُقطع. وإن أنتج مستند جوابًا لا يمكن تحليله، يُملأ صفّه بفراغات وتتابع الدفعة بدل أن تفشل في منتصف الطريق.
- إلى أين تذهب فواتيري؟
- الملفات تبقى على جهازك. أما نصّها فيذهب إلى المزوّد الذي اخترته، طلب واحد لكل مستند، يرسله متصفحك بمفتاحك أنت. ولا يوجد خادم لـ CekPDF في المسار ولا حساب لدى CekPDF، فلا شيء هنا يحتفظ بنسخة. وإن كانت المستندات تحمل بيانات شخصية - وبطاقات الهوية وكشوف الحسابات تحملها - فالشروط التي تهمّ هي شروط مزوّدك.
- ما مدى موثوقية القيم؟
- موثوقة بما يكفي لتوفير الكتابة، لا بما يكفي لتخطّي المراجعة. فحقول مستوى الترويسة في فاتورة نظيفة - الرقم والتاريخ والمجموع - تعود صحيحة في معظم الأحيان. والالتباس هو موضع الزلل: تاريخان على الصفحة وواحد منهما فقط تاريخ الفاتورة، ومجموع فرعي يبدو مجموعًا كليًا، ورقم ضريبة مقسوم على سطرين. ودرجة الحرارة صفر، فالأجوبة على الأقل متسقة بين التشغيلات.