Chat with a PDF
Ask questions about a document and get answers that cite the page.
Your PDF is never uploaded. The text is extracted in your browser, shown to you, and only sent after you approve it.
What this does
This puts a document's text in front of a language model and lets you ask it questions. The text is extracted in your browser and shown to you before anything is sent; the file itself never leaves your device. Every question and every answer is written into a Markdown transcript you can download, so a conversation is something you keep rather than something that vanishes with the tab.
It suits documents you want an answer out of rather than a reading of: a lease where the notice period is somewhere in the small print, a tender with a deadline buried in eighty pages, minutes you need one decision from, or five supplier contracts you want to compare on the same clause.
How it works
- Drop up to five PDFs onto this page. Their text is extracted and shown, with the total that will be sent.
- Choose your provider and paste your API key. You only do this once - the AI tools share it.
- Leave cite page numbers on. The text is sent with page markers and the model is told to reference them.
- Type your question and press Ask. The answer streams in as it is written, and a follow-up carries the earlier turns with it.
- Download the transcript, which holds every question, every answer and the citations.
Citation is the point of the design rather than a feature bolted onto it. A model handed a document will answer confidently whether or not the answer is in there, and nothing about the prose tells you which kind of answer you received. So the extracted text carries a marker at the start of every page, the instruction is to cite a page for every factual claim in the form (p. 12) and to say plainly when the document does not answer the question, and the transcript preserves those references. One answer you can check on page 14 is worth several that read well.
Several documents share one character budget rather than each receiving the full allowance. Ask across five files and each contributes a fifth - 30,000 characters apiece out of 150,000 - because the request that goes out is the combined one, and dividing it is what stops a five-file question quietly costing five times what you agreed to. Each document is labelled with its file name in what the model sees, so an answer can tell you which one it came from.
When a document is longer than its share, the surplus is not sent, and the transcript opens with a line saying the answers do not draw on the whole document. That line is in the file you download rather than in a message that disappears - a month later, reading the transcript again, the caveat is still attached to the answers it applies to.
Creativity defaults to 20%, low enough that the same question twice gives close to the same answer; raising it loosens the wording rather than improving the reasoning. None of that fixes the real risk, which is that a model can misread a table, miss a clause contradicting the one it found, or cite a page that says something adjacent to what it claims. Open the page it cites.
What this tool cannot do
- The documents' text is sent to the provider you choose, together with the conversation so far, on every question.
- Up to five documents at a time, and they share one character budget - a fifth each by default.
- A page citation says where the model looked, not that it read the page correctly. Check anything that matters.
- Scanned PDFs have no text to send. Run OCR a PDF over them first.
Questions people ask
- Are my documents uploaded?
- The files are not. Their text is, to the provider you chose and to nobody else, sent by your browser with your own key. Every question sends the document text plus the conversation so far, so a long thread sends more than the first question did. CekPDF has no server in that path and never sees any of it, and the panel shows you the text before the first request goes out.
- How do I know an answer is really in the document?
- By the page reference. The text is sent with a marker at the start of each page, and the model is instructed to cite one for every factual claim, in the form (p. 12), and to say when the document does not contain the answer instead of inferring it. Those references stay in the transcript. Turning citations off makes answers read more smoothly and makes them much harder to check.
- Can I ask about several documents at once?
- Yes, up to five. Each is labelled by file name in what the model sees, so an answer can say which document it came from - which is what makes comparing the same clause across contracts practical. The five share one character budget rather than each getting the full amount, a fifth each by default, so a five-file question does not cost five times a one-file question.
- My document is longer than the limit. What happens?
- The text is cut at the limit, 150,000 characters by default, and the transcript begins with a line saying the answers do not cover the whole document. Raise the limit under advanced options if you want to send more and are willing to pay for it, or work through a long document in sections using the page range in Summarise a PDF.
- Is there a charge?
- Not from CekPDF - we are not in the payment path and never see your key. Your provider charges per request, and each question re-sends the document text along with everything said so far, so the tenth question in a thread costs more than the first. The run itself counts against your CekPDF allowance, like any other task.