Tóm tắt một PDF
Nắm nội dung chính của một tài liệu dài, dùng khóa AI của riêng bạn.
PDF của bạn không bao giờ được tải lên. Văn bản được trích xuất trong trình duyệt, hiển thị cho bạn, và chỉ được gửi sau khi bạn chấp thuận.
Công cụ này làm gì
This extracts the document's text in your browser, sends it to the AI provider you choose using your own key, and writes the answer out as a Markdown file. The file itself never leaves your device, and the panel shows you the exact text and its size before anything is sent. There is no CekPDF server in that path, which also means no CekPDF account and no CekPDF bill - your provider charges you directly.
Use it when a document is longer than the time you have for it: a hundred-page tender you need to decide whether to bid on, minutes from a meeting you missed, an annual report where you only want the shape of the thing, or a contract you would like a map of before reading the clauses properly.
Cách hoạt động
- Drop the PDF onto this page. Its text is extracted and shown to you, with a character count and a rough token estimate.
- Choose your provider and paste your API key. The key is kept in this browser's local storage and sent to nobody but the provider it belongs to.
- Pick a length - roughly 120 words for brief, 350 for standard, 900 with a section per topic for detailed - and a style: bullets, paragraphs, or a nested outline.
- Set the output language if you want something other than the document's own, and narrow the page range if only part of it matters.
- Press Summarise. The answer streams in as the model writes it, and Cancel stops the request.
Long documents are summarised in two passes. The text is cut into chunks of about 60,000 characters, each chunk is summarised on its own, and a final request writes one summary from those notes. The alternative most tools take - truncate to fit and summarise whatever survived - produces a confident account of the first thirty pages that reads exactly like an account of the whole document. This one costs more requests and reports how many it made.
There is a ceiling on how much text goes out, 120,000 characters by default, and it exists to stop someone accidentally posting a 900-page book along with the bill that would come back. When a document exceeds it, the text is cut there and the summary file opens with a line saying so - in the output you keep, not in a notification that disappears. Raise the limit under advanced options if you would rather pay for the whole thing.
The instruction the model receives is deliberately narrow: use only what the document says, note where it is unclear rather than filling the gap, and skip the preamble about being an assistant. Temperature is fixed at 0.2, low enough that the same document twice gives close to the same summary. None of that makes a summary trustworthy by itself, so treat it as a way into the document rather than a replacement for the paragraph you are about to quote.
This is one of the few pages here with an amber badge instead of a green one, because something does leave your device. To be exact: the file does not, the images in it do not, and the extracted text does - posted by your own browser, straight to the endpoint you named, with your key in the header. Nothing of ours sits between the two, so there is nothing on this side that could log it or train on it.
Công cụ này không làm được gì
- The document's text is sent to the AI provider you choose. The file itself is not, but its words are.
- You need your own API key, and your provider bills you for every request the tool makes.
- Long documents are summarised in chunks and then summarised again, so a detail mentioned once in passing may not reach the final summary.
- Tables and columns are flattened into reading order when the text is extracted, so a summary of a document that is mostly tables will be weaker than one of a document that is mostly prose.
Câu hỏi thường gặp
- Tài liệu của tôi có được tải lên không?
- Tệp thì không. Văn bản của nó thì có, và chỉ tới nhà cung cấp bạn đã chọn. Việc trích xuất diễn ra trong trình duyệt của bạn, bảng hiển thị văn bản và kích thước của nó trước khi bất cứ gì được gửi đi, và yêu cầu sau đó đi thẳng từ trình duyệt của bạn tới điểm cuối của nhà cung cấp bằng khóa của bạn. CekPDF không có máy chủ nào trong đường đi đó và không thể thấy, ghi lại hay giữ bất cứ phần nào.
- Vì sao tôi phải mang khóa API của riêng mình?
- Vì không có máy chủ CekPDF, và một dịch vụ tóm tắt tài liệu cho bạn sẽ cần một máy chủ - cùng với một tài khoản, một hạn mức, và một bản sao văn bản của bạn đi qua nó. Mang một khóa giữ cho sự sắp đặt trung thực: trình duyệt của bạn nói chuyện với nhà cung cấp của bạn và không ai khác dính líu. Khóa nằm trong bộ nhớ cục bộ của trình duyệt này, dùng chung giữa các công cụ AI nên bạn chỉ nhập một lần.
- Nó tốn bao nhiêu?
- Không tốn gì cho CekPDF; ở đây không có tài khoản và không có hạn mức. Nhà cung cấp của bạn tính phí bạn cho các token trong yêu cầu và trong câu trả lời, và bảng ước lượng kích thước trước khi bạn gửi - khoảng một token cho mỗi bốn ký tự. Một tài liệu dài tốn hơn một yêu cầu, vì nó được tóm tắt theo từng đoạn trước.
- Nó xử lý một tài liệu 300 trang như thế nào?
- Bằng cách cắt văn bản thành các đoạn khoảng 60.000 ký tự, tóm tắt từng đoạn, rồi viết một bản tóm tắt duy nhất từ những ghi chú đó. Số yêu cầu được báo cáo cùng kết quả. Nếu tài liệu dài hơn giới hạn ký tự, mặc định 120.000, phần dư không được gửi và bản tóm tắt bắt đầu bằng một dòng cho bạn biết điều đó thay vì ngụ ý rằng nó bao trùm mọi thứ.
- Tôi có thể tin bản tóm tắt không?
- Hãy coi nó như một lần đọc đầu tiên nhanh, không phải như một trích dẫn. Mô hình được yêu cầu chỉ dùng những gì tài liệu nói và đánh dấu chỗ chưa rõ, và nhiệt độ thấp nghĩa là cùng một tài liệu cho câu trả lời gần như nhau hai lần. Nó vẫn có thể đọc sai một bảng hay cho một nhận xét thoáng qua nhiều trọng lượng hơn nó xứng đáng. Nếu một câu trong bản tóm tắt quan trọng, hãy tìm nó trong tài liệu trước khi dựa vào.
- Nó báo PDF của tôi không có văn bản. Giờ làm sao?
- Tài liệu là một bản quét - một ảnh của một trang, không có từ nào trong đó để trích xuất. Hãy chạy nó qua OCR một PDF trước, vốn nhận dạng văn bản và thêm nó lại thành một lớp tìm kiếm được, rồi mang kết quả tới đây. Việc kiểm tra diễn ra trước khi bất kỳ yêu cầu nào được thực hiện, nên một bản quét không tốn gì để phát hiện.