Chuyển đến nội dung

OCR một PDF

Thêm một lớp văn bản tìm kiếm được vào một tài liệu quét.

Xử lý trên thiết bị của bạn

Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.

Công cụ này làm gì

A scanned PDF is a picture of a document, so searching it finds nothing and selecting a word selects nothing. This renders each page, recognises the words on it, and draws them back over the image invisibly, so the page looks as it did and the words are suddenly there. Recognised pages are rebuilt from that rendering: they come out grayscale at the resolution you chose, and any links, form fields or annotations they carried do not survive. Pages that already have real text are left alone, byte for byte.

Reach for it whenever a document arrived as pictures rather than words: a contract someone scanned and emailed, a receipt photographed on a phone, a fifty-page report from an office machine, or an old paper you need to quote from without retyping it.

Cách hoạt động

  1. Drop the scan onto this page. PDFs work, and so do JPEG, PNG, WebP and TIFF images - an image becomes a one-page PDF first.
  2. Choose the language. English and Indonesian are installed, and the combined mode recognises both at once in a document that mixes them.
  3. Leave the mode on only pages without text unless you know for certain that every page is a scan.
  4. Set the resolution if you need to. 300 dpi suits ordinary print; go towards 400 for very small type, or down towards 150 for a long document on a slow machine.
  5. Choose whether you want the searchable PDF, a plain text file, or both, then press Run OCR.

Recognition costs something on the pages it touches, and it is better to know before than after. Each recognised page is rendered to a grayscale image at your chosen resolution, and the new page is that image with the invisible words drawn over it. Colour on those pages is gone, and their links, form fields and annotations are not carried across. Skipped pages are copied untouched. The file usually grows, sometimes a lot, because a 300 dpi page image is bigger than the compressed scan it replaced - Compress PDF is the normal second step.

The words have to land in the right place or a text layer is worse than none. The recogniser returns each word's box in the pixels of the rendered image, counting down from the top, while a PDF page counts up from the bottom, so every box is scaled by the render ratio and flipped. Words scored below 40 out of 100 are left out entirely: a confidently wrong word sends search to the wrong page, while a gap only sends it nowhere.

Nothing is uploaded, and that includes the recogniser itself. tesseract.js fetches its WebAssembly core and its language data from a third-party CDN unless told otherwise, which would have meant your scan quietly announcing itself to a third party the first time anyone here ran OCR, so all of it is served from this site instead. The cost is a one-off download. The benefit is that recognition involves nobody else at all, which you are welcome to check in the network tab.

Resolution decides both how accurate the reading is and how long you wait for it. Below about 200 dpi, thin serifs and small print break up and accuracy falls away quickly. Above 300 the gains get small while render time and output size keep climbing. 300 is the default because that is roughly where the curve flattens for ordinary body text.

Công cụ này không làm được gì

  • Handwriting is not recognised. The engine is trained on printed type, and a handwritten page comes back as scattered nonsense rather than words.
  • Only English and Indonesian models are installed, so a document in another language cannot be read properly here.
  • Recognised pages are rebuilt as grayscale images, so colour is lost on them and their links, form fields and annotations do not survive.
  • The output is a machine's reading, not a proofread transcript. Faint scans, unusual typefaces, tables and tight columns all produce mistakes.

Câu hỏi thường gặp

Nhận dạng chính xác đến mức nào?
Trên một bản quét 300 dpi sạch của văn bản in, một trang thường trở về chỉ với vài lỗi, và kết quả báo cáo điểm chắc chắn trung bình để bạn thấy công cụ tự đánh giá công việc của mình ra sao. Bản photocopy mờ, trang bị lệch, cột chật và phông chữ trang trí đều kéo con số đó xuống. Các từ có điểm dưới 40 trên 100 bị bỏ đi thay vì đoán.
Nó đọc được những ngôn ngữ nào?
Tiếng Anh và tiếng Indonesia, cộng thêm một chế độ kết hợp xử lý cả hai trong cùng một tài liệu - hữu ích cho một hợp đồng in kèm bản dịch của nó. Mô hình ngôn ngữ là thứ làm cho nhận dạng hoạt động được, và chỉ có hai mô hình đó được kèm theo, nên một tài liệu bằng ngôn ngữ khác không phải là việc công cụ này làm tốt được.
Vì sao nó tải về vài megabyte trong lần đầu?
Đó là bộ nhận dạng và dữ liệu ngôn ngữ của nó, phục vụ từ trang này chứ không phải từ một CDN của bên thứ ba để việc chạy OCR không âm thầm liên hệ với ai đó khác. Tiếng Anh khoảng 11 MB, tiếng Indonesia khoảng 4 MB. Trình duyệt của bạn lưu chúng vào bộ đệm, nên lần chạy thứ hai bắt đầu ngay, và từ đó trở đi OCR hoạt động mà không cần kết nối nào cả.
chỉ các trang không có văn bản thực ra làm gì?
Nó xem xét mỗi trang để tìm một lớp văn bản có sẵn và chỉ nhận dạng những trang không có lớp nào. Hầu hết các tài liệu quét chứa ít nhất một trang chưa từng được quét, và chạy OCR lên nó sẽ thay văn bản chính xác bằng một bản gần đúng. Nếu cả tài liệu đã có văn bản, công cụ sẽ nói vậy thay vì chạy và tạo ra thứ tệ hơn thứ bạn bắt đầu.
Nó có đọc được chữ viết tay không?
Không. Bộ máy được huấn luyện trên chữ in, và chữ viết tay trở về dưới dạng nhiễu - những chữ cái giống hình dạng chứ không giống các từ. Một chữ ký, một ghi chú ở lề hay một biểu mẫu điền bằng tay sẽ không được nhận dạng. Văn bản in trên cùng trang thì vẫn được.
Bản quét có được tải lên đâu không?
Không. Việc dựng trang diễn ra trong một Web Worker trong trình duyệt của bạn và việc nhận dạng trong một Web Worker khác, và ngay cả lõi WebAssembly lẫn dữ liệu ngôn ngữ cũng đến từ chính origin của trang này chứ không phải một CDN. Hãy kiểm chứng trực tiếp: nạp trang một lần, tắt kết nối, và chạy một bản quét qua.

Công cụ liên quan