Chuyển đến nội dung

PDF sang văn bản

Trích phần chữ thành một tệp .txt thuần.

Xử lý trên thiết bị của bạn

Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.

Công cụ này làm gì

This extracts the words from a PDF and writes them into a plain .txt file. A PDF does not store text as sentences - it stores glyphs at coordinates - so reading order, word boundaries and paragraph breaks are all reconstructed from the geometry of the page. Every extractor guesses that slightly differently, which is why the same document gives different output in different tools, and why there are two modes here rather than one.

Lifting a quotation out of a report without retyping it, feeding a document to a script or a language model, searching a book you only have as a PDF, or checking what a page actually contains when a table detector insists it found nothing.

Cách hoạt động

  1. Drop the PDF onto this page. One file at a time, up to 300 MB.
  2. Choose reading order for prose you intend to paste somewhere, or layout to keep columns and tables lined up with spaces.
  3. Decide whether pages are separated by a marker line and whether page numbers are printed into the file.
  4. Leave rejoin hyphenated words on so a word broken across a line break comes back together.
  5. Under advanced options, set a page range, or switch the encoding to UTF-8 with a BOM if Excel or Notepad will open the file.
  6. Press Convert to text and the .txt downloads on its own.

The two modes solve opposite problems. Reading order joins the lines of a paragraph back into flowing prose, which is what you want in an email, a script or a quotation. Layout mode keeps every line where it was on the page, padding with spaces so columns stay aligned - right for a table, a code listing or a form, and wrong for prose, where the padding becomes ragged whitespace you then strip out. If the output looks strange, the mode is usually why.

Rejoining hyphenated words is more careful than a search-and-replace. A hyphen at the end of a line is joined only when the character before it and the first character of the next line are both lowercase, so a word broken as inter- and national comes back as international, while a compound such as self-service, or a hyphen followed by a capitalised name, is left as written. The rule errs towards leaving a real hyphen alone rather than towards editing your text.

A scanned document is detected before any work happens. The first five pages are sampled for a text layer, and when there is none the tool stops and points at OCR, since a scan is a picture of words with nothing in it to extract. Reporting no text found would be technically true and practically useless.

The encoding option is small and saves a familiar afternoon. UTF-8 without a byte order mark is what a script, a git repository or a text editor wants. UTF-8 with a BOM is what makes Excel and Notepad recognise the file as UTF-8 instead of guessing a legacy code page and mangling every accented character. If names come out as strings of symbols, the BOM is what was missing.

Extraction runs in a Web Worker and the text file is written there too, so a confidential document is never sent anywhere to have its words read. The result reports the page count, the character and word counts and how many pages came out empty - often the quickest sign that part of a document is scanned while the rest is not.

Công cụ này không làm được gì

  • Columns, sidebars, headers and footnotes come out in the order the reconstruction decides on, which is not always the order a person reads them in. Layout mode is often clearer for those pages.
  • A scanned page contains no text to extract. That is detected and reported with a link to OCR rather than writing an empty file.

Câu hỏi thường gặp

Thứ tự đọc và bố cục khác nhau ở đâu?
Thứ tự đọc dựng lại các đoạn văn, nối các dòng của mỗi đoạn thành văn bản trôi mà bạn dán được vào bất cứ đâu. Bố cục giữ hình dạng trang bằng cách đệm dấu cách, nên cột, bảng và đoạn mã vẫn thẳng hàng trong một khung có độ rộng cố định. Dùng thứ tự đọc cho văn xuôi và bố cục khi vị trí ngang của chữ mang ý nghĩa.
Không có gì được trích ra từ PDF của bạn. Vì sao?
Gần như chắc chắn là vì nó là một bản quét. Một trang quét chứa ảnh chụp của chữ, chứ không phải chữ, nên không có lớp chữ nào để trích. Năm trang đầu được lấy mẫu trước khi bắt đầu bất kỳ việc gì, và nếu không có lớp nào thì công cụ dừng lại và hướng bạn sang OCR PDF, vốn biến bức hình thành chữ thật để rồi bạn trích ra được.
Chữ ra bị lộn xộn. Có sửa được không?
Hãy thử chế độ bố cục trước - kết quả lộn xộn thường nghĩa là một trang nhiều cột mà các dòng bị nối ngang qua các cột. Bố cục giữ mỗi cột ở đúng chỗ của nó, dễ đọc hơn và dễ dọn dẹp hơn. Một số tài liệu còn đặt đầu trang, chân trang và ghi chú bên lề theo một thứ tự không liên quan đến cách chúng được đọc, và không bộ trích nào có thể biết chắc hơn được.
Ký tự có dấu hiện sai khi bạn mở tệp. Giờ làm sao?
Hãy đổi bảng mã sang UTF-8 có BOM rồi chuyển lại. Dấu thứ tự byte cho Excel và Notepad biết tệp là UTF-8 thay vì để chúng đoán một bảng mã cũ, vốn là thứ biến một chữ có dấu thành hai ký hiệu. Hãy để nguyên UTF-8 thuần cho tập lệnh, trình soạn thảo và quản lý phiên bản.
Tài liệu có bị tải lên để trích chữ không?
Không. PDF được phân tích và tệp văn bản được tạo trong trình duyệt của bạn, và không có gì được truyền đi ở bất kỳ thời điểm nào - điều đáng biết khi tài liệu là một hợp đồng hay một hồ sơ y tế. Bạn có thể theo dõi thẻ mạng của công cụ dành cho nhà phát triển trong lúc chuyển đổi để xác nhận điều đó, và công cụ hoạt động ngoại tuyến.

Công cụ liên quan