HTML sang PDF
Chuyển một trang web đã lưu hoặc một tệp văn bản sang PDF.
Công cụ này chạy hoàn toàn trong trình duyệt của bạn. Tệp của bạn không bao giờ được tải lên, và bạn có thể tự kiểm chứng điều đó trong tab mạng của trình duyệt. Tự kiểm chứng: mở tab mạng của trình duyệt và theo dõi. Bạn sẽ thấy một yêu cầu nhỏ hỏi xem bạn còn tác vụ nào không - một tên công cụ và một mã băm, không bao giờ là tệp.
Công cụ này làm gì
This takes a saved .html or .htm file, or a plain .txt file, and lays its content out as a PDF. Headings, paragraphs, lists, tables and links are read from the document's structure and re-flowed onto the paper you choose. CSS is not executed, so the result reads correctly and does not look identical to the page in a browser. There is no field for a web address, and that is a decision rather than an omission: fetching a page for you would mean an outbound request from a tool that promises none.
Use it on a page you have already saved: an article kept for offline reading, a receipt or booking confirmation a site would only show you in the browser, documentation you want on paper, or an export from a tool that writes HTML and nothing else.
Cách hoạt động
- Save the page first. In any desktop browser, press Ctrl+S or Command+S and choose Web page, HTML only - the complete option saves a folder of assets this tool does not need.
- Drop the .html file onto this page. Plain .txt files work too, up to 20 files at a time.
- Choose the paper size, orientation and margins. Wide margins suit anything you intend to annotate by hand.
- Leave keep hyperlinks on so the page's links stay clickable in the PDF.
- Turn on page numbers for anything you plan to print and hand round.
- Press Convert and the PDF downloads.
The page is re-flowed from its structure, not rendered. A browser turns HTML into a picture by executing CSS - float, flex, grid, absolute positioning, media queries, web fonts - and shipping a layout engine into a browser tab to do that a second time is not a small addition. Instead the document is read as headings, paragraphs, lists, tables and links, and set with the same typography as every other conversion here. Two columns become one, and the article you wanted comes out readable.
That trade has a side effect people tend to like. Navigation, banners and sidebars are flattened into the flow along with everything else, so they arrive as plain lists and paragraphs rather than as furniture down the edge of the page. This is not reader-mode extraction and does not claim to be - text that was decoration is still text - but the article stops being a narrow column squeezed between two others.
Images are not embedded. A saved .html file usually points at its pictures rather than containing them, and going out to fetch them would be exactly the outbound request this tool exists to avoid. Text, tables, lists and links come through; pictures do not, including ones written into the file as data URIs. If the pictures are the point, save them separately and use Images to PDF.
The file is treated as hostile. A saved web page can carry script tags, inline event handlers and javascript: links, and all three are stripped in the parser before anything reaches the layout engine. Nothing in the page runs, and no link that ends up in the PDF can execute anything.
The missing URL box is worth one more sentence. To fetch a page for you, this site would have to make the request itself or route it through a proxy, and either way somebody's server learns which page you were reading. For a tool whose whole claim is that your document never leaves your device, that is not a trade worth making. Saving the page yourself takes one keystroke and keeps the request in your own browser, where it already was.
Công cụ này không làm được gì
- CSS is not executed, so colours, columns, positioning and web fonts are lost. The page is re-flowed as text, tables and lists.
- There is no URL field. Save the page to a file first; fetching it here would mean an outbound request on your behalf.
- Images are not embedded, including ones written into the file as data URIs. Text, tables and links come through.
Câu hỏi thường gặp
- Bạn dán một địa chỉ web thay cho tệp được không?
- Không, và đây là điều chủ ý bỏ đi chứ không phải một tính năng chưa ai kịp làm. Tải trang về sẽ đồng nghĩa với việc trang này, hoặc một proxy CORS ở giữa, gửi một yêu cầu thay cho bạn - điều đó sẽ cho máy chủ của ai đó biết bạn đang đọc trang nào và phá vỡ lời hứa duy nhất mà công cụ này đưa ra. Hãy lưu trang bằng Ctrl+S rồi thả tệp vào đây.
- Vì sao PDF không giống trang web?
- Vì các biểu định kiểu (stylesheet) không được thực thi. Cấu trúc của tài liệu được đọc - tiêu đề, đoạn văn, danh sách, bảng, liên kết - rồi được dàn theo kiểu chữ của chính trang này. Màu sắc, cột, cách định vị và phông web không được giữ lại. Thứ còn lại là nội dung của trang, dễ đọc, trên khổ giấy bạn chọn.
- Hình ảnh trên trang có được giữ lại không?
- Không. Hình ảnh không được nhúng, vì một tệp .html đã lưu thường chỉ tham chiếu đến ảnh chứ không chứa chúng, và tải chúng về sẽ đồng nghĩa với một yêu cầu đi ra ngoài. Văn bản, bảng, danh sách và liên kết đều được chuyển; hình ảnh bị bỏ ra, kể cả những ảnh được ghi thẳng vào tệp dưới dạng data URI.
- Làm sao để lưu một trang web để có thể chuyển đổi?
- Nhấn Ctrl+S trên Windows hoặc Command+S trên máy Mac và chọn "Web page, HTML only" thay vì "Complete" - lựa chọn complete sẽ tạo ra một thư mục tài nguyên mà công cụ này không dùng đến. Trên điện thoại, hãy tìm tuỳ chọn lưu-thành-tệp hoặc tải xuống trong menu chia sẻ của trình duyệt. Rồi thả tệp đã lưu vào trang này.
- Có bất cứ điều gì về trang bị gửi đi đâu không?
- Không. Tệp được giải mã, phân tích và dàn trang hoàn toàn bên trong trình duyệt của bạn, và không có ô nhập URL chính là để không có yêu cầu nào được gửi thay cho bạn. Thứ duy nhất mà trang này tải về là phông Noto, từ chính máy chủ của trang, và chỉ khi bật nhúng Unicode.