Stateless Tools

Deciding between merge and split

Merge when

You are assembling a submission packet, a contract with its annexes, or monthly reports into one file. The goal is to stop the recipient from having to open several files in the right order.

Extract a range when

You only need to send 3 pages out of 200, or you need to share a document with personal-data pages removed. The original stays untouched and only the needed part becomes a new file.

Watermark when

You want to mark a draft or record a distribution path. But this is a label, not protection — the reason why is further down.

What follows is what actually gets in the way. Keep the PDF merge/split tool open alongside if you want to check as you read.

What quietly disappears when you merge

All the pages make it through. The problem is what lived outside the pages.

Bookmarks and internal links break

A PDF's outline and its internal links ("see chapter 3") point at objects inside the document, not at page numbers. After a merge the page numbers run on, but those references still target the original document, so they jump to the wrong place or vanish. If the outline matters, click a few after merging.

Form fields collide

AcroForm fields are identified by name. Merge two documents made from the same template and names like name or date overlap, so in many viewers typing in one field fills the same-named field on the other pages. This is a real hazard when combining several forms for submission. Fill each document first, flatten it, then merge.

Digital signatures become invalid

A signature covers the exact bytes of the document at signing time. Adding even one page changes the document, so the signature reports the file as modified. Signed PDFs are things you attach, not things you merge.

Page sizes end up mixed

Merging an A4 document with a Letter-size scan leaves both sizes in place. It is invisible on screen but shows up as different margins in print. Normalise sizes first if the result will be printed.

The result can be larger than the sum

Two documents using the same font may end up embedding it twice. Conversely, a good tool deduplicates and produces something smaller than the sum. That is why the output size sometimes surprises you.

Splitting, and the illusion of deletion

Range extraction is safe

Selecting pages 3–5 and producing a new file means the other pages' content is not in the output at all. That is the right method when you must share a document without its personal-data pages.

Covering is not deleting

A PDF where black rectangles were drawn over text still contains that text underneath. Copy and paste reveals it. This has caused repeated public incidents. If information must be, exclude the page entirely or use a tool that performs real redaction.

Metadata comes along

Author name, original file path, authoring software and modification time stay in the document properties. Internal paths often reveal project names when a file leaves the company. Check the properties before sending.

Why text extraction returns nothing

"Extract text from PDF" often produces an empty result, usually for one reason.

Scanned documents contain no text

A PDF from a scanner or phone camera is one image per page. Your eyes see letters; the file holds no character data. An empty extraction almost always means this. What you need is not extraction but OCR.

It extracts, but the order is wrong

A PDF stores text as "draw this glyph at these coordinates". There is no notion of paragraphs or reading order. Two-column layouts therefore interleave line by line, and tables lose their cell boundaries. Rather than using raw output, run it through the text cleaner to fix line breaks and spacing first.

Non-Latin text comes out garbled

If an embedded font lacks the table mapping glyph IDs to Unicode (ToUnicode CMap), extraction cannot recover the characters. This is common in older documents from certain authoring tools. It is a flaw in the source document, not the extractor, so OCR is the reliable fallback.

Why watermarks and passwords are not protection

A watermark is a shape on the page

A PDF watermark is a text or image object layered over the page content. Delete that object in an editor and the original page remains. It does not prevent leaks; it leaves a trace of where the file came from. For that purpose it works fine.

Open passwords differ from permission passwords

An open password genuinely encrypts the content. Permission flags such as "no printing" or "no copying" are a request to the viewer, not encryption, and plenty of tools ignore them. If it must be blocked, use an open password — and do not put that password in the same email.

The limits of browser-side processing

The PDF tools here process files in your browser rather than uploading them, which helps with internal documents. In exchange, very large files or documents with many pages can hit device memory limits, and tasks like repairing a corrupted PDF are better served by dedicated software.

Check before you start

  1. Look for form fields or digital signatures in the documents you plan to merge. If present, attach or flatten instead.
  2. If the outline matters, click a few bookmarks after merging.
  3. For print output, check that page sizes are not mixed.
  4. Handle information that must be by excluding the page, not by drawing over it.
  5. Check the author and file path in document properties before sending externally.
  6. An empty text extraction means a scanned document. Move to OCR.

Next: work in the PDF merge/split tool, run scans through OCR, tidy the output with the text cleaner, and check length with the character counter.