Favorite places of your favorite city
Home
Add to preferred sources on Google

PDF Text Cleaner

Fix unwanted line breaks and spacing in text copied from a PDF.

Paste or type text copied from a PDF.

Updates automatically. Your original stays available.

How it works & input limits

Paste plain text to clean it automatically. Changing settings updates the result; typing updates it after a short pause. The original stays available until you edit or clear it.

Cleaning limit: 500,000 UTF-16 code units, not visible characters. Longer text stays in the editor but cannot be processed.

Text is processed locally in your browser. Drafts are not saved or uploaded by this tool. Shared site services may still make network requests.

Cleaning settings

Line-wrap repair is a heuristic. Blank lines and recognized lists, code, tables and standalone links are protected. Use spacing only for uncertain layouts.

Automatic omits the space only between adjacent Han characters; otherwise it inserts one space. It does not detect language.

Paste text to begin. Cleaning starts automatically.

How to clean copied PDF text

  1. Paste the original plain text to start cleaning automatically. Choose paragraph repair or spacing only, and enable optional changes only when needed.
  2. Review the updated excerpts and decide whether each uncertain hyphen should be removed, retained or left unchanged. Each decision updates the result automatically.
  3. Copy the cleaned result when processing finishes. Changes to the source or settings update the result and reset hyphen choices; typing waits for a short pause. Clear removes both texts and restores default settings.

What this cleaner changes

PDF copy-paste often inserts a newline at each visual line ending. This tool joins eligible lines, but a newline alone cannot reveal the author’s intent. It preserves blank paragraphs, recognized lists, fenced or indented code, tabular lines and standalone URLs, emails and paths. Sentence-ending punctuation also prevents a join. Headings, poetry and citations are not detected perfectly; review changes or use spacing only.

A line-final hyphen can split a word or belong to a compound. Each occurrence has its own decision and stays unchanged by default. Repairs always start from your current original, so resetting choices never compounds earlier edits. Cleaning an already cleaned result can expose new joins; it is not guaranteed to be idempotent.

Automatic joining omits spaces only between adjacent Han characters. It is not language detection and cannot infer every Japanese, Korean, Thai or Khmer boundary. Optional changes cover only U+00AD, ff/fi/fl/ffi/ffl, U+00A0/U+202F and repeated ASCII spaces. Protected structures remain exact. Tabs, trailing hard-break spaces, combining accents, emoji, shaping controls and other Unicode characters are preserved.

Examples: before → after

A wrapped ␊ sentence. → A wrapped sentence. (␊ denotes a line break, without the surrounding display spaces.)

inter-␊national → unchanged by default; choose removal for international, or retention for inter-national. state-␊of-the-art → state-of-the-art with retention.

中文␊文本 → 中文文本 in Automatic mode. file → file only with ligature expansion. First paragraph.␊␊Second paragraph. keeps both line breaks.

Local processing and limits

Processing is local to your browser; this tool does not upload or persist drafts. Shared site analytics and advertising can still make network requests. Editors and review excerpts use the site’s masking attributes. The 500,000 UTF-16-unit limit is not a visible-character count; emoji may use more than one unit. Oversized input is retained without truncation.

Frequently asked questions

No. Paste plain text only. Scanned PDFs need OCR elsewhere first. This tool cannot reconstruct lost columns, tables or reading order, and does not remove page numbers, repeated headers or citations.

A wrapped compound and a typesetting split can look identical. Keep the original boundary when uncertain; review each occurrence independently.

Read the joined lines in context. Use Character Counter for length, Text Size Calculator for encoded size, or SMS Character Counter for message segments using the related tools below.

SMS Character Counter & Segment Calculator

See how your message is encoded, where its segments end and what it may cost. Your text stays in your browser.

Open toolSMS Character Counter & Segment Calculator

Social Media Character Counter

Check platform-specific limits and split posts without losing text.

Open toolSocial Media Character Counter

Character Counter

Count characters, words, sentences and paragraphs. Check word density, reading time and UTF-8 size, all in your browser.

Open toolCharacter Counter

Text Size Calculator

Compare text encodings and payload sizes in your browser.

Open toolText Size Calculator

Fancy Text Converter

Convert supported fancy Unicode letters to normal text locally. Keep case, accents, emoji and your original text.

Open toolFancy Text Converter