Skip to content
Text Cleaner

Text cleaner that shows you the invisible characters, one by one

Paste text and get a report of every invisible character in it — zero-width spaces, no-break spaces, soft hyphens, direction controls, the byte order mark — each one named, counted, and removed only if you say so. Then tidy the whitespace you can see: collapse runs of spaces, trim each line, drop repeated blank lines, fix the line endings. Anything lossy stays off until you switch it on. Nothing is uploaded or saved.

  • Nothing is uploaded or saved
  • Works offline
Your text
Characters
185
Lines
9
Hidden found
12
Cleaned
Invoice No: INV‑2026–014
Client: Café Müller Sdn Bhd

Terms: net 30 days — “thirty” from the invoice date…
Phone : +603 2000 3000
Address : 12 Jalan Ampang, Kuala Lumpur
Characters
169
Lines
6
Net change
−16
What was in your text
Code pointUnicode nameCountThis runWhere it comes from
U+00A0NO-BREAK SPACE1CleanedArrives from   whenever anything is copied out of a web page.
U+00ADSOFT HYPHEN1CleanedWritten by word processors and PDF exporters to mark where a word may be split.
U+2003EM SPACE1CleanedTypographic space, one em wide. Common in text pasted out of InDesign.
U+2007FIGURE SPACE2CleanedExactly as wide as a digit, so columns of numbers line up. Replacing it will misalign them.
U+200BZERO WIDTH SPACE2CleanedMarks a place a line may break. The single most common invisible character in pasted text.
U+200ELEFT-TO-RIGHT MARK1CleanedForces the following run to read left to right.
U+202FNARROW NO-BREAK SPACE1CleanedFrench typography puts it before ; : ! and ?; some locales use it as a thousands separator.
U+2060WORD JOINER1CleanedThe opposite of a zero-width space: it forbids a line break rather than allowing one.
U+3000IDEOGRAPHIC SPACE1CleanedThe full-width space of Chinese, Japanese and Korean text. Legitimate there.
U+FEFFZERO WIDTH NO-BREAK SPACE (BYTE ORDER MARK)1CleanedThe UTF-8 byte order mark. At the head of a CSV it becomes part of the first column name.
Space runs collapsed: 4 runsLines trimmed: 1 lineExtra blank lines removed: 2 linesLeading and trailing whitespace: 1 characterCurly quotes found, kept: 2 charactersDashes and ellipses found, kept: 4 characters
Clean-up options
Invisible characters

Removed entirely. Watch the count first: U+200D joins an emoji sequence into one picture and U+200C keeps two letters apart in Persian and Devanagari, so removing them is not always a fix.

Turned into one ordinary space (U+0020) rather than deleted, so “Mr Lim” does not become “MrLim”.

Removed. A soft hyphen is invisible until a line breaks on it, which is why Ctrl+F misses the word it sits inside.

Removed. Harmless in genuine Arabic, Hebrew or Jawi text, and worth looking at anywhere else.

Turned into an ordinary line break. These two break JSON and older JavaScript parsers.

Whitespace
Line endings
Lossy — off by default

Both are one-way. They are counted in the report above as “found, kept” so you can see how many there are before deciding.

How it works

  1. 1

    Paste it, then read before you clean

    The report lists every invisible character found in your text with its code point, its official Unicode name, how many times it appears, and where that character usually comes from. The sample loaded in the box carries ten different ones. Read that list first: it is the part of this page you cannot get from a find-and-replace in your own editor, because you cannot search for something you do not know is there.

  2. 2

    Switch off the groups you want to keep

    Five switches, one per family: zero-width characters, unusual spaces, soft hyphens, direction controls, and the two line separators. Unusual spaces become one ordinary space rather than being deleted, so “Mr Lim” does not turn into “MrLim”. Everything else in those groups is removed. Each switch says what it costs, because two of these characters are load-bearing — see the emoji question below.

  3. 3

    Tidy the whitespace you can actually see

    Collapse runs of spaces and tabs to one, trim the ends of every line, allow at most one blank line in a row, and strip whitespace from the top and bottom of the whole document. Each of these is a separate switch with its own count, so “12 space runs collapsed, 4 lines trimmed” tells you what happened instead of leaving you to diff it yourself.

  4. 4

    Decide about line endings deliberately

    Keep is the default and it means keep: a CRLF file stays a CRLF file even after blank lines are removed, because the separators are carried through the clean as data rather than normalised on the way in. Choose LF for anything going into Git or a Unix pipeline, CRLF for a file that has to open cleanly in Notepad. The count tells you how many line endings were actually changed.

  5. 5

    Leave the lossy switches off unless you mean it

    Curly quotes to straight, and dashes and ellipses to ASCII, are the only one-way operations here, so they are off by default and both are listed in the report as “found, kept” until you turn them on. That way you can see how many are in your text and then decide, which is not the same thing as a tool deciding for you.

Frequently asked questions

Where do invisible characters come from in the first place?

Copying. A no-break space arrives from every   on a web page. A zero-width space is inserted by content systems and chat apps to mark where a long word may wrap. A soft hyphen is written by word processors and PDF exporters to record where a word may be split. A byte order mark is added by Excel and by several Windows editors when they save UTF-8. Nobody types any of these; they arrive attached to text that was moved from one program to another, which is why they show up in exactly the documents that were assembled from several sources.

Why not just delete everything that is not plain ASCII?

Because that deletes the text along with the junk. Every accented letter, every Chinese and Tamil name in a Malaysian address book, every emoji, and the em dash a designer put there on purpose all live above U+007F — and so does U+200B. A filter that cannot tell them apart is not a cleaner, it is a shredder, and the damage is invisible until someone downstream notices their own name is spelled wrong. That is why the list this page works from is long and specific: thirty-five named characters in five groups, each one removed on purpose.

Does a regex \s not already find these?

Only some of them, and the split is not where anyone expects. In JavaScript, \s does match U+00A0 NO-BREAK SPACE, the U+2000–U+200A quad family, U+202F, U+205F, U+3000 and even U+FEFF — check /\s/.test("\u00a0") and /\s/.test("\ufeff") in a console and both are true. It does not match U+200B ZERO WIDTH SPACE, U+200C, U+200D, U+2060 WORD JOINER, U+00AD SOFT HYPHEN, or any of the direction controls: /\s/.test("\u200b") is false. Those are precisely the characters people come looking for, which is why this page names them individually instead of leaning on \s.

What is a soft hyphen, and why does Ctrl+F miss the word?

U+00AD is a hyphen that is drawn only when a line happens to break on it. On screen “inter­national” reads as one ordinary word with nothing visible in the middle, but a search for “international” finds nothing, because there is a character between the halves. They survive copy-paste out of a PDF, which is where most of them come from, and they travel into spreadsheets and databases where the mismatch surfaces months later. The report counts them on their own line so you know whether that is your problem before you start rewriting the search term.

What are the direction controls, and should I worry about them?

They tell the renderer which way a run of text reads, and in genuine Arabic, Hebrew, Farsi or Jawi text they are ordinary and necessary. Elsewhere they are worth a second look. U+202E RIGHT-TO-LEFT OVERRIDE reverses the display order of everything after it, so a filename can be made to display with a harmless-looking extension while ending in something else entirely. A 2021 paper from the University of Cambridge called “Trojan Source” showed the same trick applied to source code, where a comment can be made to display as though it ended earlier than it does. Finding one in a filename or a pull request is worth investigating; finding one in an Arabic paragraph is not.

The first column of my CSV will not match anything. Is this why?

Very likely. A UTF-8 byte order mark (U+FEFF) at the head of the file becomes part of the first field, so the header your code reads is not “id” but an invisible character followed by “id”, and every lookup for “id” misses while the file looks perfect on screen. The other symptom is the file opening with “” in front of the first heading — those three characters are the mark’s bytes EF BB BF read as Latin-1 instead of UTF-8. Paste the first few lines here and the report will tell you whether there is one.

Why are smart quotes and dashes off by default?

Because they are the only genuinely lossy operations on this page. Both “ and ” become the same straight ", so nothing in the output records which was there; an em dash becomes two hyphens and an ellipsis becomes three full stops, so the character count moves. In published prose those characters are correct typography and converting them is a downgrade. They are offered because code, CSV files and config files want ASCII, and a curly apostrophe pasted into a shell command or a JSON string is a real bug — but that is a decision about your document, not a clean-up, so you make it rather than the page making it quietly.

Will cleaning break my emoji?

It can, and the report warns you before you press anything. U+200D ZERO WIDTH JOINER is what glues a family emoji together out of separate people, and it sits in the “Zero-width characters” group — remove it and one picture becomes three. If the count for U+200D is above zero and your text contains emoji, leave that group off. The same applies to U+200C ZERO WIDTH NON-JOINER in Persian, Urdu and Devanagari, where it stops two letters merging into a shape that spells a different word. Both are listed with their counts precisely so this is a choice and not an accident.

Is anything I paste stored or sent anywhere?

No. Every replacement on this page runs in JavaScript your browser has already loaded, there is no upload of any kind, and nothing is written to localStorage — reload the tab and the box is empty again. That matters here more than on most pages, because the text people bring to a cleaner is usually an unpublished draft, a client document or a database export on its way somewhere else.