Skip to content
Remove Duplicate Lines

Remove duplicate lines — or keep only the duplicates

Paste a list and get it back with the repeats gone, or with only the repeats showing, which is often the more useful half. You decide what counts as the same line: ignore case, ignore the spaces around it, keep the first occurrence or the last. Then sort it properly — A→Z here puts item2 before item10 and files ä beside a, which a plain sort does not. Nothing is uploaded or saved.

  • Nothing is uploaded or saved
  • Works offline
ModeEvery line once, in the order you pasted them
Your list
Lines in
10
Distinct
6
Repeated
3
Deduplicated
Ahmad Faizal
item2
aisyah.karim@example.com
Tan Wei Ming
item10
item1
Lines out
6
Removed
4
Blank kept
0
What counts as the same line
Which copy survives

Only changes the output when the duplicates are not identical — which is exactly what ignoring case or trimming causes. In data ordered oldest first, the last occurrence is the current state of the row.

Sort

Lines stay where they were

A→Z and Z→A compare with Intl.Collator using numeric ordering, so item2 comes before item10 and “ähnlich” files beside “a” rather than after “z”. Sorting happens after the duplicates are handled, so the counts above never change when you re-sort.

How it works

  1. 1

    Paste the list, one entry per line

    Email addresses out of a form, SKUs out of a spreadsheet column, tags, branch names, IP addresses from a log. The page splits on any line ending — CR, LF or CRLF — so a file copied out of Windows and one copied out of a terminal behave identically. The counts under the output update on every keystroke.

  2. 2

    Say what “the same line” means

    Two switches decide it. “Ignore case” is on by default, because most pasted lists were typed by people and “Kuala Lumpur” and “kuala lumpur” are one place. “Trim before comparing” is also on, so a line with a stray trailing space is not counted as a different entry — that single invisible space is the most common reason a dedupe appears not to work.

  3. 3

    Choose which copy survives

    Keep first or keep last. It only changes the output when the duplicates are not identical, which is exactly what the two switches above cause: with case ignored, keep-first leaves “Ahmad Faizal” and keep-last leaves “AHMAD FAIZAL”. For anything ordered oldest-first, keep-last is the current state of the row and keep-first is the original one.

  4. 4

    Turn it inside out to audit instead of clean

    “Only the duplicates” prints one copy of each line that appears more than once and drops everything unique — the answer to “which of these did I enter twice?” rather than “make this list clean”. Switch on the count prefix and each row comes out as a number, a tab, then the line, which lands in a spreadsheet as two sortable columns.

  5. 5

    Sort with a collator, not a code-unit comparison

    A→Z, Z→A, shortest first, longest first, plain reverse, and shuffle. The alphabetical sorts use Intl.Collator with numeric ordering, so item2 comes before item10 and accented words file where the alphabet puts them. Length is measured in visible characters through the same grapheme splitter the case converter uses, so an emoji counts as one and not as several.

Frequently asked questions

Why does my editor put item10 before item2?

Because a plain sort compares UTF-16 code units one at a time. Comparing “item10” with “item2”, the first four characters tie, then it reaches “1” against “2” — code units 49 and 50 — and stops there, having never read the “0”. Run ["item10","item2"].sort() in a console and item10 comes first. This page compares with Intl.Collator built with numeric ordering, which reads a run of digits as one number, so item1, item2, item10, item100 come out in that order.

Why do accented words end up after z?

Same cause, different range. “z” is U+007A and “ä” is U+00E4, so a code-unit sort files every accented word after the entire English alphabet — “zebra” lands before “ähnlich”. A collator compares letters the way the alphabet does, so ä sorts beside a. The collator here uses base sensitivity, which also makes “Apple” and “apple” tie; JavaScript’s sort is stable, so tied lines stay in the order you pasted them rather than swapping around unpredictably.

Should I ignore case or not?

It depends on who wrote the list. For names, cities, tags and free text, ignore it: “Kuala Lumpur” and “kuala lumpur” are one entry and treating them as two defeats the purpose. For anything a machine produced — SKUs, git branches, object-storage keys, base64 values, hashes — turn it off, because “AB-100” and “ab-100” can be two genuinely different things and merging them silently loses one. The default is to ignore case, because most lists that get pasted into a web page were typed by a person.

Keep first or keep last — when does it actually matter?

Only when the duplicate lines are not byte-identical, which is exactly what happens once case is ignored or trimming is on. Then the setting decides which spelling survives: “Ahmad Faizal” or “AHMAD FAIZAL”, “ali@example.com” or “ali@example.com ” with the trailing space. It also matters for ordered data. In an export sorted oldest first, the last occurrence is the most recent state of that record, so keep-last gives you the current row and keep-first gives you the original. The number of lines removed is the same either way.

How do I find which lines are duplicated without removing anything?

Switch the mode to “Only the duplicates”. It prints one copy of every line that appears more than once, in the order they first appear, and leaves the unique lines out entirely — the same job as sort | uniq -d at a shell prompt. Add the count prefix and each row becomes the count, a tab, then the line, which is sort | uniq -c and is usually what people actually wanted. Pasted into a spreadsheet that is two columns, so you can sort by frequency and see the worst offender first.

Can I use this on a CSV?

On a simple one, yes. This page works strictly line by line, so if every record is one line it does exactly what you expect. It is the wrong tool for a CSV with quoted fields containing line breaks: a record like "Lot 5,↵Jalan Ampang" is one row to a spreadsheet and two lines here, and deduping would cut it in half without saying so. Open the file and look for quotes that span lines before you paste it. The same caution applies to any format where a record can wrap — SQL dumps with multi-line strings, for instance.

What happens to blank lines?

By default they are treated as ordinary lines, which means every blank line in the document is “the same line” and only one of them survives. That is right for a list and wrong for anything with paragraphs, so “Keep blank lines” passes each blank line through untouched and leaves it out of the comparison entirely. Blank lines are counted separately in the stats bar, so you can always see which of the two behaviours you are getting rather than discovering it after you paste the result somewhere.

Is Shuffle really random?

It is a Fisher–Yates shuffle driven by a small seeded generator — not Math.random(), and not the browser’s cryptographic source. That is deliberate: this page is prerendered to static HTML at build time, so anything on the render path has to give the same answer twice, or the page a crawler receives and the page you see would not match. Pressing Shuffle picks a new seed and re-runs it, so you can keep going until you like the order. It is fine for putting a list of names in an arbitrary order; it is not something to draw a prize with.

Does the list leave my browser?

No. Splitting, comparing, counting and sorting all happen in JavaScript already running in this tab, nothing is uploaded, and no part of the list is written to localStorage — reload and the box is empty. It is worth saying plainly, because the lists that get pasted into a deduplication tool are usually email addresses, customer names or exported identifiers, which is exactly the category of data that should not be sent anywhere merely to be counted.