Built on advanced on-device AI and hardware acceleration to deliver professional-grade performance with uncompromising privacy.On-device AI — professional-grade, fully private.
Remove Duplicates from CSV
Remove duplicate CSV rows — see exactly what goes before it does.
- Runs on your device
- Match on any columns · Keep first or last
- Review every row before it goes
- Free to 10 MB · Batch with Pro
Drop a CSV here
CSV, TSV, TXT or PSV · up to 10.0 MB
Whether every cell has to agree, or only the columns you pick.
Which row of a matching group survives. Its position never changes.
These apply to whole-row and column matching alike.
A header is passed through and never counted as a duplicate.
The rows that survive, or the ones that are duplicated.
Detected from the file. Set them when a file is unusual.
How the exported CSV is formatted.
Free: one file up to 10.0 MB, and the first 200 rows in the review table.
How to remove duplicates from a CSV
- 1
Add your CSV
Drop a .csv, .tsv, .txt or .psv file, or browse for one. Nothing is uploaded.
- 2
Choose what counts as a duplicate
Match on the whole row, or tick the columns that identify a record.
- 3
Set the matching rules
Decide whether case, surrounding spaces and punctuation should matter, and whether to keep the first or last occurrence.
- 4
Review what will be removed
Every row is marked Kept or Removed and grouped with the rows it matches, so you can check the result before you commit.
- 5
Export
Download the unique rows, or switch the output to the duplicates instead.
What this remove duplicates tool offers
Deduplication is easy to do and easy to do wrong. The difference is whether you can see what is being deleted before it goes. This tool is built around that review, and around a matcher flexible enough to describe what a duplicate actually means in your data.
- Whole-row or multi-column matching — Every cell must agree, or only the columns you tick — an email, or a name and a postcode together.
- Keep the first or the last — The first occurrence survives by default; switch to the last when the newest row holds the current state.
- Four matching rules — Ignore case, ignore surrounding spaces, collapse repeated spaces, ignore punctuation. Each one applies to every matching mode.
- Row-by-row review — Every row is marked Kept or Removed and tinted by the group it belongs to, so a run of near-identical records reads as one block.
- Unique rows or duplicates — Export what survives, or invert it and export only the records that are duplicated.
- Header handling that is visible — The header row is detected, the tool tells you which way it went, and you can override it.
- Output you control — Delimiter, line endings and an optional byte order mark, so the file opens correctly wherever it is going next.
- Nothing is uploaded — The file is read on your device. It never reaches a server, so a customer list stays a customer list.
What counts as a duplicate
This is the only decision that really matters, and most tools make it for you. There are two modes, and the difference between them is usually the difference between a useful result and a destructive one.
Matching on the whole row
Two rows are duplicates only when every cell agrees. This is the safe default: it can only ever remove a row that is genuinely a repeat of another, so it is the right starting point when you do not yet know how the file was produced. It is also the mode that misses the most, because one differing timestamp or one extra trailing space in one cell is enough to make two records of the same customer look distinct.
Matching on columns
Tick the columns that identify a record and the rest are ignored. One column is the common case — an email address, an order id, a SKU — but several together are often more honest: first name, last name and postcode identifies a person more reliably than any one of the three. The columns you have chosen are highlighted in the table, because the single most common cause of a surprising result is matching on fewer columns than you meant to.
- Empty cells still count — A blank value is a value. Two rows that are both blank in the key column are duplicates of each other, which is usually right and is always visible in the review.
- Ragged rows are squared off — Rows with fewer cells than the widest row are padded, so whether a duplicate is found does not depend on whether the exporter wrote a trailing separator.
- Order never changes — The surviving row stays where it was in the file. Nothing is sorted, moved or regrouped.
Which row is kept: first or last
When several rows share a key, exactly one of them survives, and which one is a real decision rather than a detail. Excel's own Remove Duplicates makes no promise about it at all, and Power Query's documentation says explicitly that there is no guarantee which instance is chosen.
- Keep first — The earliest occurrence wins. Correct when the file is in a meaningful order and the first entry is the original record.
- Keep last — The latest occurrence wins. Correct when the file is an append-only log — an export where every change to a record was written as a new line, and the newest line is the current state.
- Position is preserved either way — Keeping the last does not move the row to where the first one was. It stays at its own position, so a diff against the original file shows deletions and nothing else.
Matching rules: case, spaces and punctuation
Exact matching finds only the duplicates that were already identical. Real data is messier than that, and four rules cover most of the distance without asking you to calibrate a similarity threshold you cannot see.
Case
Off by default, so Apple and apple are treated as one value. Turn case sensitivity on when the distinction is meaningful — product codes, passwords-adjacent identifiers, anything where letter case is part of the value rather than an accident of typing.
Spaces
Trimming removes the spaces around a value, which is on by default because a leading space is almost never intentional and is invisible in every spreadsheet. Collapsing spaces goes further and treats any run of spaces inside a value as one, so Jon Smith matches Jon Smith. That one is off by default because it changes what counts as a match inside the value, not just at its edges.
Punctuation
Ignoring punctuation drops the common ASCII marks and the curly quotation marks a word processor substitutes. Combined with collapsing spaces it makes O'Brien, J. match OBrien J — the cheap and predictable part of what a fuzzy matcher does, without the part where a threshold quietly merges two different people.
- Every rule applies to every mode — The rules affect whole-row matching and column matching identically. There is no combination where a rule is on screen and doing nothing.
- The file is never modified — Normalisation happens only while comparing. The rows that are exported are the rows you supplied, character for character.
- Turn one on and watch the count — The duplicate counter above the table moves as soon as a rule changes, which is the fastest way to see whether it is finding real matches or over-reaching.
Header rows, and how they are detected
A header row must be excluded from matching, or it can be deleted as a duplicate of a data row that happens to look like it. Getting this wrong in the other direction is worse and much quieter: treating a data row as a header exempts it from deduplication permanently, so it survives however many times it repeats, and nothing about the output looks wrong.
- Detected, not assumed — A first row qualifies as a header when no cell is empty, no cell reads as a number or a date, no label repeats, and the rows below it look like data. A file with a single row has no header — there is nothing for one to head.
- The decision is shown — The tool states which way automatic detection went, next to the control, so you never have to infer it from the row count.
- Always overridable — Set it to yes or no when your file is unusual. A header is a claim about your data, and a heuristic does not get the last word on it.
- Headers are never counted as duplicates — When a header is in play it is passed through untouched and excluded from every count.
Reviewing what will be removed
Deleting rows is not reversible once the original is gone, and a summary that says 44,063 duplicates removed tells you nothing about whether those were the right 44,063. The review table is the answer to that, and it is free.
- Every row carries a verdict — Kept or Removed, as a badge, next to the row it applies to. Removed rows are struck through so the shape of the deletion is visible at a glance.
- Groups are tinted — Rows that match each other share a background colour and show how many rows are in the group, so five near-identical records read as one block rather than as five unrelated rows.
- Duplicates-only view — A file with 200,000 rows and nine duplicates is unreadable in full and obvious when filtered. One toggle switches between the whole file and only the rows that are in a duplicate group.
- Key columns are marked — The columns the match depends on are highlighted in the table head, which is usually enough on its own to explain a surprising result.
- Four live counters — Total rows, unique rows, duplicates found and duplicate groups, all updating as you change a rule.
A 10 MB file with 146,000 rows is read and matched in about a tenth of a second, off the main thread, so the page never freezes.
Output options
The deduplicated rows are only useful if the file opens correctly in whatever reads it next, so the way it is written is yours to set rather than inherited silently from the input.
- Unique rows or duplicates — Export what survives, or invert the output and export every row that belongs to a duplicate group — including the one that would have been kept, because a group shown without its first member is not a group.
- Delimiter — Keep the file's own delimiter, or write comma, semicolon, tab or pipe regardless of what came in.
- Line endings — LF for anything Unix-shaped, CRLF for Windows tools that will not accept anything else.
- Byte order mark — Excel on Windows opens a UTF-8 CSV without a BOM as Windows-1252 and mangles every accented character. Adding the mark fixes it, and it is off by default because it confuses several other readers.
- Quoting is repaired on the way out — Values containing the delimiter, a quote, a line break or a meaningful leading space are quoted correctly, so the file survives a round trip even if the original did not.
- The dropped rows, with Pro — A second file containing exactly the rows the export left out — the two together always add back up to what you supplied.
Why remove duplicates?
Duplicate rows are rarely a cosmetic problem. They inflate the numbers people make decisions on, and they cause visible failures downstream.
- Counts stop lying — Any total, average or group-by computed over a file with repeats is wrong by however many repeats there are, and it is wrong in a direction nobody checks.
- Imports stop failing — A unique constraint on an email or an id rejects the whole batch on the first repeat. Deduplicating first turns a failed import into a successful one.
- Nobody gets mailed twice — A contact list with repeats sends the same message to the same person more than once, which costs unsubscribes and, in some jurisdictions, more than that.
- Files get smaller — A third of a customer export being repeats is not unusual. Removing them makes every downstream step faster and every file cheaper to move.
- Merged sources become usable — Two exports concatenated always overlap. Deduplication on a shared key is what turns the concatenation into a single clean dataset.
Frequently asked questions
How does removing duplicates from a CSV work?
Add a CSV, choose whether a duplicate means the whole row or just certain columns, and the tool groups every row that shares the same key. One row from each group is kept and the rest are dropped. Before anything is exported you see the result marked up row by row, so you can check what is about to be removed and change the rules if it is not what you expected.
What defines a duplicate: the full row or a single column?
You choose, and you are not limited to one column. Matching on the whole row means every cell has to be identical. Switching to columns lets you tick any combination — an email on its own, or first name plus last name plus postcode — and rows count as duplicates when those columns agree, whatever the other columns say. The columns you are matching on are highlighted in the table so it is always visible which ones the answer depends on.
Which row is kept when duplicates are found?
The first occurrence by default, and you can switch to keeping the last. Keeping the last is the right choice when the file is an append-only log and the newest row holds the current state of the record. Either way the surviving row stays where it was, so the order of your file never changes — only the removed rows disappear.
Does case sensitivity affect deduplication?
Yes, and it applies whichever way you are matching. With case-sensitive off, Apple and apple are treated as the same value; turn it on to keep them apart. Three other rules sit beside it: ignoring the spaces around a value, collapsing repeated spaces inside it, and ignoring punctuation. Together they catch the near-duplicates that exact matching walks straight past, such as O'Brien, J. and OBrien J.
Is my data private, and is the tool free?
Your file is read on your device and never sent anywhere, so it stays private even on a shared network. Every part of getting the right answer is free: multi-column matching, all four matching rules, keep-first or keep-last, header detection, the full review of what will be removed, and the output options. Pro adds scale — several files at once, delivered as one archive — and a second file containing the rows that were dropped, for when you have to reconcile the deletion.
What input options are there and can I use it on mobile?
Comma, semicolon, tab and pipe files are all read, and the delimiter is detected for you with a manual override when a file is ambiguous. Encoding is detected from the byte order mark or the content, and can be forced to UTF-8, UTF-16, Windows-1252 or ISO-8859-1. Quoted fields containing commas or line breaks are handled correctly. On the way out you choose the delimiter, the line endings and whether to add a byte order mark for Excel. It works on any modern phone, tablet or desktop.
from 79 ratings
Rate this tool
Tap a star — it takes a second