On-device AI — professional-grade, fully private.

Compare CSV Files

CSV diff by key column — see which cells changed, not just rows.

  • Runs on your device
  • Matches by key column
  • Cell-level changes
  • Two files free · 25 MB each

Drop two CSV files here

CSV, TSV or text · up to 25 MB each on Free

File A

No file yet

File B

No file yet

Load two files and press Compare.

Comparison settings
Match rows by
Match rows by

Key column pairs records regardless of order and reports edits. Whole row counts duplicates. Position absorbs inserted rows.

Key columns

Load two files with headers to choose a key.

Choose one or more columns that identify a record. Suggested columns are the ones whose values are actually unique.

Match columns by
Match columns by

By position pairs the first column with the first. By name pairs Customer ID with customer_id, and survives a column being added in the middle.

Compare these columns

Compare two files to choose columns.

Untick a column to leave it out of the comparison entirely.

Treat as the same

All off by default. Each one removes a class of false difference — and hides a real one if you did not mean it.

Reading the files
File A
File B

Detected per file. Override either side if a file was read wrongly.

Export format

Applies to every download and to the clipboard.

Download
Only in A0
Only in B0
Changed0
In both0
Change report0
Pro
Pro

Each set on its own, or the whole comparison at once.

How to compare two CSV files

  1. 1

    Load both files

    Drop or browse for two CSV, TSV or text files, or switch to Paste and paste them in. Each file is read on your device; nothing is uploaded. The delimiter, quote character, encoding and header row are detected for each file separately, and every one of them can be overridden.

  2. 2

    Choose how rows are matched

    Pick a key column such as id, SKU or email to match records regardless of row order and see which cells changed. Or match whole rows, which counts duplicates correctly. Or match by position, which absorbs an inserted row instead of shifting everything after it.

  3. 3

    Read the result

    Four tabs — changed, only in A, only in B, and unchanged — with counts. In changed records the moved cells are highlighted and show both the before and the after value, so you can see the edit rather than infer it.

  4. 4

    Export what you need

    Download any set as CSV or TSV, copy it to the clipboard, or take the long-format change report with one line per changed cell. Pro adds the whole comparison as one ZIP and a colour-coded workbook.

What this CSV Compare tool offers

Two CSV files in, and a precise account of what changed between them: which records are new, which are gone, which were edited and exactly which cells moved. Every option that decides whether the answer is correct is free.

  • Three ways to match rows — By key column, by whole-row content, or by position. Each answers a different question, and the right one depends on whether your files share a stable identifier.
  • Changed records, not just added and removed — One edited cell is reported as one change, with the before and after value side by side, instead of as a deletion plus an insertion you have to notice are the same record.
  • Columns that do not line up — Match columns by header name across different spellings, and ignore the columns that always differ.
  • Differences you did not mean to ask about — Ignore case, surrounding whitespace, empty-versus-NULL and numeric formatting — individually, so you can see what each one is hiding.
  • Five exports — Each set on its own, plus a long-format change report with one line per changed cell. Pro adds one ZIP and a colour-coded workbook.
  • Files read as they really are — Delimiter, quote character, encoding and header row detected per file, quoted line breaks handled, byte-order marks stripped.

How rows are matched, and which mode to use

This is the one decision that changes every number the tool reports, so it is worth thirty seconds. All three modes are free.

  • Key column — use this if your data has an id — Records are paired on one or more columns that identify them: an id, SKU, order number, email or a composite of several. Row order stops mattering entirely, and because the pairing survives an edit, a changed record is reported as changed rather than as a removal plus an addition. This is the mode that answers "what actually happened to this record".
  • Whole row — use this for set differences — Two rows match only when every compared cell matches. Order does not matter. Duplicates are counted rather than collapsed, so three copies against one report one match and two extras. There is no notion of a changed record in this mode, because without a key there is nothing to say two rows are versions of the same thing.
  • Position — use this when the files are in the same order — The two files are aligned line by line, but with insertions absorbed: adding a row near the top reports one insertion, not one insertion and then every following row shifted. Where a deletion and an insertion sit next to each other and the rows still agree on at least half their columns, they are re-read as one edited record.
  • A suggested key is a suggestion — When both files carry headers, the tool scores each shared column on how unique its values actually are and offers the best one. A column named id that is not unique will not be suggested, because naming a column id does not make it a key.

What "changed" means, cell by cell

A changed record is one that exists on both sides and disagrees somewhere. The tool tells you where.

  • The moved cells are marked — In the changed tab each cell that differs is highlighted and shows both values, so the edit is visible rather than something you compare by eye across two windows.
  • Key columns are never reported as changed — They are what paired the two records in the first place, so by construction they agree. They are shown for identification only.
  • The change report is the one to circulate — One line per changed cell — status, key, row numbers in both files, column, before, after — plus one line for each added and removed record. It is the form that pastes into a ticket without editing.
  • A changed header is reported separately — If the two files disagree about their column names, that is a schema change rather than a data change and it is called out on its own instead of being mixed into the record counts.

Columns that do not line up

Two exports of the same table rarely agree on column order or on how a header is spelled. Comparing by position after either of those has happened reports every row as changed — a true statement about the bytes and a useless one about the data.

  • Match columns by header name — Customer ID, customer_id, customerId and CUSTOMER-ID all fold to the same identity, so they pair up. Punctuation, spacing and case are ignored; accented and non-Latin headers fold correctly too.
  • Columns on one side only — A column that exists in only one file is listed rather than silently dropped, and it does not make every record look changed.
  • Ignore the columns that always differ — A last-modified timestamp, a row number, an export id — untick them and they stop contributing to the comparison entirely, in both the matching and the change detail.
  • Duplicate header names are handled in order — Two columns both called Notes pair with the two called Notes on the other side, first with first, rather than both collapsing onto one.

Ignoring differences you did not mean to ask about

Four switches, all off by default. Each one exists because it is a routine source of false differences between two systems' exports — and each one also hides a real difference if you turn it on without meaning to, which is why none of them is on for you.

  • Ignore case — Apple and apple match. Email addresses and product codes are the usual reason.
  • Ignore surrounding whitespace — " pear " and "pear" match. One system trims on export and the next does not.
  • Treat empty and NULL as the same — An empty cell matches the literal tokens NULL, NA, N/A and nil, in any case, plus a cell holding only spaces. The list is deliberately short and closed: a dash, a zero and the word none are all plausible real values and are not included.
  • Compare numbers by value — 1.0 matches 1, 1e3 matches 1000, and a leading plus sign stops mattering. This is the one to be careful with: it also makes 007 match 7, which is wrong if that column is an identifier — which is exactly why it is off unless you ask.

Delimiters, quoting and encodings

Each file is read on its own terms. A comma-separated export and a semicolon-separated one compare perfectly well against each other.

  • Delimiter detection that reads the file, not the first line — Comma, semicolon, tab and pipe are each tried against a sample and scored on whether they produce a consistent number of columns. Counting separators on line one — the usual approach — is fooled by a single comma inside a quoted address, and by a header row, which is the least representative line in the file.
  • Quoted fields with commas and line breaks — A quoted field may contain the delimiter, quotes doubled to escape them, and real line breaks. All three are common in exported address and notes columns, and all three break a parser that splits on newlines first.
  • Byte-order marks are stripped — Excel's "CSV UTF-8" export writes one. Left in place it makes the first header cell subtly different from its twin, so the entire header row reads as changed — which looks like two unrelated files. Handled here without being asked.
  • Encodings, including the ones Excel actually writes — UTF-8, UTF-16 in both byte orders — what Excel's "Unicode Text" export produces — plus Windows-1252 and true ISO-8859-1. A mark in the file overrides the dropdown, and the tool says so rather than silently doing the right thing under a control that now reads as a lie.
  • Ragged and blank rows are kept — A row with an extra column, or a blank line in the middle, is part of the file. Dropping blank lines silently shifts every row number after them, which makes the report point at the wrong lines.

Duplicate rows, and why the counts matter

A duplicated record is usually the thing you were looking for, so a comparison that smooths it over has failed at the job.

  • Whole-row matching counts, it does not collapse — Three identical rows in the first file against one in the second give one match and two rows present only in the first. A tool that treats each file as a set of distinct rows reports all three as present in both, which is wrong and looks right.
  • Repeated keys are paired in order and reported — If a key value occurs twice on each side, the first pairs with the first and the second with the second, and the repeated key is listed so you know the pairing was arbitrary.
  • Rows with an empty key are counted — They cannot be matched meaningfully, so the tool says how many there are rather than quietly pairing them with each other.

The exports

Five files, each of which answers a different question, plus two ways to take the whole thing at once.

  • Only in A and only in B — Complete rows in their own file's shape, with the original header, so they re-import into whatever produced them.
  • Changed — One row per changed record, every compared column present as a before and after pair, and a final cell naming which columns actually moved.
  • In both — The records that did not move, for when that is the interesting set.
  • Change report — Long format: one line per changed cell plus one per added or removed record. Status, key, both row numbers, column, before, after.
  • One ZIP, and a colour-coded workbook — Pro takes every non-empty set plus a manifest as a single archive, or a spreadsheet with added rows green, removed rows red and the changed cells amber. Both hold exactly the data the free downloads do, gathered into one file.
  • Output written the way you need it — Choose the delimiter, the line ending and whether to write a byte-order mark, and name the files with a prefix. Exports are written with your settings, not with the input's.

Both files are read once into a byte index and compared without ever building a string per cell, so a pair of 25 MB exports finishes in about a second.

Why compare CSV files?

Two exports of the same table, taken at different times or from different systems, and a question about what moved between them.

  • Reconciling two systems — A CRM export against a billing export, matched on email or account number, to find the records one has and the other does not.
  • Checking what an import did — Export before, export after, match on the primary key, and read the changed list. It is faster than trusting the import log and it names the cells.
  • Reviewing a data change before it ships — The change report is a reviewable artefact: a colleague can read it without opening either file.
  • Auditing a price or stock list — Match on SKU, ignore the timestamp column, and the changed tab is the list of what actually moved.
  • Working on data you cannot upload — It all runs on your device, not our servers, so the tool is usable on customer records, payroll and anything else that cannot go to a server.

Frequently asked questions

How does the CSV comparison work?

You choose how rows are matched. With a key column — an id, SKU, order number or email — records are paired on that value regardless of row order, and the tool reports which are new, which are gone, and which changed, down to the individual cell. With whole-row matching every cell must be identical for two rows to count as the same, and duplicates are counted rather than collapsed. With position matching the two files are aligned line by line, and an inserted row is reported as one insertion instead of shifting everything after it.

Does the comparison consider column order or headers?

Both, and you choose. Matching columns by position pairs the first column with the first, which is right when the two files came out of the same export. Matching by header name pairs them by what they are called, so Customer ID and customer_id line up and a column added in the middle of the newer file does not push everything sideways. You can also mark columns to ignore, so a last-modified timestamp that always differs stops drowning out the changes you care about.

What is the difference between case-sensitive and case-insensitive mode?

With case-sensitive comparison on, which is the default, Apple and apple are different values. Turn it off to ignore letter case so those match — useful when two exports differ only in capitalisation, as email addresses and product codes often do. Three related switches sit beside it: ignore surrounding whitespace, treat an empty cell and NULL as the same thing, and compare numbers by value so 1.0 and 1 match. All four are off by default, because each one also hides a real difference if you did not mean to ask for it.

Can I download or copy the results?

Yes. Each set — changed, only in A, only in B, and unchanged — has its own download and copy button, written with your chosen output delimiter, line ending and optional byte-order mark. There is also a change report in long format: one line per changed cell, naming the record, the column, the value before and the value after, which is the version that goes into a ticket. An optional prefix names the files. Pro adds the whole comparison as one ZIP and a colour-coded workbook where added rows are green, removed rows red and changed cells amber.

Which delimiters and encodings can I use?

Each file has its own delimiter, quote character and encoding, detected automatically and overridable. Comma, semicolon, tab and pipe are offered, and detection scores each one by whether it actually produces a rectangle rather than by counting separators, so a comma inside a quoted address does not fool it. Encodings are UTF-8, UTF-16 in both byte orders, Windows-1252 and ISO-8859-1; a byte-order mark in the file overrides the dropdown, because the file knows and the dropdown is a guess.

Is my data private, and is the tool free?

Both files are read and compared on your device. Nothing is uploaded, and the tool keeps working offline once the page has loaded, which is what makes it usable on data you are not allowed to send anywhere. It is free for two files of up to 25 MB each, with every match mode, every normalisation option and the full cell-level detail included — the free tier is capped on scale, never on whether the answer is correct. Pro adds larger files limited only by your device, the one-ZIP download and the colour-coded workbook.

Why does my comparison show every row as different?

Almost always one of four things, and the tool names each of them on screen. The delimiter was detected differently for the two files, so one of them was read as a single column. One file has a header row and the other was not recognised as having one, which shifts every record by one. The columns are in a different order, which position matching cannot see through — switch to matching columns by header name. Or one file carries a byte-order mark that makes its first header cell subtly different from the other's; that one is handled automatically here, and is the single most common cause elsewhere.

Can it handle duplicate rows?

Yes, and correctly, which is less common than it sounds. Whole-row matching treats the files as multisets: three copies of a row in the first file against one in the second reports one match and two rows present only in the first, rather than declaring all three present in both. With key matching, a key value that appears more than once is paired in order and the repeated keys are listed as a warning, because a duplicated key is usually the thing you were looking for rather than a detail to smooth over.

How large a file can it compare?

Two files of 25 MB each on the free tier, which is roughly 390,000 rows of eight columns per file. Both files are held at once because a comparison needs them both, and they are held as a byte index rather than as several million separate strings, which is what makes a pair that size finish on a phone. Pro removes the imposed cap; above it the limit is your device's memory, and the tool tells you when a file is larger than it usually handles rather than refusing it.

4.8out of 5

from 72 ratings

Rate this tool

Tap a star — it takes a second