Skip to content
ALL DEVELOPER TOOLS56
JSON FormatterTXT → JSONJSON to CSVJSON → CSVCSV to JSONCSV → JSONSQL FormatterSQL → SQLMarkdown to HTMLMD → HTMLHTML to MarkdownHTML → MDXML FormatterXML → XMLBase64 Encode/DecodeTXT → B64URL Encode/DecodeTXT → URLJWT DecoderJWT → JSONHTML Entity Encode/DecodeTXT → HTMLUUID Generator— → UUIDPassword Generator— → TXTHash GeneratorTXT → HASHLorem Ipsum Generator— → TXTQR Code GeneratorTXT → PNGColor Picker & Converter— → HEXCSS Gradient Generator— → CSSBox Shadow Generator— → CSSRegex TesterTXT → MATCHCron Expression GeneratorTXT → CRONTimestamp ConverterNUM → DATEText DiffTXT → DIFFText Case ConverterTXT → TXTWord CounterTXT → STATSSubtitle ConverterSUBS → SRT · VTT · ASSYAML to JSONYAML → JSONJSON to YAMLJSON → YAMLYAML FormatterYAML → YAMLJSON MinifyJSON → JSONSort & Dedupe LinesTXT → TXTLine Ending ConverterTXT → CRLF · LF · CRXLSX to CSVXLSX → CSVCharset ConverterTEXT → UTF-8HAR ViewerHAR → TABLE · HARJWT EncoderJSON → JWTHMAC GeneratorTXT → MACFile ChecksumFILE → VERDICTWi-Fi QR Code Generator— → QRChmod CalculatorOCTAL → RWXHTTP Status CodesCODE → MEANINGByte ConverterSIZE → UNITSTime Zone ConverterTIME → ZONESAspect Ratio CalculatorSIZE → RATIOURL ParserURL → PARTSExtract ArchiveARCHIVE → FILES · ZIPCreate ZIPFILES → ZIPUnzip FilesZIP → FILESDOCX to MarkdownDOCX → MDSVG OptimizerSVG → SVGBarcode Generator— → BARCODEQR Code ReaderIMG → TEXTVCF to CSVVCF → CSVICS to CSVICS → CSVCSV to Markdown TableCSV → MDBcrypt Generator— → HASH
ENGINE INTL.COLLATORACCEPTS TXT CSV LOG

Sort & Dedupe Lines

INTL.COLLATORTOOL 150 OF 190

Sort and Deduplicate Lines Online

Sort a list, drop the duplicates and the blank rows, and see exactly how many lines left and why.

ENGINEINTL.COLLATOR
ACCEPTSTXT CSV LOG
MAX SIZENONE
UPLOADNEVER
01Drop a .txt, .csv or .log file into the drop zone, or switch SOURCE to Paste.
02Choose an order under SORT, then pick Keep all, Keep first or Keep last under DUPLICATES.
03Check the lines in and lines out counts under OUTPUT, then copy the result or click Download.

About Sort & Dedupe Lines

A list of hostnames, an export of email addresses, a log of user agents, a column pasted out of a spreadsheet: the job is always the same few operations, and the part that goes wrong is always the counting. This page runs them in one fixed order — trim, drop blanks, de-duplicate, sort, reverse — and prints what each step cost. Lines in, lines out, duplicates removed and across how many repeated values, blank lines removed, whitespace trimmed from how many lines. A run that quietly took a hundred rows out of your list cannot pass for a clean one. Ordering uses the browser's own Unicode collation rather than raw code-point comparison, which is why apple sorts before Banana here and after Zebra in a naive sort, and why the natural mode puts file2 ahead of file10. De-duplication runs before the sort, so keeping the first occurrence means the first one in your file, not the first one after a re-order. Case sensitivity governs both the comparison and the duplicate test together. A leading UTF-8 byte-order mark is detached before any of that and written back at the front of the result, or dropped if you tick the box — it is never left attached to line one, where a sort would carry it into the middle of the file and where it would stop a marked line from matching its unmarked twin. There is no shuffle: it needs a random source, and a result you cannot reproduce is not one you can check.

Questions

It sorted file1, file10, file2 — that is not the order I want.

That is alphabetical order doing exactly what it says: comparing character by character, 1 comes before 2, so file10 lands between file1 and file2. Switch the sort to Natural. That mode turns on numeric collation, which reads a run of digits as a number rather than as characters, and gives you file1, file2, file10. It handles digits anywhere in the line, so IMG_2 before IMG_10 and v1.9 before v1.10 both work. Everything else about the sort is unchanged; only the treatment of digit runs differs.

Why does apple sort before Banana instead of after Zebra?

Because the ordering is the browser's Unicode collation, not a comparison of raw code points. In code-point order every capital letter comes before every lower-case one, so a naive sort produces Banana, Zebra, apple — which is almost never what anyone means by alphabetical. Collation orders letters the way a dictionary does and treats accented characters as variants of their base letter. If you want case to matter, leave Case-sensitive ticked and letters that differ only in case will still sort in a stable, defined order; untick it and API and api become the same line for both sorting and de-duplication.

Which copy of a duplicate line survives?

Whichever you choose, and the choice means what it says because the operations run in a fixed order: trim, drop blanks, de-duplicate, sort, reverse. De-duplication happens before the sort, so Keep first means the first occurrence in your original file rather than the first one after a re-order, and Keep last leaves that copy in the position it already occupied instead of moving it to the front. The counts on screen show how many lines were removed and across how many distinct repeated values, so a list with one line repeated fifty times reads differently from fifty lines each repeated once.

Can I sort a CSV by one of its columns?

No. Everything here compares whole lines, so a CSV sorts by its first column and then by whatever follows, and the header row sorts along with the data. There is a worse trap: a CSV field may legally contain a line break inside quotes, and this tool splits on every break, so such a record would be torn into two lines and could then be sorted apart. If your data has quoted multi-line fields, do not run it through here. Use a spreadsheet, or convert to JSON first with the CSV to JSON tool.

My output has fewer lines than I expected.

Read the four counters — they exist for exactly this. Lines in and lines out bracket the run, and the two causes are itemised separately: duplicates removed, and blank lines removed. Blank means empty or whitespace-only, so a line of spaces counts. Trim runs first and can turn two lines that looked different into the same line, which then de-duplicates. And with Case-sensitive off, API and api collapse to one. A warning above the panes spells out each of these in words whenever it applies, so nothing leaves silently.

Is my file uploaded to a server?

No. Transmute processes everything locally in your browser using JavaScript and WebAssembly. Your files never leave your device — there is no server, no upload, no cloud processing.

Related