ALL DEVELOPER TOOLS56
Charset Converter
Charset Converter — Legacy Encodings to UTF-8
Read a file in Windows-1252, Shift_JIS, KOI8-R or any of 25 encodings, see the accents come out right, and save it as UTF-8.
About Charset Converter
A plain text file records nothing about its own encoding. It is a run of bytes, and the same bytes mean different letters depending on which table the reader uses — which is why a Czech subtitle file, a Russian CSV or an old Japanese log opens as a screenful of question marks and diamonds. This tool reads the bytes in whichever encoding you name and writes them back out as UTF-8, the encoding everything modern agrees on. Detection believes a byte-order mark, spots UTF-16 by its pattern of zero bytes, and otherwise tests the bytes strictly against UTF-8; bytes that fail that test are read as Windows-1252 and flagged as a guess, because a Cyrillic, Greek or Turkish file is equally valid non-UTF-8 and nothing can tell them apart from the bytes alone. That is why the preview is the point: you read the accented characters, and if they are wrong you name the real encoding and watch them change. The conversion runs one way only, and that is a browser limit rather than a missing feature. The browser can decode about forty encodings, but its encoder writes UTF-8 and nothing else, so there is no in-browser path to Shift_JIS or Windows-1251 bytes. The page reports bytes in, bytes out, whether a mark was found, and how many characters the decoder could not read at all.
Questions
My CSV opens in Excel with é where the accents should be.
That is a file written in one encoding and read in another. é is what the two UTF-8 bytes for é look like when something reads them one byte at a time as Windows-1252, and the reverse mistake gives you a diamond or a question mark instead. Drop the file here and the preview shows you what each candidate encoding actually produces, so you can settle it by looking rather than guessing. Save the result as UTF-8, and if the file is a CSV you are about to open in Excel on Windows, tick the byte-order mark box as well — Excel needs that mark before it will treat a CSV as UTF-8.
Can I convert a file TO Shift_JIS or Windows-1251?
No, and no browser tool can. The browser gives pages two halves of the encoding standard and they are not symmetrical: TextDecoder reads about forty encodings, and TextEncoder writes UTF-8 only — its constructor takes no argument at all. So a page can read Shift_JIS bytes and cannot produce them. This tool therefore converts legacy encodings to UTF-8 and states that plainly rather than offering a target list it cannot deliver. If you genuinely need legacy bytes out, that is a job for iconv or Python on your own machine, where the full encoding tables are available in both directions.
How does it know which encoding my file is in?
Mostly it does not, and it says so. Nothing inside a plain text file records its own encoding. Detection runs in order: a byte-order mark is a statement by whoever wrote the file, so it is believed; BOM-less UTF-16 is recognised from its pattern of zero bytes, which has to be checked before UTF-8 because a strict UTF-8 decode accepts those bytes and returns text with a gap between every letter; then the bytes are tested strictly against UTF-8. Anything that fails that test is read as Windows-1252 and labelled a guess, because a Cyrillic, Greek or Turkish file is equally valid non-UTF-8. Reading the preview is the only way to resolve it, which is what the picker is for.
What does the UNDECODABLE count actually mean?
It counts the U+FFFD replacement characters in the decoded text — one for each byte sequence the decoder could not use. A number above zero on a multi-byte encoding such as UTF-8, Shift_JIS or Big5 means the label is wrong for these bytes, and picking the right encoding takes it to zero. Zero does not prove the opposite, though. A single-byte codepage like Windows-1252 maps all 256 byte values, so no byte in it can ever be invalid: a Russian file read as Windows-1252 scores zero undecodable characters and is still completely wrong on screen. That case is what the preview and the guess warning are for.
Should I tick the byte-order mark box?
Only for software that asks for it. The download is UTF-8 either way; the mark is three extra bytes at the front that announce the fact. Excel on Windows is the main reason to add one — without it, Excel opens a CSV in the system codepage and mangles every accent, no matter how correct the file is. Most other software ignores the mark, and a few things break on it: it is invalid at the start of a JSON document, it shows up as stray characters in some shell pipelines, and web pages and browsers never need it. When in doubt, leave it off.
Is my file uploaded to a server?
No. Transmute processes everything locally in your browser using JavaScript and WebAssembly. Your files never leave your device — there is no server, no upload, no cloud processing.