Toolvore

Text Encoding Converter

Decode supported text encodings and save UTF-8 or UTF-16 with strict byte checks and BOM options.

This tool runs entirely in your browser. Your data is never uploaded, never stored, and never leaves your device.

Convert explicitly selected UTF-8, UTF-16 or Windows-1252 text bytes to Unicode output, with strict validation and an optional BOM.

How to use it

  1. 1Choose one text file up to 5 MiB or use the Windows-1252 example.
  2. 2Select the known input encoding and the Unicode output encoding.
  3. 3Choose whether to add an output BOM, then convert.
  4. 4Inspect the text preview and download the complete converted bytes.

Example

Input
Windows-1252: 43 61 66 e9 20 80 0d 0a
Output
UTF-8: 43 61 66 c3 a9 20 e2 82 ac 0d 0a

Café € with its CRLF preserved; output BOM off.

What happens to your data

Decoding and encoding stay local without file-content upload or browser-storage persistence. No encoding is guessed. Page assets and anonymous usage events may use the network, and downloads remain on your device.

Last updated October 2026

Convert the character encoding of a local text file while retaining its text and line endings. Explicitly choose UTF-8, UTF-16 little-endian, UTF-16 big-endian or Windows-1252 as input, then export UTF-8 or either UTF-16 byte order. You can include or omit the output byte order mark, commonly called a BOM.

This is for files whose original encoding you know. It does not guess an encoding, translate languages, fix already corrupted text, normalize Unicode or convert a document format. Choosing the wrong encoding can produce readable-looking but incorrect text even when decoding succeeds.

How it works

Select one text file up to 5 MiB. UTF-8 decoding rejects malformed, truncated, overlong and surrogate byte sequences. UTF-16 requires an even byte count and valid surrogate pairs. Windows-1252 maps its defined single-byte characters; undefined bytes 81, 8D, 8F, 90 and 9D are rejected rather than silently mapped to control characters. Windows-1252 is not ISO-8859-1.

A matching leading Unicode BOM is consumed as an encoding marker. A BOM that conflicts with your selected input encoding is an error: choose the matching encoding and convert again. A UTF-32 BOM signature is rejected because UTF-32 is outside the supported formats. Without a BOM, the selected input encoding is still used explicitly; the absence of a marker does not validate that selection.

Output BOM is a separate setting. UTF-8 with BOM begins EF BB BF; UTF-16LE with BOM begins FF FE; UTF-16BE with BOM begins FE FF. The rest of the bytes encode the same Unicode text. Existing CRLF, LF and CR line endings are retained without trimming. No Unicode normalization is applied.

The text preview is limited to 12,000 characters and is only an inspection aid; the download contains the entire converted byte sequence. Output may be larger than input, particularly when converting Windows-1252 to Unicode. Empty text creates an empty file when output BOM is off, or a file containing only the chosen BOM when it is on.

Changing the input encoding, output encoding, BOM setting or selected file removes the old download. Clear releases the temporary result. The file is processed locally, without content upload or browser-storage persistence. Page assets and anonymous usage events may use the network; downloaded copies remain on your device.

Worked example: Windows-1252 café and euro sign

Input
43 61 66 e9 20 80 0d 0a; Windows-1252
Result
43 61 66 c3 a9 20 e2 82 ac 0d 0a; UTF-8, no BOM

The text is Café € followed by CRLF. The original line ending is preserved byte-for-byte in the Unicode output.

Common use cases

  • Migrating an old Windows-1252 text export to UTF-8
  • Preparing UTF-16 text for an application that requires a specific byte order
  • Adding a required Unicode BOM to a known text file
  • Converting log or CSV text encoding without changing line endings

Frequently asked questions

Can the converter detect my file's encoding?+

No. You select the input encoding. The tool checks whether a leading Unicode BOM conflicts with that choice, but it does not infer an encoding from the remaining bytes.

What does UTF-16 little-endian or big-endian mean?+

UTF-16 stores each code unit as two bytes. Little-endian puts the lower byte first; big-endian puts the higher byte first. Choose the byte order the source file actually uses.

Should I include a UTF-8 BOM?+

That depends on the receiving application. UTF-8 has no byte-order ambiguity, but some software uses its BOM as a marker. Follow the application's requirements; this tool lets you choose either output.

Will converting Windows-1252 to UTF-8 fix mojibake?+

Only if you still have the original Windows-1252 bytes and select that encoding. Text that was already decoded incorrectly and saved again may require a separate repair process. A successful decode alone cannot establish the intended text.

Does this convert CSV delimiters or line endings?+

No. It changes character encoding only. Commas, tabs, CRLF, LF and CR remain as text characters. Use a CSV-specific tool for delimiter or table changes.

Why does a Windows-1252 file fail on byte 81?+

Byte 81 and four other positions are undefined in the Windows-1252 character table used here. The converter rejects them explicitly rather than inserting replacement or control characters. Confirm the source encoding or re-export from the original application.