Skip to main content
I
Uni
UNICODE
Tools/UTF Endianness Inspector

UTF Endianness & Raw Byte Serialization Inspector

Inspect, compare, and convert text across all 5 standard Unicode serialization formats: UTF-8, UTF-16LE, UTF-16BE, UTF-32LE, and UTF-32BE with optional Byte Order Mark (BOM) inclusion.

UTF-8 (Web Standard)

Width: 1–4 BytesEndianness: N/A (Byte Stream)
55 6E 69 63 6F 64 65 20 F0 9F 9A 80

Backward-compatible with ASCII. Uses 1 byte for ASCII, 2 for Latin/Greek, 3 for CJK/Arabic, 4 for emojis.

UTF-16LE (Little-Endian)

Width: 2 or 4 BytesEndianness: Least Significant First
55 00 6E 00 69 00 63 00 6F 00 64 00 65 00 20 00 3D D8 80 DE

Default in Windows NT kernel, JavaScript V8 string memory, and Java JVM.

UTF-16BE (Big-Endian)

Width: 2 or 4 BytesEndianness: Most Significant First
00 55 00 6E 00 69 00 63 00 6F 00 64 00 65 00 20 D8 3D DE 80

Standard Network Byte Order for UTF-16 transmission in legacy protocols.

UTF-32LE (Fixed 4-Byte)

Width: 4 Bytes (Fixed)Endianness: Least Significant First
55 00 00 00 6E 00 00 00 69 00 00 00 63 00 00 00 6F 00 00 00 64 00 00 00 65 00 00 00 20 00 00 00 80 F6 01 00

Direct 1-to-1 array indexing for every code point in memory without surrogate pairs.

UTF-32BE (Fixed 4-Byte)

Width: 4 Bytes (Fixed)Endianness: Most Significant First
00 00 00 55 00 00 00 6E 00 00 00 69 00 00 00 63 00 00 00 6F 00 00 00 64 00 00 00 65 00 00 00 20 00 01 F6 80

Fixed 32-bit dwords serialized with most-significant byte first.

Character-by-Character Byte Breakdown (9 Glyphs)

Anatomical inspection of each code point with its individual UTF-8 byte stream and UTF-16 LE/BE encodings.

GlyphCodepointUTF-8 HexUTF-16LEUTF-16BEUTF-32LE
UU+00555555 0000 5555 00 00 00
nU+006E6E6E 0000 6E6E 00 00 00
iU+00696969 0000 6969 00 00 00
cU+00636363 0000 6363 00 00 00
oU+006F6F6F 0000 6F6F 00 00 00
dU+00646464 0000 6464 00 00 00
eU+00656565 0000 6565 00 00 00
U+00202020 0000 2020 00 00 00
🚀U+1F680F0 9F 9A 803D D8 80 DED8 3D DE 8080 F6 01 00

The Engineering Science of UTF Serialization & CPU Endianness

In computer architecture, endianness refers to the order in which bytes of a multi-byte word are stored in computer memory. The terms originate from Jonathan Swift’s Gulliver’s Travels, where rival factions fought over whether soft-boiled eggs should be cracked from the larger end (Big-Endian) or smaller end (Little-Endian).

When serializing 16-bit (UTF-16) or 32-bit (UTF-32) Unicode integers, the byte order matters critically:

Little-Endian (LE): Least Significant Byte (LSB) first — native to Intel x86, AMD64, Apple Silicon (ARM), Windows NT, and JavaScript V8.
Big-Endian (BE): Most Significant Byte (MSB) first — native to traditional TCP/IP Network Byte Order and mainframe architectures.

UTF-8, by contrast, encodes characters as an 8-bit stream of bytes. Because every encoding unit is already a single byte, UTF-8 is completely immune to endianness issues and does not require a BOM header.

Byte Order Mark (BOM U+FEFF) Signatures Matrix

UTF-8 BOM:EF BB BF
UTF-16 Little-Endian:FF FE
UTF-16 Big-Endian:FE FF
UTF-32 Little-Endian:FF FE 00 00

Frequently Asked Questions (FAQs)

Why does Windows Notepad add a UTF-8 BOM by default?+

Legacy versions of Windows Notepad used the EF BB BF byte signature to distinguish UTF-8 text from Windows-1252 ANSI. Modern standards recommend omitting BOM in UTF-8 to prevent shell script execution failures.

How does UTF-32 provide O(1) constant-time character indexing?+

Because every single Unicode code point in UTF-32 consumes exactly 4 bytes (32 bits), accessing character at index N is a simple pointer calculation: buffer + (N * 4).

Can I copy raw hex bytes directly into C/C++ or Rust code?+

Yes! Click the "Copy Hex Bytes" button next to any encoding to copy formatted hex sequences for uint8_t arrays or buffer initialization.