Base64 Explained: What It Actually Does and Where It Breaks

A practical, from-the-bytes-up guide to Base64: why it exists, how the 3-to-4 character mapping works, why UTF-8 trips up naive encoders, what URL-safe means, and the security mistake almost everyone makes at least once.

Base64 is one of the most misunderstood little formats on the internet. Developers paste a string into it to "encrypt" a password, embed images directly in CSS, and ship tokens in HTTP headers — often without a clear mental model of what the encoding is doing or where it quietly falls apart. This article walks through the actual mechanics, then covers the four places where Base64 causes real bugs.

Why Base64 exists at all

A lot of systems were designed to carry text, not arbitrary bytes. Early email (SMTP) could only reliably transport printable 7-bit ASCII. A URL query string has characters that mean something else — a plus, a slash, an equals sign. JSON string literals cannot contain raw control bytes. If you try to shove a PNG file, an encryption key or a UTF-8 emoji through any of those pipes, the raw bytes either get mangled or break the framing of the format itself.

Base64 solves this by rewriting any sequence of bytes using only 64 characters that every one of those systems agrees are safe: A–Z, a–z, 0–9, plus + and /. Because it uses exactly 64 symbols, each character carries 6 bits of information. That is the whole idea — a transport-safety shim, nothing more.

The 3-to-4 byte mapping, concretely

Base64 reads input in groups of three bytes (24 bits) and emits four characters (4 × 6 = 24 bits). No bits are added or lost; the data simply gets re-chunked from 8-bit groups into 6-bit groups. When the input length is not a multiple of three, the final group is padded with = characters so the output length is always a multiple of four.

Take the three letters "Man". In ASCII those are 0x4D, 0x61, 0x6E. Write them as 24 bits and split into 6-bit groups:

M  a  n
77 97 110  (decimal)
01001101 01100001 01101110  (binary, 24 bits)
010011 010110 000101 101110  (re-split into 6-bit groups)
19       22      5      46     (each group as a number)
T        W        F       u     (look up in the Base64 alphabet)

So "Man" encodes to "TWFu". Add one byte at a time and watch the padding appear: "Ma" becomes "TWE=" and "M" becomes "TQ==". This is why Base64 always expands your data by roughly 33% — three bytes in, four characters out. If you are inlining a large asset as a data URI, budget for that overhead.

Where naive encoders break: UTF-8

The single most common bug is treating Base64 as a character-level operation when it is a byte-level one. In JavaScript, btoa() operates on "binary strings" — strings where each character must fall in the 0–255 range. Feed it a character above U+00FF, like an accented letter or an emoji, and it throws.

The fix is to encode to UTF-8 bytes first, then Base64 those bytes. On the decode side, Base64 back to bytes, then decode those bytes as UTF-8:

// Encode (handles emoji, CJK, accents)
const bytes = new TextEncoder().encode('héllo 🌍')
let bin = ''
bytes.forEach((b) => (bin += String.fromCharCode(b)))
const encoded = btoa(bin) // 'aMOpbGxvIPCfjI4='

// Decode
const back = new TextDecoder().decode(
  Uint8Array.from(atob(encoded), (c) => c.charCodeAt(0))
)

Our Base64 tool does exactly this round-trip, which is why it handles multilingual text and emoji without mangling them. If a tool you have used elsewhere corrupts non-Latin input, this byte-vs-character confusion is almost certainly the cause.

URL-safe Base64 and why the + and / are a problem

The standard alphabet includes + and /, both of which are reserved in URLs. In a query string, + is interpreted as a space by many parsers, and / changes path structure. Putting standard Base64 into a URL or a filename is therefore asking for trouble.

The URL-safe variant (RFC 4648 §5) swaps + for - and / for _, and usually drops the trailing = padding. This is what you will see in JWTs, in cookie values, and in most modern APIs. Decoding is just the reverse substitution, plus re-adding padding to a multiple of four. Toggle the URL-safe option in the tool to encode or decode this variant.

The security mistake everyone makes once

Base64 is encoding, not encryption. It has no key, no secrecy, and is reversible by anyone in a single step. Storing a password "as Base64" in a database, or Base64-ing a token "so it is not readable", protects nothing — it only stops you from reading it at a glance. Anyone who finds the string can decode it in milliseconds.

That said, Base64 shows up constantly in security-adjacent places, and knowing what you are looking at matters. A JWT is three Base64url segments separated by dots; the signature is not encrypted, only encoded, so anyone can read the payload claims. A Basic-auth HTTP header is literally "Basic " plus Base64 of "user:password" — which is why Basic auth must always run over TLS. Recognising Base64 on sight (long string, ends in = or ==, only uses the 64-character alphabet) lets you understand what a system is really doing with your data.

Quick reference

  • Alphabet: A–Z a–z 0–9 + / (standard) or - _ (URL-safe), with = padding.
  • Size: output is always a multiple of 4 characters, about 4/3 the input size.
  • It encodes bytes, not characters — always UTF-8-encode text first.
  • It is reversible and public: never treat it as a security measure.
  • Padding can be stripped for transport as long as both sides agree to re-add it.

If you just need to move text or bytes through a system that only speaks ASCII — a data URI, a JSON field, an HTTP header — Base64 is the right tool, as long as you respect the four rules above. You can encode, decode and try the URL-safe variant right here, entirely in your browser.

Base64 Encode / DecodeText to Base64 — open the tool