Hashing vs. Encryption vs. Encoding: Three Different Jobs Wearing Similar Clothes
These three terms get used interchangeably often enough that it's worth being precise about why that's a problem: they solve different problems, and picking the wrong one doesn't just mean using suboptimal terminology — it means building something insecure while feeling confident it's secure.
These three terms get used interchangeably often enough that it's worth being precise about why that's a problem: they solve different problems, and picking the wrong one doesn't just mean using suboptimal terminology — it means building something insecure while feeling confident it's secure.
Here's a way to keep them straight that doesn't require memorizing definitions: ask what each one is for.
- Encoding is for compatibility. It answers: "how do I represent this data in a format some other system can read?"
- Encryption is for confidentiality. It answers: "how do I hide this data from anyone without the key, while still being able to get the original back?"
- Hashing is for integrity and verification. It answers: "how do I prove this data is what I think it is, or check a value against a secret, without ever needing the original back?"
The single most useful fact connecting all three: encoding and encryption are both reversible by design. Hashing is deliberately, permanently one-way. That one distinction explains almost every mistake people make when they're used in the wrong place.
Encoding: no secret involved at all
Base64, URL encoding, and Hex are the common examples, and none of them involve a key or a secret of any kind. Encoding exists because not every system can safely transmit or store every byte value — Base64 exists because email systems and JSON couldn't reliably carry arbitrary binary data, so binary gets translated into a restricted set of printable characters instead.
Anyone can decode Base64. There's no key required, because there was never anything to protect — encoding was never trying to hide the data, only to make it transportable. This is exactly why "we Base64-encoded the password before storing it" is not a security measure; it's a formatting choice that a five-second lookup reverses completely.
flowchart LR
A["Hello"] -- encode --> B["SGVsbG8="]
B -- decode --> A
Encryption: a secret you can get back out
Encryption transforms data using a key, such that the original can be recovered — but only by someone who holds the correct key. This is the tool for confidentiality: data at rest on a disk, data in transit over a network, a file you want only specific people to be able to open.
There are two broad families. Symmetric encryption (like AES) uses the same key to encrypt and decrypt, which is fast but requires both parties to already share that key somehow. Asymmetric encryption (like RSA or ECC) uses a mathematically linked key pair — encrypt with one, decrypt with the other — which solves the "how do we agree on a shared secret in the first place" problem, at the cost of being significantly slower. In practice, most real systems use both together: asymmetric encryption to safely exchange a symmetric key, then symmetric encryption for the actual bulk data, because it's fast enough to use at scale. This is exactly what happens at the start of every TLS connection.
flowchart LR
A["Plaintext"] -- "encrypt (key)" --> B["Ciphertext"]
B -- "decrypt (key)" --> A
The property that matters most: encryption is meant to be reversed, by the right party, using the right key. If nobody can ever decrypt it again, that's not encryption working correctly — that's a lost key and a data loss incident.
Hashing: no going back, on purpose
Hashing takes an input of any size and produces a fixed-size output — a digest — through a one-way function. There is no key to reverse it with, because reversing it isn't supposed to be possible at all. Given the hash, you cannot recover the original input; that's not a limitation of current hash functions, it's the entire design goal.
This is why passwords should be hashed, never encrypted. Encrypting a password implies there's a key somewhere that turns it back into the plaintext password — which means anyone who gets that key (an attacker, an insider, a misconfigured backup) can recover every user's actual password. Hashing a password means that even the system storing the hash has no way to recover the original, ever. When a user logs in, the system doesn't decrypt anything to check the password — it hashes what the user just typed and compares the two hashes.
flowchart LR
A["correct-horse-battery-staple"] -- "hash (one-way)" --> B["e3b0c44...af434c"]
B -. "no path back" .-> A
Modern password hashing goes further than a plain hash function like SHA-256, because SHA-256 is fast — deliberately fast, since it was designed for integrity checks, not password storage — and fast means an attacker with a stolen hash database can try billions of guesses per second. Purpose-built password hashing algorithms like bcrypt, scrypt, and Argon2 are deliberately slow and memory-intensive, specifically to make large-scale guessing expensive even after a breach. They also incorporate a per-password salt automatically, which defeats precomputed rainbow-table attacks by ensuring two users with the same password don't produce the same hash.
Hashing shows up beyond passwords too: verifying a downloaded file hasn't been corrupted or tampered with (compare its hash to a published one), detecting whether a document has changed (a single-bit change produces a completely different hash — the avalanche effect), and as a building block inside HMAC, digital signatures, and blockchain structures.
Where people actually mix these up
The classic mistake is "we encrypted the passwords" as a security claim — which, if true, is actually worse than it sounds, because it means there's a key somewhere capable of turning every stored password back into plaintext. The correct claim is "we hashed the passwords with bcrypt," which describes a system where recovering the original isn't just hard, it's structurally impossible.
The second common mixup is treating Base64 encoding as if it provides confidentiality — "we encoded the API key before putting it in the config file" — when Base64 provides zero protection against anyone who bothers to decode it, which takes one line of code or a five-second web search.
The third is using a general-purpose hash function for passwords because "SHA-256 is secure" — true for integrity checking, badly wrong for password storage, precisely because speed is a feature for checksums and a vulnerability for anything an attacker might want to brute-force offline.
The short version
If you need to get the original data back, and you control who's allowed to, that's encryption. If you need to represent data in a format another system can handle, and there's no secret involved, that's encoding. If you never need the original back — you only ever need to verify or compare — that's hashing, and for passwords specifically, that means a slow, salted, purpose-built hash, not a general-purpose one and never encryption at all.