Approving a Login on Your Phone for a Smart TV Is an OAuth Flow

You buy a television. You open the YouTube app. It shows you a URL and eight characters — BQDJ-MXWP — and asks you to type them into your phone.

You buy a television. You open the YouTube app. It shows you a URL and eight characters — BQDJ-MXWP — and asks you to type them into your phone.

Thirty seconds later the TV is logged in. Nobody typed a password into a remote control, the credential never touched the television, and the whole exchange worked over a channel the TV could not read.

That is RFC 8628, the OAuth 2.0 Device Authorization Grant, and it is the most interesting flow in OAuth for a reason that has nothing to do with televisions. Every other OAuth flow assumes a browser redirect: the client hands the user to the authorization server and the authorization server hands them back. The device flow is the one where that handoff is impossible, and watching what has to be rebuilt in its absence tells you what the redirect was actually doing for you.

Why it exists, and why "no keyboard" is the smaller half

The obvious motivation is input. A TV remote has a D-pad. Entering k7$Qm!tvR9 on an on-screen grid keyboard takes ninety seconds and three mistakes, and if the user has a passkey or a hardware token there is no way to use it at all — the TV has no authenticator, no platform biometric, and no browser to broker WebAuthn.

The less obvious motivation is the screen. A television is a shared, observable, untrusted display. It is in a living room, a hotel lobby, a conference hall, a shop window. Whatever appears there is visible to everyone in the room and, increasingly, to the room's camera. You do not want a password prompt on it, and you very much do not want the account holder's email address, MFA challenge, or recovery options on it either.

So the device flow makes a clean split: the untrusted device gets a display-only role, and all of the security-sensitive interaction moves to a device the user already trusts and has already authenticated on — their phone. The phone has the password manager, the biometric sensor, the existing session cookie, the ability to do real WebAuthn. The TV gets a code.

This is why the flow generalizes so far beyond TVs. CLI tools use it (gh auth login, aws sso login, kubectl oidc-login) not because terminals lack keyboards, but because a terminal is a bad place to conduct a federated login with MFA — and because piping a password through a CLI process is exactly the thing you spent a decade teaching people not to do. Printers, POS terminals, VR headsets, IoT provisioning, and kiosk enrolment all land in the same category: the thing that needs the token is not the thing that can safely conduct the login.

The shape of the flow

sequenceDiagram
    participant TV as Device (TV/CLI)
    participant AS as Authorization Server
    participant User as User on phone
    TV->>AS: POST /device_authorization (client_id, scope)
    AS-->>TV: device_code, user_code, verification_uri, interval, expires_in
    TV->>User: Display "go to example.com/activate, enter BQDJ-MXWP"
    loop every `interval` seconds
        TV->>AS: POST /token (grant_type=device_code, device_code)
        AS-->>TV: authorization_pending
    end
    User->>AS: Opens verification_uri, authenticates, enters code
    AS-->>User: "Approve sign-in to YouTube on Living Room TV?"
    User->>AS: Approves
    TV->>AS: POST /token (device_code)
    AS-->>TV: access_token, refresh_token

Two codes come back from that first request, and the distinction between them is the entire security design of the flow:

  • device_code — long, high-entropy, secret. Never displayed, never typed. It lives only in the device's memory and is the bearer credential the device uses to claim the resulting token.
  • user_code — short, low-entropy, human-transcribable. Displayed prominently, typed by the user, and not sufficient on its own to obtain a token.

The device polls with the device_code. The user authorizes with the user_code. Because the token can only be collected by the party holding the long secret, shoulder-surfing the eight characters off a lobby screen does not get you a token. It gets you the ability to approve someone else's pending request, which is a different problem — and a real one, which we'll come back to.

Why polling, and not a callback

The first question every engineer asks is why the device sits there hammering the token endpoint instead of the authorization server calling it back when the user approves. Polling feels primitive. It is, in fact, the only thing that works.

A callback requires the authorization server to open a connection to the device. A television sits behind residential NAT with no reachable address, no public DNS name, no TLS certificate, and possibly no inbound path at all. A CLI tool on a laptop on a corporate VPN is worse. Web OAuth solves this by using the user's browser as the delivery vehicle for the redirect — the browser is a component both parties can reach. The device flow has no shared browser: the TV's "browser" and the phone's browser are different user agents in different security contexts, and nothing joins them.

You could push the result down a WebSocket the device holds open. Some vendors do, as an optimization. But that is a proprietary channel with its own scaling and reconnection semantics, and it doesn't remove the need for the polling path as a fallback — so RFC 8628 standardizes the thing that always works and lets you layer the optimization on top.

Polling brings its own contract, which implementations get wrong constantly. The authorization server returns an interval (typically 5 seconds), and the device is required to respect it. If the device polls faster, the server returns slow_down, and the device must increase its interval — the RFC's language is that the client adds 5 seconds to its current interval, not that it retries once and carries on. There are shipped clients that treat slow_down as a transient error and retry immediately, which converts a well-behaved flow into a self-inflicted denial-of-service against your own token endpoint. If you are building the server side, rate-limit per device_code and expect to have to.

The four terminal-ish responses are worth memorizing because they are the whole state machine: authorization_pending (keep going), slow_down (keep going, slower), access_denied (the user said no — stop, don't retry), and expired_token (the device_code timed out — start over with a fresh request).

The short code is a deliberate entropy compromise

This is the part of the flow that repays close attention, because it is a place where the spec knowingly trades cryptographic strength for human factors and then spends its remaining budget on mitigations.

A user_code has to be read off a screen across a room and typed on a phone. That constrains it to roughly 6–9 characters. It also constrains the alphabet: RFC 8628 suggests a case-insensitive set of 20 letters that excludes vowels and visually confusable characters. Excluding vowels is not an aesthetic choice — it prevents the generator from producing profanity in any of the several dozen languages your users speak, which is a real support-ticket category. Excluding 0/O and 1/I/l prevents the transcription failures that make users think the flow is broken.

Do the arithmetic on what that leaves you. An 8-character code from a 20-symbol alphabet is 20⁸ ≈ 2.6 × 10¹⁰, about 34 bits. That is not a secret by any modern standard. A 6-character code is 20⁶ = 64 million — 26 bits, which a single machine can enumerate in an afternoon if you let it.

And the code space isn't the number that matters. What matters is the live code space: how many codes are simultaneously valid. If you have 100,000 pending authorizations at peak and a 26-bit space, a random guess hits a live code roughly one time in 640. An attacker who can make a few hundred thousand guesses is not brute-forcing a specific victim — they are trawling for anyone's pending device authorization, and the flow hands them whatever account approves next.

So the entropy has to be bought back somewhere else, and the RFC is explicit about where:

Short lifetime. expires_in should be minutes, not hours. Ten minutes is typical; it shrinks the live set and caps the guessing window.

Rate limiting on the verification endpoint. The code-entry page is the brute-force surface, and it is the one people forget to protect because it doesn't look like a login form. It needs limits per IP, per session, and globally — plus a CAPTCHA or equivalent once a client starts producing failures. Global limits matter here in a way they don't for password login: an attacker guessing at the code space rather than at an account is invisible to per-account controls.

Single-use and immediate invalidation. A code that has been submitted — successfully or not — should not remain guessable. Some implementations burn the code after a small number of failed attempts against it, which is cheap and closes the "keep hammering one code" path.

Never authorize on entry alone. The verification page must require an authenticated session and an explicit approval step. If entering a valid code silently completes the grant, the code is your entire authorization decision and 26 bits is what protects the account.

Longer codes are, of course, available: if your device can render a QR code, verification_uri_complete (from the same RFC) lets you embed the code in a URL the phone camera scans, and now the transcription constraint is gone and you can use a 128-bit value. This is what most streaming apps and modern CLIs do, and it is strictly better where it's available. Keep the short code as the fallback for the user whose phone camera won't focus — but recognize that you are then running two flows with very different entropy, and the short one sets your security floor.

The attack the flow is genuinely vulnerable to

Everything above concerns an attacker guessing codes. The real-world attack runs the other way, and it is worth understanding precisely because it is not fixable with entropy.

The attacker initiates the device flow themselves, against your legitimate authorization server, with a legitimate client_id. They receive a real device_code and a real user_code. They then send the victim a message: "Your Microsoft 365 session needs re-verification. Go to microsoft.com/devicelogin and enter code BQDJ-MXWP."

Every element the victim can check is genuine. The domain is correct. The TLS certificate is correct. The page is the real authorization server. There is no lookalike domain to spot, no credential harvesting proxy, nothing that a password manager's domain binding would refuse to fill, and — critically — nothing that phishing-resistant authentication prevents. The victim authenticates with a passkey, correctly, to the real site, and then approves a request that originated on the attacker's machine. The attacker's device polls, collects the token, and now holds a session for the victim's account.

This class of attack has been used at scale against Microsoft 365 and Google Workspace tenants. It is the reason device-code phishing gets its own detection rules in most EDR and identity-protection products, and the reason several vendors now let you disable the device grant per-application by policy.

What actually mitigates it:

Make the consent screen describe the device, in the user's language of trust. Not "Approve sign-in for client a1b2c3." Something closer to "A device in Frankfurt, Germany is requesting access to your account. If you are not currently setting up a device, do not continue." Show the requesting IP's geolocation, the user agent if you have one, the time, and the scopes in plain words. The victim's only defence is noticing that they didn't start this, so the screen's job is to make that noticing easy.

Compare the two contexts. You know the IP and rough location of the device that requested the code and of the browser approving it. When those disagree materially — different countries, hosting-provider ASN on one side and residential on the other — that is a strong signal. It is not proof (legitimate hotel-TV and VPN cases exist), so treat it as a reason to escalate the warning or require step-up, not to hard-fail.

Restrict the grant. The device grant should be enabled per-client, and only for clients that are actually input-constrained. If your tenant has no smart-TV app and no CLI, the grant should be off. Most breaches in this class involve an attacker using a first-party client ID that the tenant never intended to allow the device flow for.

Bind and rate-limit at the tenant level. Alert on device-authorization requests arriving from datacenter IP ranges, and on a single account approving multiple device codes in a short window.

There is a structural alternative worth knowing about: CIBA (Client-Initiated Backchannel Authentication, from OpenID Connect). In CIBA the client names the user it wants to authenticate, and the authorization server pushes a notification to that user's registered device. Because the server chooses who gets asked, the attacker cannot redirect an approval prompt to an arbitrary victim — they'd have to name the victim, and the notification arrives unsolicited in a channel the user associates with their own actions. CIBA solves a slightly different problem (call-centre and high-assurance transaction flows), needs a registered authenticator per user, and is nowhere near as widely deployed. But if you are choosing between them for a new high-value use case, that difference in who initiates is the thing to weigh.

Implementation details that bite

The device polls before the user has finished, for minutes. Every poll is a full token-endpoint request with client authentication. A million televisions being set up on Christmas morning, at 5-second intervals, is a load pattern worth capacity-planning for — and the reason interval is server-controlled is so you can raise it under pressure. Make sure yours is configurable at runtime and not a constant in the client SDK.

Public clients still need PKCE-equivalent thinking. A TV app is a public client; its client_id is not a secret and its binary is extractable. The device_code is what protects the exchange, which means it must be generated with a CSPRNG, be at least 128 bits, and never be logged. device_code values landing in access logs or APM traces is a live vulnerability, not a hygiene issue — anyone with log access can complete a pending authorization.

Refresh tokens on devices you cannot reach. The device flow's whole purpose is to authorize something long-lived, so it issues a refresh token to a device that has no user present to re-authenticate. That refresh token is now the account's exposure for as long as the TV lives — including after the TV is sold, or left in the hotel room. Two things follow: refresh tokens issued via the device grant deserve rotation and reuse-detection, and your account UI needs a per-device revocation list that a user can actually understand ("Samsung TV, living room, last used yesterday"). The device flow's failure mode isn't the login; it's the eight-year-old grant nobody remembers making.

Scope should be narrower here, not wider. The device is the least trustworthy holder of a token in your system. It sits in physical spaces you don't control, runs firmware that stops receiving updates in three years, and cannot be re-consented interactively. If a TV app needs to read a watch history and play video, it should not hold a token that can change the account's email address.

What the flow teaches

Strip the television away and the device grant is a general pattern: split a transaction across two channels when one of them cannot be trusted with the whole thing. One channel carries a low-entropy human-readable reference; the other carries the high-entropy secret and collects the result. The reference is deliberately weak because a human has to move it, and every other part of the design exists to compensate for that weakness.

Once you see it in that shape, you'll notice it elsewhere — in payment confirmation codes, in Bluetooth pairing, in the six digits your bank reads out on the phone. And you'll notice that in every one of those systems, the interesting failures are not in the cryptography. They're in whether the human on the trusted channel understood what they were approving.

That's the lesson worth taking from RFC 8628. The flow is a triumph of pragmatic design, and its one serious weakness is the one place where engineering stops and a person has to make a judgment. Design the consent screen as if that judgment is the only control you have, because in the phishing case, it is.