Writing /
Interviewing SREs in Tokyo: what I actually ask
I’ve sat on the hiring side of something like fifty SRE interviews in the last two years, most of them in a small meeting room with the aircon fighting a Tokyo September and losing. It’s still 33 degrees outside as I write this. Candidates arrive ten minutes early, apologize for the heat as though they arranged it, and accept a can of coffee from the machine down the hall. I like that ritual.
I’ve also gotten much less clever about the loop itself, which is what this post is about.
The market you’re actually hiring in
A few structural facts shape everything here, and none of them are complaints.
The funnel splits on language before it splits on skill, and the overlap between the Japanese-required pool and the English-fine pool is small enough that everybody is hiring the same forty people. The senior pool is also smaller than it looks, because a lot of the strongest operations people in this city have never had “SRE” in a title. They came out of on-prem infrastructure, or datacenter work, or a corporate 情報システム team where they quietly ran a rack room with three people and no Kubernetes. Screen on title vocabulary and you filter out exactly the people who’ve been woken up the most.
Tenure norms differ too. Eight years at one company is normal here, sometimes loyalty and sometimes a sign of never having seen a second way to do anything, and the interview has to tell you which. And candidates are, almost without exception, prepared and respectful. They rarely trash a previous employer, which is pleasant but means you work harder to hear about failures, because the cultural default is to under-claim. I’ve learned to ask about mistakes twice, in different words, and to sit through the pause.
What I dropped
LeetCode-hard for ops roles
I’m the last person on earth who gets to be sniffy about contest problems. I love them and still do them for fun. That’s exactly why I stopped asking. The best on-call engineer I’ve ever worked with would have been visibly annoyed by a segment tree question and would not have finished it, and one of the most confident whiteboard performers I ever passed through a loop froze the first time a real incident had no known answer. What I was measuring was whether someone had practiced a specific sport recently, which tells me about their last three months of evenings, not about how they behave when the graph goes flat.
I kept one small exercise: read a log format, aggregate it, handle the malformed lines. Twenty minutes, no exotic data structures.
Port numbers and protocol trivia
“What port does etcd listen on.” “Name the TCP states in a handshake.” I thought those proxied for depth, and they proxy for recall, specifically for having crammed the week before. Worse, in a room where at least one of us is working in a second language, trivia punishes retrieval speed under stress and calls it knowledge. Everyone I want to hire looks that up. The good ones look it up fast and know which answer smells wrong, and I can test that a better way.
What I ask instead
Walk me backward through an incident you owned
Backward is the whole trick. Start from the fix and work back to the first signal you saw.
Told forward, an incident becomes a story, and stories clean themselves up. They acquire a protagonist, and the dead ends fall out because dead ends are boring to tell. Told backward, you get the actual search. Why did you look there? What made you rule the deploy out?
I’m listening for the “so then I checked” moves, and for cost-awareness in the ordering. A strong answer sounds like: “I checked the deploy log first because it’s one click and it’s the usual cause, it was clean, so I went to the dependency dashboards, and the auth p99 was flat, which surprised me, so I stopped trusting the dashboard.” That’s someone who ranks hypotheses by how cheap they are to eliminate.
I do not care whether they fixed it. Some of the best answers I’ve heard ended in “and then a vendor fixed it and I still don’t fully know what happened,” from someone who could name which four things they’d eliminated and why. That beats a clean resolution with no visible search.
One capacity estimate, out loud
I pull a number out of our own estate and ask them to Fermi it aloud. We ingest a couple of terabytes an hour of logs and want ninety days of it: how much storage, and what does it cost?
Order of magnitude is the pass mark; the process is what I’m watching. Do they say the assumption out loud before using it. Do they write anything down. Do they sanity-check against something they know from real life (“that’s about a rack, which feels right” or “that’s more than our entire cloud bill, so I’ve messed up somewhere”).
The failure mode isn’t bad arithmetic, it’s refusing to guess. Some candidates ask three clarifying questions and then say they’d need real numbers. Defensible in a meeting, fatal at 2am when somebody has to say an approximately right number inside ninety seconds. I tell them guessing is permitted, and half of them are fine once they know that.
Here’s a real alert rule, critique it
I bring a real PromQL rule out of our repo. It’s one of mine, from early on, and it’s bad in four distinct ways, which is why I like it.
I want to hear: does this deserve to wake a human, and what does that human do next. What’s the evaluation window, and what does it do during a rolling restart. Does the threshold describe something a user can feel, or an internal number that merely correlates with one.
The best answers get there and then ask something I didn’t plant: “what’s the runbook link on this?” My favorite response ever came from a candidate who read it and said, “I don’t think this should page at all. I’d make it a ticket and tell you why in the PR.” She was right. Pushing back on the premise of my own bad alert is the strongest pass available.
Tell me about a system you deleted
Nobody prepares for this one, which is exactly why I keep it last.
It sorts hard. People who’ve only built things go quiet, or reach for a rewrite and call it a deletion. People who’ve owned something long enough to see the end of its life light up, because removing a system in a company with other humans in it is hard work nobody gets credit for. How did you find the callers. How did you prove the traffic was gone, and how long did you watch. Did you ever actually delete the copy you kept.
Someone who has removed a load-bearing thing has shown me ownership, patience, stakeholder work, and measurement in one story, without being asked about any of those words.
Red flags
The hero arc. “I was the only one who understood the system.” Sometimes true, and even then it describes a failure being presented as a strength. I follow up with “what did you do about that,” and the answer sorts people fast.
Blame. Not frustration, which is fine and human, but a story where every problem was upstream of the speaker. Ops work is a permanent condition of inheriting other people’s decisions, and if nothing in three years was partly theirs, they either weren’t close enough to it or won’t tell me when they break something.
Zero questions about on-call, the flag I weight most heavily. If a senior candidate asks nothing about rotation size, page volume, escalation, or what a bad week looks like, either they don’t intend to carry the pager or they’ve never been hurt by a bad rotation and don’t yet know that’s a thing that can hurt you.
Green flags
Curiosity about the boring parts. The questions that make me sit up are all unglamorous. How long does CI take. How noisy are the alerts, honestly. What happens when the deploy is bad at 6pm on a Friday.
One candidate asked how long our pipeline took, I told him, and he asked why. Not to be difficult, genuinely interested. We spent six minutes on test parallelism and I learned something. People who ask about the boring parts intend to improve them.
The one I still think about
A candidate blew the capacity estimate. Badly, gigabytes where she wanted terabytes, and she moved on without noticing. I let it go, because one arithmetic slip wasn’t going to decide this.
Twenty minutes later, mid-answer to an unrelated question about deploy strategy, she stopped. “Sorry, I need to go back. Earlier I said gigabytes. That was wrong by a factor of a thousand, and if you’d used my number you’d have bought a fraction of the disk you need.”
Nobody prompted her. Some part of her brain had kept that thread open for twenty minutes and then interrupted her own answer to close it.
That’s the job. Not the estimate, not even the recovery, but the background process that keeps chewing on a number that didn’t feel right. We hired her. She’s caught two things this year that nobody else was going to catch.
What the loop is for
An interview loop samples one afternoon of someone’s life, in a room they’ve never been in, in a language that may not be their first, with a stranger deciding their next two years. It’s a terrible measurement apparatus and there’s no version of it that isn’t. All I’ve really done is stop asking questions that measure interview performance and start asking ones where the honest answer and the impressive answer are the same sentence.
Because what I’m evaluating isn’t whether someone can be brilliant while being watched. It’s whether they’re the person I want on the other end of the call at 3am, in a Tokyo January, when neither of us knows what’s wrong yet.