In a room, I read your page upside down. Before you look up from your issue tree I know whether it has two branches or four, whether you wrote "fixed" and "variable" or just "costs", and whether the pen stopped because you were thinking or because you were stuck. On a video call I know none of that. I see a face in a rectangle and I hear a voice.
That is the whole difference between a virtual case and a live one, and most candidates get it wrong. They treat the Zoom case as the same case with worse Wi-Fi and spend their preparation on ring lights. The real change is that the interviewer has lost every channel except your voice, so anything you wrote down and did not say out loud did not happen, as far as the scoring sheet is concerned.
By the end you will be able to measure your own silence, narrate structure and math without slowing down, choose your medium before anyone asks, and tell published virtual rules from forum folklore.
What the interviewer loses when the case moves to a screen
Scoring a case in person uses four channels: your voice for the reasoning, your page for the structure and arithmetic, your body for whether you are working or frozen, and the exhibit on the table, where I can watch your eyes. Video keeps the voice, slightly delayed, and strips the other three down to a face from the collarbones up. Everything the page and the pen used to tell me now has to arrive as words.
A 2016 meta-analysis of twelve studies found interviewer ratings in technology-mediated interviews lower than face-to-face by a moderate effect (d = -0.41), larger in real hiring settings (d = -0.59) than in labs, with video and telephone tied for worst medium (d = -0.46). Most pooled studies date from 2001 to 2004, so treat the size of the gap as dated. The direction surprises nobody who has scored both formats: candidates who rely on being watched lose the most.
Bain says it in plainer words. Its case-interview page tells virtual candidates to "let your interviewer know when you're taking notes and clearly describe your structure". That is not etiquette. It is the firm telling you which channel is still open.
The silence budget: measure yours before the interviewer does
Here is a calculation you can reproduce with a stopwatch. Take a 25-minute candidate-led case and add up the moments when a well-prepared candidate is working but not talking. My estimates, not a study.
| Moment | Silence | In the room, I see | On video, I see |
|---|---|---|---|
| Building the structure | 90 s | A tree taking shape | A face looking down |
| Two pauses before answers | 45 s each | Pen moving | A frozen frame |
| Main math block | 150 s | Columns of numbers | Nothing |
| Preparing the synthesis | 30 s | Underlining | Nothing |
| Total | 360 s | Six minutes of visible work | Six minutes of dead air |
Six minutes is 24 percent of the interview. In person that quarter is evidence in your favor. On video it is a quarter of the case in which I have no information about you, and my attention drifts to the chat window, the scoring sheet in another tab, and my fourth call of the day.
You do not fix this by talking constantly. You fix it by labeling every pause and narrating headings instead of paragraphs.
- Before the structure. "Give me about a minute. I'm splitting this into revenue and cost." Then take the minute in silence. The label turns dead air into a promise with a deadline.
- During the structure. Say the branch names as you write them, nothing more: "Revenue: price, volume. Cost: fixed, variable." Five seconds per branch, and the tree can be scored.
- Before math. "Two minutes on this. Stop me if I go wrong." Then narrate the calculation itself.
- Before the synthesis. "Thirty seconds to pull this together." Then answer first, reasons second.
Ned's rule. On video, a pause you have labeled is thinking; a pause you have not labeled is a connection problem. Say how long you need, then take exactly that long.
Case math when nobody can see your paper
The exhibit arrives in the chat as a screenshot. Northgate Garden Centers, a 40-store chain in the upper Midwest, wants to know where its profit comes from.
| Region | Stores | Revenue per store | Operating margin |
|---|---|---|---|
| Lakes | 12 | $2.5M | 8% |
| Prairie | 20 | $1.8M | 12% |
| River | 8 | $3.0M | 4% |
In the room, most candidates go quiet, fill half a page, and look up with "$7.7 million, blended margin around 8.5 percent." I glance at the page, see the columns, and tick the box. On video the same behavior gives me two minutes of nothing followed by two numbers I cannot check. If one is wrong, I do not know which step failed, and the accuracy box stays empty.
The narrated version takes about the same time:
"Revenue first, then profit. Lakes: 12 stores at 2.5 is 30 million. Prairie: 20 at 1.8 is 36. River: 8 at 3 is 24. Total 90 million. Profit: 8 percent of 30 is 2.4; 12 percent of 36 is 4.32; 4 percent of 24 is 0.96. Total 7.68 million, a blended margin of about 8.5 percent. The so-what: River is 27 percent of revenue and only 12.5 percent of profit. If the question is where the margin problem lives, it is the eight River stores."
Check it: 24 / 90 is 0.267, and 0.96 / 7.68 is exactly 0.125. Every step was spoken with its unit, so the interviewer could have stopped me at any of them. That is what "show your work" means when the work is invisible.
Two more habits. When the exhibit lands, say what you see before you analyze it; if it is too small to read, ask for it again, because guessing a unit costs more than five seconds. And if you are asked to type into a shared spreadsheet, keep narrating. A cell filling in silence is silent math in a worse font.
Paper, whiteboard, or voice: decide before they ask
Ask in the first minute if you have not been told: "Would you like me to share a whiteboard, or work on paper and talk you through it?" Decide in advance what you will do under each.
| Medium | Works when | Fails when | From my chair |
|---|---|---|---|
| Paper, narrated | Almost always; natural speed, no software risk | You forget it is invisible and go quiet | The default. Hold the page up for one number, never to read a tree |
| Digital whiteboard | The interviewer opened it and is watching | You are learning the tool live or writing too small | Fine for a long final; painful when you spend 40 seconds finding the pen |
| Voice only | Phone rounds, or when video drops | You have nothing to point back to | Signpost every move: "back to the cost branch" |
Switch to a board only when the interviewer opens one. Do not surprise a partner with a Miro link.
A point most guides skip: the second monitor. McKinsey's Assessment Integrity Expectations say that in any interview or assessment, including virtual ones, candidates will "only use one screen" and "discard any notes taken", and will not record, screenshot, or use "any applications or websites, generative AI, a calculator, or prepared notes". A whiteboard habit that needs a second display collides with that text. Other firms have published nothing as specific that I could find, which is not the same as having no rule; ask the recruiter.
What the firms actually publish about virtual rounds
Most "virtual interview rules" online are folklore dressed as policy. Here is what the three largest strategy firms had on their own careers pages as of 2026-09-24; "not stated" means I could not find it, not that it is allowed.
| Item | McKinsey | Bain | BCG |
|---|---|---|---|
| Camera | "Keep your camera on" | "Maintain strong posture and engagement on screen" | Not stated |
| Background | Ensure "any virtual backgrounds or blurring features are disabled" | "Neutral (or blurred) background" is fine | Not stated |
| AI tools | Disable "note-taking tools or virtual assistants"; no generative AI, calculators, or prepared notes | Not stated | Not stated |
| Screens | "Only use one screen" | Not stated | Not stated |
| Format | Not stated | Not stated | Varies by office; Switzerland runs a virtual 45-minute first round and three in-person finals |
Sources: McKinsey's interviewing page and integrity page, Bain's case-interview page, and BCG's Switzerland page. Formats change by office and season; the invitation email outranks this table.
Two cells deserve a second look. McKinsey and Bain disagree about blur; if you interview at both in the same two weeks, check which call you are on. And McKinsey's phrase about AI note-takers, "may unintentionally record or assist", is aimed at the assistant you forgot you installed. A meeting bot that auto-joins calendar invites will announce itself before you do. Disable it at the account level, not just the window.
Bandwidth, lag, and the beat you have to leave
Most of the technical part is arithmetic.
Zoom's own system requirements put group video at 2.6 Mbps up and 1.8 Mbps down for 720p, audio at 60 to 80 kbps, and screen sharing without a video thumbnail at 50 to 75 kbps. Your voice needs about 3 percent of the bandwidth your face does. If the connection is marginal, lower the video quality, not the microphone; a case scored on audio does not need you in HD.
Latency matters more than resolution. The ITU's recommendation on one-way transmission time treats mouth-to-ear delay under 150 milliseconds as essentially transparent and 400 milliseconds as the planning ceiling. A consumer connection sits somewhere in that band, so each of you hears the other a fraction of a second late. When the interviewer finishes a sentence, wait one full beat; the gap that feels rude on your side is normal on theirs. When you are interrupted, stop mid-clause. Talking over a partner because of lag is still talking over a partner.
Camera at eye height, light in front of you, notes directly below the lens so a glance at them reads as a glance at a page, not at a second monitor. Then hide your own tile. Stanford's Jeremy Bailenson argued in a 2021 paper on "Zoom fatigue" that the default self-view is an all-day mirror that raises self-evaluation; he calls it an argument rather than a finding, but you do not need a study to know that watching yourself do mental math is a poor use of attention.
Do the ten-minute dry run on the application named in the invitation, from the same chair, with a friend reading a sentence back to you. A guest link into a firm's tenant is not your employer's Teams, and 9:02 on interview day is the wrong moment to learn that the whiteboard is disabled for guests.
The rep to run tonight
Sit at the desk you will use for the real thing, camera on, and run the live voice case on CoachNed: UrbanBrew, about eighteen minutes, seven-score debrief. Record your own side on your phone; this is practice, so no policy is offended. Then do the silence audit: seconds of unlabeled silence divided by case length. Over ten percent means you are still relying on being watched.
With five minutes instead of twenty, the first rep is three typed turns with instant scores and needs no account. The structure drill below is the same muscle under a clock.
Everything is open for seven days, no card; then $120 for a recruiting season or $49 a month. CoachNed is independent and has no affiliation with McKinsey, BCG, Bain, or any other firm named here.
Frequently asked questions
Do virtual case interviews score lower than in-person ones?
The best public evidence, a 2016 meta-analysis, found lower interviewer ratings in technology-mediated interviews (d = -0.41), with video among the worst media. Its studies are mostly from 2001 to 2004, so the gap has probably narrowed, and you are compared with other Zoom candidates, not with an in-person version of you.
Can I use a second monitor in a virtual case interview?
At McKinsey, no: its Assessment Integrity Expectations require candidates to "only use one screen". No other firm publishes a comparable rule that I could find, but eyes drifting off-camera read as reading, so one screen is the safer habit.
What happens if my internet drops during a case interview?
Say so, switch to phone audio if the platform offers it, and ask where to pick up. A twenty-second reconnect costs nothing on the scoring sheet; ninety seconds of apologizing costs you the thread of the case.
Can I blur my background in a consulting interview?
McKinsey asks that "any virtual backgrounds or blurring features are disabled"; Bain's case page says a "neutral (or blurred) background" is fine. A plain wall needs no policy.
Sources
- Interviewing at McKinsey — camera on, disable AI note-takers, disable backgrounds and blur. Checked 2026-09-24.
- McKinsey Assessment Integrity Expectations — one screen, discard notes, no recording or generative AI. Checked 2026-09-24.
- Preparing for the Case Interview, Bain & Company — blurred background allowed, narrate notes and structure. Checked 2026-09-24.
- Application and interviews at BCG Switzerland — virtual first round, in-person finals. Checked 2026-09-24.
- Blacksmith, Willford, and Behrend (2016), Technology in the Employment Interview: A Meta-Analysis — d = -0.41 on ratings, moderators, date caveat. Checked 2026-09-24.
- Zoom system requirements: Windows, macOS, Linux — bandwidth for video, audio, and screen sharing. Checked 2026-09-24.
- ITU-T Recommendation G.114, One-way transmission time — 150 ms and 400 ms delay thresholds. Checked 2026-09-24.
- Bailenson (2021), Nonverbal Overload: A Theoretical Argument for the Causes of Zoom Fatigue — self-view as an all-day mirror. Checked 2026-09-24.
