A credential issued in 2022 for a limited pilot integration at Klue sat active for roughly four years after the pilot itself was quietly abandoned. Nobody owned it. Nobody reviewed it. Nobody put a date on the calendar for it to stop working, and that last part is the whole story. When the extortion group Icarus went looking for a way in, that single dormant token turned into the entry point for the Salesforce environments of close to 200 companies, according to TechCrunch’s reporting on the incident. There was nothing sophisticated about the attack itself, no zero-day, no clever pivot. It was a deviation from standard access practice that simply outlived every reason it was ever granted.
And it’s not a one-off story, much as I’d like to file it under “freak accident” and move on. Black Kite’s 2026 Third-Party Breach Report tracked 136 major third-party breach events last year that hit 719 companies directly, with an estimated 26,000 additional downstream companies affected but never individually named, because disclosures only reported impact in aggregate. The average gap between breach and public disclosure ran 117 days, per Black Kite’s report. Multiply the Klue scenario across a portfolio of vendor integrations, unpatched dependencies, and legacy access grants and this is what you get: risk that got accepted once, informally, and never came back up for a second look.
The sentence versus the system
“All exceptions must be documented and time-boxed” reads fine in a standards document. It says nothing about who submits a request, what information is mandatory, who has authority to approve it at what risk tier, how often it comes back for review before it expires, or, most importantly, what happens automatically when the clock runs out. Leave those mechanics unwritten and “time-boxed” ends up meaning whatever the requesting team feels like six months later, which in practice is indefinite. I’ve watched that gap close in real time, on a board, over a request that looked completely reasonable in the room.
I sit on the board that adjudicates the exceptions that clear our architecture review board’s blast-radius threshold, and the recurring failure isn’t bad-faith requests, I want to be clear about that. It’s exceptions that were entirely reasonable at the moment of approval and never got revisited because nothing in the process forced a revisit. A usable intake form has to do more than collect a justification paragraph. At minimum, it needs:
- The specific standard or principle being deviated from, named, not “general security concerns”
- A named individual owner, not a team or a distribution list, accountable for the exception’s status
- The business or technical justification, and why the standard can’t be met on the normal timeline
- The compensating control in place for the duration of the exception, not just an acknowledgment that risk exists
- A blast radius assessment: what systems, data, or customer scope this touches if the underlying risk is realized
- A hard expiration date, set at approval, not “to be determined later”
Cadence, ownership, and the control everyone skips
The compensating control line is the one teams leave blank most often, and it’s the one that matters most, which tells you something about where the discipline actually breaks down. An exception that accepts risk with no mitigating control isn’t an architecture decision, it’s a shrug with a due date attached. If there’s no meaningful compensating control available, that’s a signal the request belongs in full risk acceptance with executive sign-off, not a routine exception.
Risk tier should set both who approves and how often the exception gets reviewed before it expires. A low-blast-radius exception might warrant a single check-in at the 30-day mark, and that’s usually fine. Anything touching customer data, external access, or a regulated system needs a real review, not a rubber stamp, at 90, 60, and 30 days out, with the reviewer required to either re-justify with fresh information, close it because the underlying gap is fixed, or escalate it if it needs to persist past its original scope. Renewal should never be a checkbox that regenerates the same expiration date with the same stale justification, and I’ll admit I used to think a lighter-touch renewal was fine as long as someone technically clicked approve. It isn’t. If the reason for the exception hasn’t changed since approval, that’s usually a sign the underlying problem was never actually being worked, just tolerated.
The default when the clock hits zero
Here’s the part that turns this from a formality into an actual control: what happens by default the moment the clock hits zero. If the answer is “nothing, it just stays open,” you don’t have a time-boxed exception, you have a new unwritten standard. The default on expiry has to be closure, revocation, or automatic escalation, never silent persistence, which means the exception needs to live somewhere capable of enforcing that state change rather than a spreadsheet nobody reopens after the approval email goes out.
That’s why I track exceptions inside LeanIX as our EAM system of record rather than in a GRC ticket that closes and gets forgotten. Every exception links to the architecture component or capability it deviates from, carries its own lifecycle state, and shows up on the same portfolio view the board uses for everything else. When an exception expires without action, it’s visible in the same tool where we track fact sheets and decommission status, not buried in an inbox. That visibility is the control, full stop. A policy that says “time-boxed” without a system that surfaces the countdown is a policy that gets quietly ignored the first time a delivery deadline is tight, and it will be tight, it’s always tight.
An exception process earns its keep by making risk acceptance visible, owned, and finite. Think of it the way you’d think of a carton of milk left in the back of the fridge: the date on the label doesn’t do anything by itself, someone still has to open the fridge, check it, and throw it out. The Klue credential wasn’t a failure of policy language. A standard existed somewhere that said temporary access should be revoked, I’d bet money on it. What was missing was the mechanism that made revocation the thing that happened automatically instead of the thing somebody was supposed to remember.
Photo by Markus Spiske on Unsplash