The question I hear in every engagement is some version of the same thing: are we exposed? The honest answer requires a real assessment. Real assessments feel like six-week consulting engagements. So the question stays open, the worry persists, and nothing gets done. I've watched that loop run in organizations that had already experienced an incident. The gap between "I know we have a problem" and "I know what the problem is" was still too wide to cross.
So we built the smallest possible thing that closes that gap.
What We Were Trying to Build
The constraints were strict. It had to run conversationally, one question at a time, so a non-specialist could complete it without a glossary. It had to produce a score, because a number you can track beats a paragraph you skim. It had to be honest: surfacing the gaps that actually cause incidents, not the ones that are comfortable to discuss. And it had to work for a stranger who downloaded it with none of our internal context, because a tool that only runs in its author's environment isn't a tool. It's a script.
The mental model underneath it is one I keep coming back to: security isn't a castle wall, it's a set of zoning laws. The auditor doesn't ask whether your perimeter is high. It asks whether you have rules about which data your own employees are allowed to put into AI tools, because that is where most incidents actually start.
The Six Dimensions, and Why Each One
The skill scores six things, and each one is the source of a real class of failure I've seen in the field.
- Data classification: you cannot protect what you haven't labeled. In every audit I've run, this is where I find the first crack.
- Shadow AI: the tools in active use that nobody approved. These almost always outnumber the sanctioned ones, and they're processing your data under terms nobody reviewed.
- Usage policy: the constraint layer that defines what data goes where, under what conditions, with what audit trail.
- Access controls, incident response, and governance ownership: the last one is the question that ends most meetings, which is exactly why we built it in. Who is actually accountable for AI security specifically, not IT security generally?
Each dimension is a single question with four answer levels worth one to four points. We chose four deliberately. Three collapses into a vague middle. Five invites everyone to pick the center. Four forces the only distinction that matters: the gap between "we have this on paper" and "we've actually operationalized it."
What Broke in Testing, and What It Taught Us
Two things broke, and both were worth breaking.
The first version assumed our environment. It referenced internal frameworks and folder structures that a stranger wouldn't have, so it failed silently when someone outside MBL ran it. We rebuilt it to be fully self-contained. The generalizable lesson: a tool that assumes its author's context is a personal script. A product works without you.
The second version was too polite. A team scoring badly on data classification got a gently worded note about areas for improvement. That's not an honest read. That's flattery. We rewrote the output to lead with the exposure and name the cost in plain language. The tradeoff is that the output can sting. That's the point. An assessment that softens the truth is worse than no assessment at all, because it lets people believe they've done the work.
What Running It Actually Does
The score is not the value. The value is the decision the score forces. Either you build your security posture in parallel with adoption, or you build it after an incident. The second path costs multiples of the first. A twenty-minute diagnostic that tells you exactly where you're exposed is the cheapest entry point into that decision you'll ever find.
We're giving it away because the build is the argument. Run it. Judge the output. Then ask what else in your security posture is one afternoon away from being a system instead of a worry.
Three Questions to Ask Before You Trust Any AI Tool You Build
1. Does it assume context the end user will not have?
2. Does it tell the user the comfortable thing or the true thing when the two diverge?
3. Can a non-specialist complete it without help, or does it secretly require an expert to interpret the output?
Download the skill: AI Security Posture Check → Get it on GitHub. Install in Cowork: download the file, drag it into your Cowork window.
MBL Partners · mblpartners.com · [email protected]