AI Compliance Fails When Questions Reward Prose Not Proof

Learn why current AI compliance questionnaires reward creative writing over evidence and how standardized checklists can fix the problem.

By Central
Current AI compliance methods fail because they ask for prose instead of requiring proof, like surgical checklists that save lives.
Highlights
  • Current AI security questionnaires reward creative writing over evidence, failing to reduce actual risk.
  • Standardized model cards, similar to surgical checklists, can streamline AI compliance assessments.
  • The EU AI Act and ISO 42001 require documentation that could be satisfied by a common model card format.

In 2009, surgeon Atul Gawande and a team backed by the World Health Organization demonstrated that a 19-item surgical checklist could slash complication and death rates dramatically across eight hospitals worldwide. Not a thousand-page protocol. Not a comprehensive framework. Nineteen items, printed on a single card. Aviation learned the same lesson decades earlier: the pre-flight checklist fits in a pilot’s hand, not in a binder. Nearly two decades later, security teams send AI vendors questionnaires with 300 questions, half of which begin with “describe your approach to…” and almost none of which would catch a real failure. We have the frameworks. What we don’t have is the checklist.

The timing matters. The EU AI Act’s enforcement teeth for general-purpose AI arrive this August, high-risk obligations phase in behind them, and ISO/IEC 42001 now appears by name in third-party risk questionnaires. NIST’s AI Risk Management Framework has become the default answer for “show me you have an AI risk program” in North America. Add the OECD Principles, HITRUST’s AI assurance work, sector regulators like the FDA, and a growing patchwork of US state laws, and most enterprises operate under two or more frameworks simultaneously. The frameworks themselves largely agree. Published crosswalks show substantial overlap between ISO 42001, NIST AI RMF, and the EU AI Act. An organization that builds its program thoughtfully can satisfy all three with a single set of processes and documentation. The problem isn’t the frameworks. The problem is what happens downstream, when those frameworks translate into the questionnaires, audits, and attestations that land on real desks.

How Current Questionnaires Reward Creative Writing Instead of Compliance

If you’ve been on the receiving end of an AI security questionnaire lately, you know the artifact being described. Hundreds of questions. Free-text answers. Prompts like “Describe how your AI system ensures fairness” or “Explain your approach to responsible AI.” These questions have three fatal flaws.

First, they cannot be answered with evidence, only with prose. And prose isn’t compliance; it’s creative writing. A vendor with a mature program and a vendor with a good technical writer produce indistinguishable answers. The exercise rewards confident fiction and punishes honest uncertainty. When a question can be answered without producing a single artifact, that question is not reducing anyone’s risk.

Second, they ignore the nature of the systems they assess. Large language models are stochastic systems that are extremely difficult to replay and troubleshoot. A point-in-time attestation about model behavior becomes stale the moment a model version changes, a system prompt is updated, or a temperature setting moves. Asking “Does your model produce biased outputs?” as a yes/no compliance question fundamentally misunderstands what these systems are. The right question is whether you can measure it, log it, and show the trend.

Third, they do not scale with risk. The same 300-question addendum is sent to a vendor running a marketing chatbot and a vendor deploying clinical decision support. When everything is high-risk, nothing is. The EU AI Act got at least this much right: its four-tier risk classification exists precisely so obligations scale with consequences. Most homegrown questionnaires have no tiers at all.

Five Tests Every AI Compliance Question Must Pass

A framework is only as good as its worst question. Before any question makes it into an AI assessment—whether a vendor questionnaire, an internal review gate, or an audit program—it should pass five tests.

1. Answerable with an artifact. Every question should map to evidence: a log, a config, an eval report, a data flow diagram, an architecture document. If the only possible answer is an essay, cut it or rewrite it. “Describe your approach to model security” becomes “Provide your logged inference parameters (model version, temperature, top P, token limits) for your production deployment.”

2. Scoped to risk tier. Classify the system first, then ask the questions that tier deserves. A limited-risk internal tool gets ten questions. A system touching patient data or financial decisions gets the full treatment. Do not send Annex IV-depth documentation requests to a chatbot vendor.

3. Measurable or binary where possible. “Do you run evals? At what cadence? What was your last pass rate on your safety benchmark?” beats “Describe your testing philosophy” every time. Reviewers can score it, trend it, and compare it across vendors.

4. Decision-relevant. For each question, ask: if the answer came back bad, would it change the decision? If removing the question would not change any outcome, remove the question. This single test eliminates half of most questionnaires.

5. Mapped once, reused everywhere. Build your control set once, then use the published crosswalks to answer NIST, ISO, and EU AI Act asks from the same evidence base. One process, multiple regulatory readings. If teams produce separate documentation for each framework, they pay a triple tax on the same work.

The Questions That Actually Matter for AI Vendor Assessment

If AI vendor assessment were compressed down to a single card, it would look something like this: Where is the model deployed, and who is the upstream provider? What data flows in, and what flows out, and where is that logged? Where are copies of that data stored, for how long, and who can access them? What inference parameters are logged per request, and can you replay an incident? What is your eval suite, what does it cover, and how often does it run against production? Where are the human oversight points, and what can the system do without one? What is your incident process when the model does something it should not? How do you swap or change models, what testing gates a switch, and do downstream customers get notified?

Notice these are largely the same questions that apply to red teaming: before deploying, know where the model runs, what inputs it processes, what the outputs look like, and how business risk gets revisited over time. That is not a coincidence. Good security questions and good compliance questions converge, because both are ultimately about whether you understand the system you operate.

What Are the Critical Flaws in AI Compliance Questionnaires Today?

The critical flaws are threefold. First, questionnaires rely on prose rather than artifacts, making it impossible to distinguish genuine compliance from well-written fiction. Second, they treat AI systems as deterministic, ignoring that LLMs require continuous monitoring rather than point-in-time attestations. Third, they fail to differentiate risk tiers, applying the same burden to a marketing chatbot as to a clinical decision support system, diluting attention and resources.

How Standards Bodies and Procurement Pressure Can Drive a Model Card Schema

There is one industry fix that would eliminate half of these questions before they are ever asked: a standardized model card. Today, every model provider publishes something different—different eval benchmarks, different safety disclosures, different levels of detail on training data, different definitions of the same terms. That inconsistency forces every downstream customer to run their own bespoke questionnaire, and forces every vendor to answer the same questions a hundred different ways.

Imagine instead a model card with a fixed schema: model version and lineage, training data provenance categories, where and how long data is retained, a common core of eval benchmarks with published scores, documented safety mitigations, default inference parameters, and a change log tied to every model swap or version update. SOC 2 did not succeed because it was clever. It succeeded because everyone agreed on what the report looks like, so one artifact could answer a thousand customers. The model card should be AI’s equivalent: produce it once, keep it current, and let it stand in for the first fifty questions of every assessment.

Standards bodies are circling this idea, and both ISO 42001’s documentation requirements and the EU AI Act’s transparency obligations gesture at it. But until the schema is common, buyers can force the issue by asking for the same fields, in the same format, every time. Procurement pressure standardized SOC 2. It can standardize the model card too.

How to Build Compliance That Outlasts the Next Generation of AI Standards

Compliance is an evidence problem, and for AI, evidence is an observability problem. The organizations that will sail through the enforcement wave now arriving are not the ones with the thickest binders. They are the ones whose logging, evals, and documentation were designed so that any reasonable question can be answered in minutes with an artifact rather than in weeks with an essay. This is what timeless compliance looks like. Frameworks will keep multiplying, regulators will keep diverging, and the models themselves will be unrecognizable in three years. But the principles underneath do not move: know your system, log what matters, measure continuously, scale scrutiny to risk, and never ask a question you cannot act on. Those held true for surgical checklists and pre-flight cards, they held true for SOC 2 and ISO 27001, and they will hold true for whatever comes after the current generation of AI standards. Comprehensive coverage is a moving target. Simplicity, done honestly, is permanent.

Share This Article