AlpacaX

Insights

AI vendor questionnaire checklist: 6 runtime questions to score

The six runtime questions to run this week, scored so a failure can't average into an approve.

Marco Kwak
Marco KwakHead of GTM · September 11, 2026

The six runtime questions to run this week, scored so a failure can't average into an approve.

The 2026 vendor questionnaire templates have started asking some of these questions. Reco AI's own template has an "autonomy and guardrails" section with a kill-switch question and a per-action approval question, and it scores every answer on a 0-to-2 rubric that totals to a decision band. But the only red flag that forces an automatic decline sits in the data-handling and model-training domains—a red flag on the autonomy questions just costs points, and the pass answer it scores for is whether approval gates are configurable per action type, not whether the gate pauses anything at execution time. A vendor can score well on "configurable approval gates" and still ship an approval flow that never actually stops mid-execution. This is the part you can run this week: the six runtime questions, scored so a failure at execution time can't average into an approve.

Each question below gets a pass bar—what a real answer looks like—and the red flags that should stop the evaluation until you get a real answer. A vendor that clears all six on paperwork alone hasn't cleared anything; make them answer in specifics.

The six runtime questions

1. Which commands or actions can be blocked at runtime, and how? (NIST AI RMF, EU AI Act risk controls) Pass: names a real mechanism—a rules engine, a policy gate, a human-approval step—and can describe a command it actually stopped. Red flag: "we log everything" as the entire answer, or "we haven't needed to block anything yet."

2. When is human approval required, and how does the human intervene? (EU AI Act Article 14 human-oversight obligation for high-risk systems) Pass: a defined trigger (risk score, action type, target system) routes to a named approver role before the action runs. Red flag: "approval" turns out to mean a policy document gets reviewed quarterly, not that anything pauses at execution time.

3. Can you terminate or intervene on a running agent, and how fast? Pass: the vendor commits to a specific bound in writing and can describe the mechanism that meets it—EU AI Act Article 14's stop obligation for high-risk systems requires the capability to intervene, but names no time window, so the number is the vendor's own commitment, not a standard to check them against. Red flag: "we can shut it down eventually," or a mechanism with no time attached.

4. What did the agent do, why, and under whose identity? (EU AI Act Article 12 record-keeping, SOC 2, ISO 42001) Pass: every action ties back to an authenticated human, and the vendor can produce a sample record. Red flag: attribution stops at an API key or a shared integration user—that's not an accountable identity.

5. Are the agent's permissions standing, or scoped and re-checked per action? (least privilege; NIST AI RMF govern/manage) Pass: the permission is issued for the task, and elevation is re-checked at each privileged action rather than resting on a grant that was issued once. Red flag: "the credential doesn't expire, but we rotate it periodically"—rotation is not scoping.

6. What's the inbound attack surface of the tool itself? Pass: the vendor can produce a specific inventory of their own network exposure—every port that's open, every protocol that's listening, and what's outbound-only versus inbound—not a general assurance. Red flag: "that's covered under our SOC 2," which says nothing about the agent's own footprint.

Three more red flags worth screening for

These aren't runtime questions, but they belong on the same sheet. Palavir's 2026 vendor red-flag list names the first; Reco AI's questionnaire carries the other two, as data and access questions.

  • Documentation deflection. "We can get that to you later" or "we're working on it" in response to a SOC 2, DPA, or BAA request is the answer, not a delay—good vendors have these ready before the first sales call; if it isn't ready, dig into what else isn't.
  • Training on your data by default. If the vendor trains on customer data unless you opt out, and the opt-out is buried in account settings, that's a red flag independent of anything above. Ask for it in writing, not as a checkbox you have to go find.
  • Long-lived, manually rotated API keys. Static, hand-rotated keys are the same red flag pattern applied to the vendor's own infrastructure instead of the agent's—it's a smaller version of question 5's standing-privilege problem.

How to use this

Don't average the answers. Reco AI's own rubric already scores every question 0 to 2 and totals to a decision band—but the only automatic decline it forces is a red-flag answer in the data-handling or model-training domains. Extend that same veto logic here: a vendor can pass four of six and still be a bad bet if the two failures are questions 1 and 4—block and attribution are the ones that determine whether you can act when something goes wrong, not just whether you have a nice paper trail after. The fuller argument behind these six—why runtime is the layer most questionnaires skip—is in AI vendor risk assessment: the runtime questions your questionnaire is missing.

Marco Kwak
About the authorMarco KwakHead of GTM

Marco Kwak is Head of GTM at AlpacaX, where he leads enterprise go-to-market and partnerships for Alpacon, an AI-native PAM platform with runtime execution control for AI agents. He previously held senior roles at H2O.ai and VMware, spanning AI cloud presales, global enterprise partnerships, and infrastructure software. He brings together engineering depth and commercial experience to help emerging infrastructure technologies move from technical validation to global adoption.


AI vendor questionnaire checklist: 6 runtime questions to score | AlpacaX