Key Takeaways
- A demo is built to hide exactly what matters at scale: multi-client behavior, real volume pricing, failure recovery, data handling, and exit cost.
- Ask what happens when the vendor's model or API changes underneath you, not just what happens when your own usage goes up. We run our own workflow fleet and have lived a 22-day silent outage — the question isn't rhetorical for us.
- A good vendor answer sounds specific and slightly uncomfortable. A red-flag answer sounds confident and generic.
- We evaluate vendors from both chairs — as a buyer choosing tools and as a company that sells assembled AI workflows. That bias is disclosed below, not hidden.
An AI vendor's demo is the best version of their product that exists. Clean input, one user, one context, thirty minutes, no deadline pressure — every condition that makes software look good and none of the conditions that make agency work hard. That's not a criticism of the vendor. It's what a demo is for. The problem is that agencies buy based on the demo and then operate the tool under completely different conditions: eight client voices instead of one, production volume instead of a sample query, a launch week with two client accounts live at once instead of a scheduled sales call with nothing else competing for attention.
We wrote about the build, buy, or assemble decision as the frame that comes before this one — which parts of your stack should be a subscription, which should run on an orchestration platform you control, and which are rarely worth building from scratch. This piece assumes you've landed on "buy" for a specific capability, or you're deciding between buying and assembling, and now you're sitting across from a vendor whose product looks good. It usually does. The checklist below is how to find out what the demo isn't showing you, structured as six dimensions, each with the question to ask, what a good answer sounds like, and what a red-flag answer sounds like.
One disclosure before the checklist, because it changes how you should read everything that follows: Kelva sits on both sides of this table. We buy vendor tools for capabilities that don't differentiate us, the same way any agency should. We also sell assembled AI workflows, which puts us in the vendor's chair in other conversations. We're not neutral in the abstract sense, and we're not pretending to be. What we are is specific about where a tool earns its keep and where it doesn't, because that specificity is what a demo can't fake and a sales deck won't volunteer.
How Does It Behave With Multiple Clients Running at Once?
Ask the vendor to run the tool live with three different client voice profiles active in the same session, not switched between separately staged demos. A tool architected from the start for multiple tenants shrugs the exercise off — profiles load, switch, and stay walled off from each other, with nothing from Client A's session bleeding into Client B's output. A tool that wasn't built for it will show it fast: a login screen between profiles, an awkward pause while someone resets state, or the confession that "profiles" is really just re-typing the same settings for whichever client is up next.
This matters because most AI tools are designed for a single user with a single context, and agencies are structurally the opposite: one team, many clients, many voices that must never bleed into each other. A vendor's roadmap slide about "multi-tenant support coming soon" is a tell, not a reassurance — it means the architecture wasn't built for this from the start, and retrofitting isolation into a single-context tool is a harder problem than most roadmaps admit.
Good answer: a live, unscripted switch between three active profiles, with a straight answer about how isolation is enforced at the data layer, not just the UI layer.
Red-flag answer: "most of our customers just use one workspace" or a demo that only ever shows one voice active, framed as if that's the whole product.
What Does the Bill Actually Look Like at Your Volume?
Ask for the calculation against your actual monthly usage pattern, in writing, before you sign — not the vendor's example numbers, and not a generic tier from the pricing page. Agency usage is bursty and concentrated: a handful of people touch the tool dozens of times a day during a launch week and barely at all the rest of the month, and total volume scales with client count more than with headcount. Per-seat and per-generation pricing assumes individual, steady usage, not a roster that spikes and slumps with client load — and the distance between what the pricing page quotes and what your real month bills out to is exactly where buy-path budgets quietly blow up.
The honest version of this question has two parts: what does a normal month cost, and what does a crunch month cost, given how your team's usage spikes. A vendor who can answer both from your numbers has priced this before with a customer shaped like you. A vendor who can only answer the first has priced this against their own average customer, which may look nothing like an agency running eight client accounts through a handful of power users.
Good answer: a written estimate built from your usage pattern, with the crunch-month number shown separately from the average-month number, and no hedging about whether seat count or generation count is the real driver.
Red-flag answer: "it depends" without a follow-up question about your actual pattern, or a quote that only ever cites the published tier price.
What Happens When It Breaks — and Who Finds Out First?
Ask the vendor to walk you through what happens when the tool fails mid-run: who notices, how fast, and what the client-facing consequence looks like while it's down. The answer tells you whether the team behind the tool has lived through running it under load, or simply shipped it and moved on. A team that has run the thing in production has a specific story — a monitoring layer, an alerting path, a rollback procedure, maybe even a past incident they can describe plainly. A team that hasn't will answer in the abstract: "we have monitoring in place," with nothing behind it when you ask what the monitoring actually watches.
We take this question seriously because we've paid for the answer ourselves. A scheduled workflow we depended on went silent for twenty-two days — no errors, no crash, nothing in the logs, just a trigger that stopped firing while everything around it still looked healthy. It stayed invisible that long because the watchdog and the workflow it was supposed to catch were running on the same infrastructure — when one went dark, so did the alarm. The fix wasn't a cleverer check bolted onto that same system; it was an external watchdog running on separate infrastructure, so two independent systems were watching each other instead of one system watching itself. That's not a vendor story — it's what happened running our own stack — but it's exactly the failure mode this question is designed to surface in someone else's tool before you're the one who discovers it twenty-two days late.
Ask specifically what happens when the model or API behind the vendor's product changes or goes down upstream — a provider deprecates a version, changes a rate limit, or has an outage of its own. A vendor with a real answer has a fallback path or at minimum a communication plan. A vendor without one will discover the gap the same day you do, in production, on a client account.
Good answer: a specific description of monitoring, alerting, and what a client sees during an incident, ideally with a past example the vendor is willing to name.
Red-flag answer: "we have monitoring in place" with no detail about what it watches, or visible surprise that you asked about upstream model failures at all.
Where Does Client Data Go, and Can You Get It Back Out?
Ask exactly what happens to client-confidential material once it enters the tool: where it's stored, who can access it, how long it's retained, and whether it's used to train anything beyond your own account. Then ask the harder, less-asked question — what does exporting your data back out look like, in full, in a usable format, on your schedule rather than theirs. Agencies routinely think through the first question because a client's data-handling terms or a framework like GDPR forces the conversation. Fewer think through the second until they're already trying to leave.
A vendor built for this gives you specifics without needing to check with legal first: retention periods, a data processing agreement, a clear answer on training use, and an export path that doesn't require a support ticket. A vendor who hasn't will point you to a privacy policy page and call it answered, which is not the same as answering it. We've written the fuller GDPR framing separately in GDPR and AI: what you can send — worth reading before this conversation, not after, so you know which categories of client material change your calculus before you're mid-negotiation with a vendor who hasn't thought about it either.
Good answer: specific retention terms, a clear no (or clearly scoped yes) on training use, and an export mechanism you can test before signing.
Red-flag answer: a link to a general privacy policy in place of a direct answer, or hesitation on whether your data trains their model.
Can It Model Your Approval Chain, Not Just a Single Pass?
Ask the vendor to show a workflow with a rejection built in — not the happy path where the first draft is approved, but the branch where a reviewer sends something back, a client's specific rule triggers ("this account needs legal sign-off before publish"), or a revision ceiling is hit and the item escalates to a person. Real agency work is full of these branches, and they differ client to client. A demo, almost by definition, shows the clean pass: prompt in, output out, approved on the first try. That's the version that fits in thirty minutes. It's also the version that tells you nothing about whether the tool can hold your actual process.
Most single-purpose tools were not built to model this, because modeling it requires exactly the kind of conditional logic and state-holding that a workflow platform handles and a generation tool typically doesn't. That's not automatically disqualifying — plenty of tasks genuinely are single-pass, and buying a tool for those is the right call. But if the capability you're evaluating sits inside an approval-heavy, client-specific process, and the vendor can't show you a rejection branch on request, you've learned something the demo alone wouldn't have told you.
Good answer: a live rejection branch, a client-specific rule applied mid-workflow, and a description of how escalation to a human actually routes.
Red-flag answer: "you'd handle that outside the tool" for a step that's core to how your agency works, offered as if it's a minor gap.
What Does It Cost You to Leave?
Ask what switching away from this tool involves a year from now — not because you're planning to leave, but because the honest answer tells you how much leverage the vendor thinks they'll have over you later. Get specific: can you export your configuration, your prompts, your client profiles, your history, in a format another tool could use? Or does everything of value live in a proprietary format that only this vendor's interface can read? Lock-in isn't always a red flag by itself — every tool creates some switching cost, including the ones you'd assemble yourself. What matters is whether the vendor is upfront about the size of that cost or lets you discover it later, at the worst possible time, mid-negotiation over a price increase.
This is also where "buy" and "assemble" carry genuinely different exit profiles, worth naming plainly since it cuts against our own commercial interest to leave unsaid. A subscription tool's lock-in is usually about data portability and configuration — you can generally leave, but you'll redo setup work. An assembled stack's lock-in, if you built it right, should be lower, because the process logic lives in workflows you own on a platform you chose, not inside a vendor's proprietary interface. It's a different exit profile again if the switching cost you're weighing is a full custom build rather than a vendor subscription — we've priced that path separately in the real cost of building an internal AI tool, and it's rarely the cheaper way out. Neither path is switching-cost-free. The point of this question is to make the cost visible before you commit budget, not to discover it in the renewal email.
Good answer: a direct description of what's exportable, in what format, and an acknowledgment of what genuinely doesn't travel.
Red-flag answer: "we haven't had many customers leave" offered as an answer to a data-portability question — that's a retention statistic, not an export mechanism.
The Checklist as a Table
Six dimensions, the question worth asking on each, and the shape of a good answer versus a red flag, side by side.
| Dimension | Ask | Good answer | Red flag |
|---|---|---|---|
| Multi-client behavior | Run three voice profiles live, not staged separately | Smooth live switch, isolation explained at the data layer | "Most customers use one workspace" |
| Pricing at your volume | Written calculation against your real usage, average and crunch month | Specific numbers built from your pattern | Only the published tier price, no follow-up questions |
| Failure and recovery | Walk through a mid-run failure, including upstream model or API issues | Named monitoring, alerting, and a real past example | "We have monitoring in place," no detail |
| Data handling and export | Retention, training use, and a testable export path | Specific terms, export you can try before signing | A link to the general privacy policy |
| Workflow depth | Show a rejection branch and a client-specific rule mid-workflow | Live branch, clear escalation routing | "You'd handle that outside the tool" |
| Lock-in and exit cost | What's exportable a year from now, and in what format | Direct list of what travels and what doesn't | "We haven't had many customers leave" |
Frequently Asked Questions
How long should a vendor pilot run?
Long enough to cross at least one real deadline crunch, not just a quiet week. A pilot that only ever runs during calm conditions never tests the thing that actually breaks tools — a launch week, a client escalation, two accounts needing output at once. If the pilot window can't stretch that far, weight the evaluation toward the six dimensions above more heavily than toward the pilot's smooth output.
Should we evaluate more than one vendor at once?
Where the budget and timeline allow, yes. Running two vendors through the same six questions side by side makes red flags easier to spot — a hedge that sounds normal in isolation reads very differently next to a competitor's direct answer to the same question. It also removes the pressure to accept a mediocre answer just because it's the only one you've heard.
What if the vendor won't answer a question?
That silence is the answer. A vendor who can't or won't get specific about multi-client behavior, real volume pricing, or failure recovery isn't holding back a trade secret — they're telling you the tool hasn't been tested against that condition. Treat a dodge with the same weight as a bad answer, not as a neutral non-event.
Ask This Before the Feature Checklist
Most vendor evaluations start with a feature checklist, and feature checklists are where demos win, because features are exactly what a demo is built to show. Start with these six dimensions instead, and the feature checklist becomes a much smaller conversation afterward — mostly confirming what you already suspected once you'd watched how the vendor handled the harder questions. A tool that survives multi-client behavior, real volume pricing, an honest failure story, clean data handling, genuine workflow depth, and a fair exit path has earned the feature conversation. One that can't clear those six has told you what you needed to know before you got there.
None of this is an argument against buying. Plenty of agency capability is correctly bought, and we say so plainly in the build, buy, or assemble piece this checklist supports — single-user tasks, standard output, one brand voice, buying wins on every axis that matters. If the capability in question is content generation specifically, our category-by-category review of AI writing tools for agencies maps where each type of off-the-shelf tool hits the same walls. The checklist above is for the moment right before you commit budget to any tool that has to survive contact with real agency volume, not just a sales call. If you've run it and the answers were good, buy with confidence. If the tool you're evaluating keeps hitting red flags on the dimensions that matter most to how your agency operates, talk to us about what assembling that capability around your own process would look like instead.
