How to judge a pilot proposal: scope, price, success criteria, and exit terms
The pilot proposal is where consulting firms reveal themselves. A firm that intends to earn the larger relationship writes a tight, measurable, exit-friendly pilot. A firm that intends to embed writes a vague one, because vagueness converts to change orders and change orders convert to permanence. You can read the firm's entire business model in six pages of proposal language if you know what to look for.
This deep-dive covers the four sections that decide whether a pilot proposal deserves a signature: scope, price, success criteria, and exit terms. It assumes you have already run the firm through the seven criteria on the main guide and the twelve questions on the interview page. The proposal is where their answers become contractual, or quietly fail to.
Scope: one problem, one system, one owner
A good credit union AI pilot runs four to eight weeks and attacks exactly one problem. Longer than eight weeks and you are not piloting, you are implementing without having earned the evidence. Shorter than four and nobody learns anything defensible.
The scope section should name three things precisely. First, the process being changed, with its current baseline: "indirect lending stipulation review, currently averaging 26 minutes per file across 340 files per week." Second, the system boundary: which data sources the pilot touches, which it explicitly does not, and whether anything member-facing changes (for a first pilot, the right answer is almost always no). Third, a single named owner on your side with real decision authority, because pilots without an internal owner drift into week twelve with nothing shipped.
Watch the scope language for load-bearing vagueness. "Explore opportunities across the lending lifecycle" is not a scope; it is a fishing license. So is any deliverable described as a "roadmap" inside a pilot that was supposed to produce working evidence. A roadmap is a fine deliverable for an assessment engagement. Inside a pilot, it usually means the firm hedged against its own build failing.
Price: sanity checks by engagement type
Pricing tells you whether the firm is right-sized for you, which is criterion six on the scorecard rubric. Reference points from the current market:
- Audit-style and assessment engagements (readiness reviews, use case prioritization, governance gap analysis) commonly run $10,000 - $50,000. These are diagnostic, fixed-scope, and short. Anything above that range for pure assessment work needs a persuasive explanation, usually institution size or unusual data complexity.
- Working pilots that ship something (a document automation workflow, a back-office copilot with your data, a fraud triage assist) typically land between $30,000 and $150,000 depending on integration depth. The number should map visibly to named people and weeks, not to "program management."
- Larger builds and platform work start around $100,000 and climb fast. If a firm proposes this as your first engagement, ask why the evidence step is being skipped. Global firms often cannot price below this floor; that is a fit problem, not a moral failing.
Two sanity checks catch most bad pricing. Divide the price by the stated team weeks and compare the implied weekly rate to what the staffing section actually promises: a $90,000 pilot delivering six person-weeks of a senior engineer plus oversight implies roughly $15,000 per person-week, which is defensible for specialist work and indefensible if the roster turns out to be one junior analyst. Then check fixed versus variable: a pilot should be fixed price. Firms that insist on time and materials for a six-week pilot are pricing their own scoping failure into your invoice.
One more test. Ask what happens to the pilot fee if you proceed to a larger engagement. Firms confident in their pipeline often credit some portion forward. Firms that refuse to discuss it at all are telling you the pilot is the product.
Success criteria: numbers someone can lose an argument over
The success criteria section is where most proposals collapse. "Demonstrate the value of AI for member service" cannot fail, which means it cannot succeed either. Insist on criteria with three properties: a number, a measurement method, and a named source of truth.
Strong examples from real credit union pilots: reduce stipulation review time from 26 minutes to under 10 on at least 80% of files, measured over the final two weeks against LOS timestamps. Achieve extraction accuracy above 95% on a 500 document holdout set your team selects, not the vendor's. Cut manual GL reconciliation exceptions by 40% month over month, per your finance team's existing exception report.
Note who controls measurement in each example: you do. A pilot where the consultant grades their own homework will pass. Also insist on one explicit failure condition in writing, such as "if accuracy on the holdout set is below 90%, the parties agree the approach did not validate." Firms with real confidence accept failure conditions readily. The ones who fight hardest against defining failure are the ones planning to declare victory regardless.
Baseline measurement belongs in week one of the pilot, not in a prior paid engagement. If a firm says it cannot commit to success metrics without a $40,000 discovery phase first, that is sometimes legitimate for genuinely messy data, but ask them to name the specific unknown that blocks commitment. "We need to see your data" is fair. "We never commit to outcomes" is a different business model wearing a pilot costume.
Exit terms: IP, data, and the right to walk away clean
Exit terms get the least attention during the honeymoon and cause the most damage later. Four items to verify before signing:
Work product ownership. Everything built with your money and your data (prompts, configurations, workflow definitions, evaluation sets, documentation) should be yours or perpetually licensed to you at no additional cost. Firms may retain their pre-existing tools and general methods. The line between those two categories should be written down, not assumed.
Data handling on exit. The proposal should state where your data lives during the pilot, whether any of it touches model training, and certify deletion or return within a stated window after the engagement ends, thirty days is typical. This is also third-party risk documentation your examiner may ask about, so get it in the contract, not in an email.
No hostage dependencies. If the pilot runs on the firm's proprietary platform, ask what happens at exit. A pilot that only proves value on infrastructure you must keep renting has mostly proven the rental. Portable builds on your tenant or on mainstream cloud services exit clean.
Continuation is optional and re-scoped. The pilot contract should end at the pilot. Auto-renewal clauses, pre-committed phase two language, or discounts contingent on signing the larger engagement before pilot results exist all convert your evidence-gathering exercise into a sales funnel you already paid to enter.
Red flags and green flags at a glance
Red flags: scope described in capabilities rather than outcomes; time and materials pricing on a short pilot; success criteria the firm measures itself; no failure condition; IP language that licenses your own work product back to you; a phase two signature required before pilot results exist.
Green flags: a single measurable problem with your baseline data quoted back to you correctly; fixed price mapped to named staff; a holdout evaluation you control; an explicit failure condition; clean IP and data deletion terms; a written option, not obligation, to continue.
Scoring it on the rubric
A pilot proposal gives you fresh evidence for at least three scorecard criteria: sequencing philosophy (did they pick a back-office problem with real numbers), right-sized delivery (does the price fit your budget with room to grow), and evidence (are the success criteria measurable by you). Score the proposal itself, not the meeting where they presented it. Charisma does not survive contact with a holdout set.
For a reference point on what a bounded, fixed-scope first engagement looks like when it is designed to be judged this way, see the structure of an AI readiness assessment, then apply every test in this article to it and to every competing proposal with equal force.