Procurement documents describe what's required. Quotes describe what's promised. The first project describes what the vendor actually does. Structuring the first project deliberately — rather than as a generic small job — turns 'let's see how it goes' into an evaluation you can defend to anyone.
Every buyer who has ever onboarded a new survey vendor has run a test project, whether they framed it that way or not. The first delivery tells you whether the quote was honest, the QA was real, and the operator can actually deliver what they said.
The difference between buyers who learn this well and buyers who don't is whether the test is deliberate — structured to expose the things that matter — or accidental, hoping that the small job reveals what a larger one would.
This article is the deliberate version. How to scope a first project that tests what needs testing, what to look for in the deliverable, and how to use the result to commit (or not commit) to programme work.
Three things only the first project reveals:
What's actually in the QA pack. Procurement specifies that the pack should include independent residuals; the quote acknowledges; the delivery either does or doesn't. You can't audit the contents until the contents arrive.
How the vendor handles change. A real project has midstream adjustments — site access shifted, coverage extended, deliverable format clarified. Vendors who handle change cleanly are different from vendors who handle it poorly, and the difference is invisible in a quote.
Whether the on-time promise is reliable. Quote says ten days from capture; first project lands in ten days or in fifteen. Both happen routinely; only the first project tells you which this vendor does.
For a one-off small project these three matter modestly. For programmes of recurring captures or large multi-stage projects, they're the entire point of vendor selection.
The test should be:
Small enough to fail cheaply. A $10k test project that reveals a vendor problem is a tolerable lesson. A $50k test project with the same lesson is more expensive than necessary. Pick a scope where the budget is acceptable as evaluation cost even if the delivery is rejected.
Real enough to exercise the workflow. A 0.5 hectare bare paddock doesn't test classification, doesn't test hydro-enforcement, doesn't test tile structure (only one tile). The test should hit at least three meaningful workflow stages.
Representative of your actual project shape. If your real work is corridor capture, test on a short corridor segment, not an area site. If your real work involves dense canopy, test on a vegetated site. The capability you need to validate is the capability the test should exercise.
Not so trivial it doesn't expose anything. A test where every default works is a test that didn't test anything. Worth deliberately including at least one non-default requirement.
Realistic test scopes: 10-30 hectare area capture or 2-5 km corridor with vegetation, engineering accuracy spec, full QA pack, deliverables in your target format. Project value typically $8k-25k.
What to check. The QA pack contains a base station RINEX log and post-processed trajectory documentation. The covariance plot through the flight shows consistent solution quality.
What to look for. Pack with base log + trajectory pass. Pack without either fails. Pack with RINEX but no covariance plot is intermediate — suggests PPK was run but the operator doesn't ship the documentation by default. Worth asking for.
What to check. The QA pack distinguishes between control marks used in calibration and marks withheld for validation. Reported residuals are at the withheld marks.
What to look for. Pack with explicit "used for calibration" vs "independent checkpoint" categorisation passes. Pack with single aggregated "control residuals" without the distinction fails — operator is reporting calibration consistency, not absolute accuracy.
What to check. Ten standard sections per the QA pack article — metadata, capture parameters, trajectory quality, strip alignment, checkpoint residuals, classification quality, density coverage map, surface QA, deliverable manifest, sign-off.
What to look for. All ten sections present with substantive content passes. Sections present but skeletal ("classification: done") fails. Sections missing entirely fails.
What to check. Delivered files open natively in your CAD, GIS or engineering software without conversion. Tile structure matches scoping. Manifest documents what's there.
What to look for. Native ingestion in target tools = pass. Conversion required for any specified format = partial fail (worth asking why). No manifest or undocumented naming = fail on documentation grade even if files work.
What to check. Throughout the project, how quickly does the vendor respond to scope clarifications, format change requests, schedule adjustments? Does the response include rationale or just compliance?
What to look for. 24-48 hour substantive response with rationale and trade-offs articulated = pass. Multi-day silence or non-substantive replies ("we can do that, no problem") = warning. Resistance to reasonable mid-project clarifications = fail.
What to check. Did the delivery land when the quote said it would, or later? If later, was the slip flagged proactively with reasons and revised timeline?
What to look for. On-time delivery = pass. Slip flagged proactively with explanation = partial pass (depending on context). Slip with no notice or delivered with no comment on the timing = fail.
Beyond the standard test items, include at least one deliberately non-default requirement in the brief — something that's a bit unusual and exposes how the vendor handles friction:
The stress test isn't about whether the vendor can do it — most can. It's about how they respond when asked: do they push back constructively (good), do it without comment (acceptable), or quietly skip it and hope you don't notice (fail)?
A vendor who proactively raises a concern about a stress-test item, explains the implication and proposes a path is more valuable than one who silently complies and leaves you to discover the issue at QA.
The single most diagnostic deliverable in an onboarding test is the QA pack. A few specific things to check, beyond completeness:
Honesty of self-assessment. Does the pack report worst-case residuals alongside RMSE? Does it call out areas of lower quality (sparse density, marginal classification)? Vendors who hide problems in the pack signal something about how they'll handle problems on bigger projects.
Sign-off practice. Is the sign-off by the same person who ran the processing, or by an independent reviewer? Independent sign-off is the convention and the indicator of mature internal QA.
Documentation depth. Are processing parameter values documented? Software version numbers recorded? The deeper the documentation, the more reprocessable the archive (see re-fly vs reprocess article) and the more transparent the workflow.
Voice of the pack. Does the pack read like a defensible engineering document or a marketing brochure? "Accuracy achieved" with residuals attached vs "exceeded all requirements" with no numbers.
Beyond the deliverable, the invoice carries information:
Cost breakdown clarity. Does the invoice itemise the quoted work, or arrive as a single lump sum matching the quote? Both are acceptable for fixed-price work, but the itemisation makes later renegotiation easier.
Itemisation of any extras. If anything was added scope, is it clearly attributed and justified? Or buried in a generic adjustment line?
Match against quoted scope. Does the invoice match the quote, or include "expected" extras that weren't communicated during the project?
Response when asked about line items. A vendor who can explain every line clearly = pass. Defensive or hand-waving responses = warning.
Three patterns we see when buyers run first projects without explicit evaluation framing:
Test project too small to reveal anything. A 2 hectare bare site captured at planning accuracy doesn't test the workflow. Pick a test that exercises real capability.
Skipping the QA pack evaluation. Deliverable looks fine in the design tool, project proceeds. The QA pack is the actual scorecard; the deliverable is the consequence. Evaluate the pack regardless of how the cloud looks.
Treating it as "just a small job". The first project sets the working pattern for everything that follows. Treating it casually generates a casual working pattern; treating it as an evaluation generates a more rigorous one.
Not having an exit criterion before starting. Decide in advance what would cause you to not proceed with programme work. Evaluate against that criterion when the test completes. Without the criterion, every result becomes "good enough" by inertia.
A test project that lands well produces:
A test project that lands poorly produces:
Both outcomes are useful. The wasted outcome is the test that neither passes nor fails because the buyer didn't have criteria to evaluate against.
The first project with any new vendor is a test whether you frame it as one or not. Frame it deliberately. Pick a scope small enough to fail cheaply but real enough to exercise the workflow. Test six things explicitly: PPK + base log delivery, independent checkpoint methodology, QA pack completeness, deliverable format alignment, communication response, on-time delivery.
Include at least one deliberately non-default requirement as a stress test. Evaluate the QA pack rigorously — honesty, sign-off, documentation depth, voice. Evaluate the invoice too.
Decide your exit criterion before starting and apply it when the test completes. A test without a decision framework isn't a test.
If we're being considered, the briefing above is the test we'd recommend you run on us. Send through your evaluation criteria and the project shape; we'll structure the first project to expose whether we're the right fit before either side commits to programme work.