There is a class of product failure that lab testing never catches. The unit passes drop test, temperature cycling, and every functional check on the bench — then fails in week three at a customer site because somebody mounted it upside down, or the Wi-Fi there uses a captive portal, or the enclosure gets pressure-washed every morning and nobody told you. A field pilot exists to surface exactly that while it is still cheap to fix.

The mistake is treating the pilot as a marketing event or a soft launch. It is neither. It is a structured experiment with a hypothesis, a sample, instrumentation, and a decision at the end. Run it casually and you end up with twenty units in the field, a pile of anecdotes, and no basis for a go decision.

What a field pilot is actually for

A pilot answers questions that only reality can answer:

  • Does it survive the real environment? Dust, vibration, humidity, temperature swings, power quality, cleaning chemicals, and the specific abuse of the actual job.
  • Do real users operate it correctly? Without you in the room, without a demo script, after the novelty wears off.
  • Does it deliver the value you claim? Measured, not asserted. If your pitch is "cuts inspection time by half," the pilot has to measure inspection time.
  • What breaks, and how often? Early failure rate, failure mode distribution, and whether failures are random or systematic.
  • What does support actually cost? Calls per unit per month is a number that decides your business model.

Note what is not on that list. A pilot is not the place to test whether people want the product — that question belongs much earlier, in validating the idea before you spend on development. By pilot stage you have already committed real money.

How many units, and for how long

There is no universal number, but there are useful anchors. For a consumer product, ten to thirty units across a deliberately varied set of users gives you enough coverage to see the common failure modes. For a B2B or industrial product, three to eight sites is often more informative than thirty consumer units, because each site runs the product hard and reports precisely.

Duration is driven by the failure mechanism you are worried about. A wear failure at a few thousand cycles needs a pilot long enough to accumulate them, or an accelerated bench equivalent running in parallel. Four to twelve weeks is the usual working range; under two weeks you catch only infant mortality, and over four months you are delaying a decision you could already make.

Product typeUnits in fieldDurationPrimary risk being tested
Consumer electronic device15–304–8 weeksMisuse, connectivity, battery life
Industrial or B2B equipment3–8 sites8–12 weeksDuty cycle, environment, integration
Wearable or body-worn20–406–12 weeksFit range, skin contact, sweat and washing
Outdoor or field hardware5–15Full seasonWeather, UV, ingress, temperature range

Recruit for variance, not for enthusiasm

The instinct is to hand units to friendly early adopters who will be gentle and forgiving. That is the opposite of useful. You want the extremes: the largest and smallest user, the least technical one, the site with the worst network, the operator who does not read manuals, the climate that is hottest and the one that is wettest. Enthusiastic users under-report problems because they do not want to disappoint you.

Be explicit with participants that the unit is pre-production, that you expect failures, and that reporting a failure is the entire favor you are asking. Some teams pay a nominal amount for participation specifically so the relationship is transactional rather than social — it measurably improves honest reporting. The mechanics of recruiting and managing participants overlap heavily with running a hardware beta program.

Instrument the units before they leave

The single biggest difference between a pilot that teaches you something and one that does not is instrumentation. Users describe symptoms badly and remember timing wrong. Data does not.

If the product has a microcontroller, log it: power cycles, resets and their cause, watchdog events, error codes with timestamps, battery voltage under load, temperature, and every state transition. If it connects, upload that log; if it does not, store it and retrieve it at the end. Building in remote diagnostics and logging before the pilot rather than after is one of the highest-return decisions in the whole program.

For purely mechanical products, instrument differently: a cycle counter, witness marks on wear surfaces, a torque check at intervals, and photographs of every unit at the start so you can compare wear at the end.

Define exit criteria before you start

Write down, in advance, what result means go and what result means stop. Otherwise you will rationalize whatever happens. Reasonable criteria look like this:

  • No safety-related failure of any kind. One is disqualifying.
  • Field failure rate below a stated threshold — often 5% over the pilot window for a first product.
  • No systematic failure mode repeating across more than one unit without an understood root cause and a verified fix.
  • Core value metric met at a defined percentage of sites.
  • Support contacts per unit under a level your margin can absorb.
  • All units returned and physically inspected, not just reported on.

That last one matters more than it sounds. Teardown of returned pilot units reveals loosening fasteners, corroded contacts, cracked bosses, and chafed wires that no user ever noticed and no log ever recorded.

Handle the legal side properly

Pre-production units in the hands of the public carry real exposure. Use a written participation agreement covering the pre-production status, permitted use, data collection and privacy, confidentiality where relevant, and return obligations. Check whether your product needs certification before it can lawfully be operated at all — an unlicensed radio, for instance, has limits on what may be marketed or operated before authorization. Confirm your insurance covers pre-launch units, and read up on product liability and insurance for a new product before the first unit ships.

Run change control from day one

Pilots generate fixes, and fixes generate confusion. Track which unit has which hardware revision and which firmware build, because otherwise a failure report becomes uninterpretable. Resist the urge to quietly patch units in the field without recording it. When a change is real, push it through a proper engineering change order so the drawing, the BOM, and the factory instructions all move together.

What to do with the results

Sort every finding into three buckets: fix before launch, fix in a running change, and accept and document. Founders tend to put too much in the first bucket and stall, or too much in the third and ship a known defect. The discriminator is severity multiplied by frequency, with anything safety-related automatically in bucket one.

Then decide honestly. A pilot that produces a stop decision has done its job and saved you the cost of a full production run plus the reputational damage of a bad launch. A pilot that produces a go decision gives you something more valuable than confidence: a documented basis for the tooling spend and a realistic warranty reserve.

A field pilot naturally sits between prototype field trials and the first real manufacturing build, so it pairs closely with field testing a prototype beforehand and a pilot production run after. For B2B products, the same units often double as the proof that closes your first commercial customer — see closing your first pilot with a business customer.

If you are approaching a launch and want a pilot designed properly — sample plan, instrumentation, exit criteria, and a real teardown of the returned units — Projects House can build and run it with you. Get in touch through our contact form.