Training
What a Two-Group Test Can Prove in the Field
Comparing two setups on a training field is easy to do badly and possible to do well. The difference is a protocol that lets only one thing change at a time.
A two-group test on a training field can prove one thing: that under these conditions, with this dog, on this ground, the changed variable moved the result in a direction. It cannot prove the variable will do the same for a different dog, on different ground, in different weather. That is not a weakness. It is the honest shape of what field evidence is, and a test run within those limits is worth more than a season of impressions.
The pattern is the one every measuring discipline uses: hold everything still except the thing being compared. In archery, a register of equipment notes ran exactly that experiment on arrows, and published it as a bare and wrapped flight test: same bow, same shafts, the only difference the film of vinyl wrapped around the shaft. The value of the result is not that wraps always cost the same; it is that on that day, for those arrows, the difference was measured rather than argued about.
How do you keep a field test fair when only one variable changes?
By deciding everything else in advance and writing it down. The dog, the ground, the thrower or launcher, the distance, the order of the runs and the scoring rule all get fixed before the first rep. The variable being compared is the only thing allowed to move: two collars, two whistle timings, two line positions at the same station, two drying routines between identical swims.
Order is the variable nobody intends to introduce. A dog is fresher on the first runs than the last, so the two setups should alternate rather than run in blocks, and the count of runs should be even. Fatigue, learning and daylight all move through a session, and a test that does not account for them measures them instead of the variable.
What can a score sheet say, and what can it not?
A score sheet can say what happened: times to the fall, distances off line, whistles used, deliveries clean or not. It can say the direction of a difference and roughly its size. What it cannot say is why. A dog that ran two seconds faster under one setup might have been faster because of the setup, because the wind changed, or because it had stopped worrying about the gun. The sheet records the result, and the reader has to resist promoting it into a cause.
The useful habit is to write the conclusion as a sentence with conditions attached: on this ground, with this dog, the difference ran this way. A conclusion without its conditions is where field testing turns into folklore. The table that measures what a level actually requires, so that the thing being tested is worth testing, sits in the stakes and levels table.
How many runs does a result need?
Enough that the day stops being the variable, which in practice means more than a handler wants to run. A difference that shows in four runs is a suggestion; the same difference holding across twelve runs, on two days, is a result. There is no magic count, and pretending otherwise is the way field tests get overstated. What there is instead is a discipline: run enough that the order effects and the tired runs have had their chance to flatten the difference, and if the difference survives them, it is probably real.
The count is set before the session for the same reason the measure is: a test that ends when the numbers look right is a test that was waiting for the numbers to look right.
| Decision | Fair version | What it prevents |
|---|---|---|
| Runs | Even count, setups alternating | Fatigue and freshness becoming the variable |
| Ground | Same field, same wind window, same day | A softer fall area deciding the result |
| Scoring | One written measure, fixed before the first rep | A measure chosen after the fact to fit the outcome |
| Conclusion | Stated with dog, ground and date attached | A local result hardened into a rule |
Where should a test protocol live so another club can repeat it?
In writing, in public, with the numbers left blank enough for another ground to fill them. A protocol that lives in one person's memory dies with the season. The disciplines that keep their measurements comparable, from archery registers to laboratory handbooks such as the NIST engineering statistics handbook, all do the same thing: the method is written as a method, and the result is written as a result, and the two are not confused.
A field test worth keeping
- One variable, named in a sentence before the session starts.
- An even number of runs, alternating, on the same ground on the same day.
- A single scoring measure written down before the first rep.
- A conclusion that carries its conditions rather than shedding them.
- The protocol itself saved where someone else could run it again.
Common mistakes
- Changing two things at once and crediting one of them.
- Running all of group A, then all of group B, and calling the difference an effect.
- Choosing the measure after the runs, when the numbers are already known.
- Generalizing from one dog on one day to a recommendation for every dog.
- Keeping the protocol in memory, where it quietly changes between sessions.