I thought tracking a golf club meant finding a line. Then the line disappeared.
The club looked simple until motion blur, hands, clothing, and the edge of the frame turned one object into five different evidence problems.

The first version of the Swing Trainer problem sounded almost embarrassingly simple: put a phone behind a golfer, find the club, draw its path.
Then I started looking at actual swing frames.
At address, the club is a clean dark line. Halfway down, the shaft becomes translucent. Near impact, the head is a smear. The hands cover the grip. A black shaft crosses black pants. The club exits the frame. Sometimes the thing that looks most like a shaft is a seam in the hitting net.
The object had not changed. The evidence had.
One line was hiding five questions
Calling the club a line collapses several different questions. Where does the grip end? Where does the shaft meet the head? Which side of the head is the heel? Which side is the toe? Is a missing point truly absent, merely outside the frame, or visible only in the model's imagination?
We replaced the line with a five-point contract: grip, shaft-grip junction, hosel, inner clubhead, outer clubhead. The names matter because each point carries different information. Two shaft points give orientation. The hosel anchors the transition from shaft to head. Two head boundaries expose more than a center dot can, while still stopping short of claiming that a phone image has recovered face angle in three dimensions.
That last boundary is important. Five image points are richer evidence. They are not magic.
Missing became a first-class result
The biggest change was not adding points. It was adding honest ways for points to be missing.
Every landmark can be observed, occluded, outside the frame, not present, or not annotatable. If the head is a blur, the annotator does not click the middle of the blur and call it truth. If the grip is behind the golfer's hands, the model does not receive a guessed coordinate. If the club leaves the frame, zero is not a location.
This sounds like annotation housekeeping. It is actually product architecture. A coaching system that cannot say “I did not see it” will eventually turn low-quality pixels into high-confidence advice.
The tracer is now an evidence stream
The body pose and club pose live as sibling evidence streams. That distinction keeps us from quietly turning a body skeleton into a club measurement or filling a club gap with a wrist trajectory. The streams can be compared later, but neither gets to impersonate the other.
The same rule governs rendering. Observed points can appear. Missing points create visible gaps. Predicted or filtered motion can be shown as prediction only, never smuggled into the golden labels used to judge the model.
What the phone can honestly show
A single camera can show image-plane position, timing, repeatability, and view-specific patterns. Those are useful. It cannot automatically give calibrated three-dimensional club path, true face angle, or physical clubhead speed. Perspective compresses depth. Frame timing and shutter blur change what can be located. A beautiful path overlay can still be physically wrong.
So the build order changed. First make the evidence legible. Then measure how often it is missing. Then validate errors on real footage. Only after that do we earn stronger claims.
I started this project trying to draw a line. I ended up writing a contract for when the line is allowed to exist. That contract is less impressive in a demo and far more valuable in a product.