Field Notes

Reflection
Reflection

Codex Astra did not remove the need for proof.

A more capable coding agent can cross a larger problem in one run. That makes scope, receipts, and independent verification more important, not less.

2026-09-07 · 5 min
Pattern
Pattern

The Second Brain needed a job description.

A beautiful knowledge graph becomes another inbox unless it is clear what belongs there, what stays elsewhere, and which system is allowed to call work complete.

2026-09-07 · 5 min
Postmortem
Postmortem

Polygon count could not save the course.

The environment had terrain, trees, bunkers, and mountains. It still did not feel like the concept, because technical completeness and visual authorship are different jobs.

2026-09-07 · 5 min
Pattern
Pattern

The game started as a pipeline, not a level.

Before there was a playable hole, there had to be a way to keep concepts, assets, physics, builds, and claims from collapsing into one folder called progress.

2026-09-07 · 5 min
Teardown
Teardown

A disabled button was not a broken recorder.

The Mac app was telling the truth about paired capture while hiding the working single-camera path below the preview.

2026-08-29 · 6 min
Pattern
Pattern

The copy had to become permission to delete.

Moving a golf video off the iPhone was easy. Proving that the phone copy could be safely removed was the real system.

2026-08-29 · 7 min
Teardown
Teardown

Seventy-six UI tests passed. Two range controls were still outside the window.

Swing Trainer's Live Studio passed its presentation suite, but a real 1149×772 window clipped critical capture controls. Why visual verification remains a separate release gate.

2026-08-25 · 5 min
Postmortem
Postmortem

The Mac app was healthy, signed, and running. The workstation still did not appear.

A signed macOS process, green tests, and an idle run loop did not prove that the user had a usable app. The launch investigation that separated disk work, window state, and observation error.

2026-08-25 · 7 min
Postmortem
Postmortem

I marked a swing, closed the app, and discovered the annotation was never a record.

Swing Trainer needed to preserve partial human review without calling incomplete work training truth. The revision model that keeps drafts, labels, and source video separate.

2026-08-25 · 8 min
Pattern
Pattern

I taught the app when to record a swing. I refused to let it call that motion Takeaway.

Body motion can start and stop a golf recording, but it cannot prove where the club is. The evidence boundary behind Swing Trainer's hands-free capture system.

2026-08-25 · 7 min
Pattern
Pattern

A phone can trace the club. That does not mean it measured clubhead speed.

The physics boundary inside Swing Trainer: what monocular video can support, what calibration might unlock, and what still requires an external instrument.

2026-08-24 · 8 min
Pattern
Pattern

I had enough data to spend money on a GPU. I did not have enough evidence to train.

Why the Swing Trainer RunPod gate stops at provenance, frozen splits, and real five-point labels—even with thousands of images already staged.

2026-08-24 · 7 min
Teardown
Teardown

We found 8,295 golf images. Almost none were the data we needed.

GitHub, Roboflow, Kaggle, and Hugging Face produced plenty of files. The hard part was refusing to call mismatched labels a dataset.

2026-08-24 · 8 min
Reflection
Reflection

I thought tracking a golf club meant finding a line. Then the line disappeared.

The club looked simple until motion blur, hands, clothing, and the edge of the frame turned one object into five different evidence problems.

2026-08-24 · 7 min
Playbook Update
Playbook Update

Automation needs a job description.

The useful question is not “can Airtable automate this?” It is “what may this automation change without a person reviewing it?”

2026-08-16 · 3 min
Pattern
Pattern

Do not put the club in the body frame.

A pose system can see a body, a club detector can see a club, and a launch monitor can report a ball. Calling all three one measurement is how a golf product starts inventing facts.

2026-08-16 · 5 min
Reflection
Reflection

The performance number we need is the one we do not have.

The Range spike has a full instrument, a protocol, and a hard exit bar. It still has no iPhone measurement. That is not a gap in the story. It is the story.

2026-08-16 · 5 min
Pattern
Pattern

You can falsify a simulator without a single real measurement.

We had no measured data and a simulator we needed to trust. A model-vs-model audit produced decision-grade findings anyway. The five moves, in order, and what this method can never tell you.

2026-08-04 · 7 min
Pattern
Pattern

One prompt clause turned my analysis agent into an experimentalist.

Most analysis sessions return confident prose. This one returned reproduction scripts, error bars, and a ranked list of experiments. The difference was a one-page brief written like a lab contract.

2026-07-21 · 6 min
Teardown
Teardown

I asked Claude to audit my simulator. It killed every high score.

We had a tens-of-millions-of-throws simulator and an overnight Monte-Carlo run nearly three times larger. One Saturday-night physics audit showed the two models agree on every trajectory and disagree on half the outcomes — and every max-score throw was an artifact.

2026-07-09 · 8 min
Postmortem
Postmortem

I called three sign-recognition models failures. The recipe was the failure.

A Parley notebook reported three landmark architectures as broken on cross-signer ASL. A warmup and a gradient clip brought all three back, and two matched the best model. The ranking had measured my training recipe, not the models.

2026-06-08 · 7 min
Teardown
Teardown

Recipe beats architecture: lottery tickets in sign models

We trained seven landmark architectures three times each. Three of them worked on one seed and collapsed to near-random on the others. A single-seed comparison would have called two of those collapses a result.

2026-05-31 · 6 min
Pattern
Pattern

The 38-point gap: one accuracy number, twenty-one very different users

Our sign model averages 42% across signers. That average hides a range from 26% to 64% — and the thing that decides where a person lands is not the signs they make, it is who they are.

2026-05-31 · 7 min
Reflection
Reflection

45%, not 90%: the only sign-recognition number I trust

Our best landmark-only sign model scores 45% on signers it has never seen. The field routinely reports numbers twice that high. The lower number is the honest one, and it is the one we publish.

2026-05-31 · 6 min
Pattern
Pattern

Six ways hearing-built sign-language AI fails the Deaf community

I keep a running catalog of how hearing-led sign-language AI fails. It is not a list of other people's sins. It exists so Parley can catch itself the moment it starts to look like one of them.

2026-05-31 · 7 min
Reflection
Reflection

Picking the glasses

One AR-glasses decision for two ventures. The criteria that survived the cut, and why doing it once was the right move.

2026-05-25 · 7 min
Reflection
Reflection

Why I'm running Parley

I started a Kaggle research project in Q2 2026 while running two startups. The decompression channel, the four open questions, and what makes it survive.

2026-05-25 · 8 min
Postmortem
Postmortem

ESP32 firmware with Claude: the gap between 'it compiles' and 'it works on the bench'

Claude is excellent at writing ESP32 firmware that compiles. It is not reliable at predicting what that firmware will do when the hardware is actually in front of you. Three incidents and the gate I added.

2026-05-17 · 7 min
Pattern
Pattern

31 million systematic throws or 91 million random ones. Which one taught us more.

The answer isn't obvious. Systematic grid search gives you coverage. Random sampling gives you reality. You need both, and you need to know what each one is telling you.

2026-05-17 · 6 min
Teardown
Teardown

We ran 124 million simulated cornhole throws. Here's what it cost and what we got.

A parametric physics engine, 8 parameters, two overnight runs. The database exists. Here's the honest accounting of what building it took and what we learned that we couldn't have learned any other way.

2026-05-17 · 8 min
Teardown
Teardown

What it actually takes to build an AR overlay on a physical object in real time.

AR on physical objects is 80% coordinate system problems. Claude is great at helping you think through the geometry. You still have to understand it yourself.

2026-05-17 · 8 min
Pattern
Pattern

Training a custom CV model with Claude: the data quality lesson we learned the hard way.

Clean data beats model size. Every time. Don't upgrade the model until you've audited the labels.

2026-05-17 · 7 min
Teardown
Teardown

We put a language model inside a hardware device. Here's every decision we made.

LLMs in real-time hardware aren't ChatGPT. Latency budget is the constraint that changes everything.

2026-05-17 · 8 min
Postmortem
Postmortem

Claude built our CV pipeline. Then it lied about being done.

Agents are unreliable judges of their own work. Here's how a structural fix — not a smarter model — stopped QC's CV pipeline from shipping silent failures.

2026-05-17 · 8 min
Postmortem
Postmortem

I pasted my session tokens into a chat. Here's the gate I built.

Sixteen cookies that together are my whole Google account, dropped into a chat window — while building a security playbook. The how is the whole point.

2026-05-14 · 5 min
Pattern
Pattern

Computer vision already runs elite sports. It's about to run the rec league too.

Twelve-camera tracking rigs and Hawk-Eye are infrastructure at the top. The same capability now fits on a $249 board and a commodity camera. What that unlocks across every sport — and why the smartest way in is the narrowest one.

2026-05-14 · 7 min
Playbook Update
Playbook Update

I run three ventures from one Obsidian vault. Here's the 13-folder template.

Why a markdown vault outperforms a Notion plus Asana plus Drive plus Slack stack when AI agents are part of the work. Battle-tested across QC, MHG, and Parley.

2026-07-07 · 8 min
Playbook Update
Playbook Update

Why I write a postmortem for every meaningful incident, and the template that makes it take 45 minutes.

Twenty-six postmortems across QC and Parley in six weeks. The template, four worked examples, and the discipline that makes the same mistake stop happening twice.

2026-06-23 · 7 min
Playbook Update
Playbook Update

I shipped a Sprint Contract template. Here's why my AI agents kept declaring done when they weren't.

A 48-contract system born from agents that praised their own work. The fix wasn't a smarter model. It was structural.

2026-06-09 · 7 min
Reflection
Reflection

I'm starting a publication. Here's what it is and why.

Field notes from an operator running three ventures on Claude Code. Biweekly. No theory. Receipts only.

2026-05-26 · 5 min
Pattern
Pattern

How I run six AI agents across three ventures without them stepping on each other.

The routing model that holds up under load. Built for TruPath, tested on Mile High Golf, Quantum Caddy, and Parley.

2026-05-12 · 6 min
Postmortem
Postmortem

I tried to build my own AI-native business OS. Here's why I scrapped it.

A 20-agent Electron app that didn't ship. The decision to throw it away. What I'd tell anyone tempted to build the same thing.

2026-04-28 · 7 min