Notes from the field.
Inspiration and learned experience while building, shared to be used.
Entries
2 pieces- July 31, 2026Evaluation
Evaluating LLM agents: how would you know it had stopped working?
What a grader is, what a harness is, what a protocol decides, and why the difference matters the first time a system is allowed to block a release. Built from two products, with the gaps named.
- April 20, 2026Agentic AI
A wrong forecast is absorbed. A wrong action reaches a guest.
Prophet is a statistical forecasting library. Claude is a large language model. Keeping them as two components, rather than letting one model both predict and act, is the decision that shaped everything after it.