Lab notes.
One mechanism a day, written the day it bit us. Shorter and rougher than the essays — published while the bruise is fresh.
These are the daily tier of our research: small, honest write-ups of what running a company on an AI operating system actually breaks and teaches. Each one is a single generalizable mechanism, not a diary entry.
08.02The error message is a witness, not a judge — A 404 told us the database wasn't shared. The lookup had failed — but the explanation was wrong, and believing it would have silenced an ingest forever.
07.30Funny is a skeleton, not a garnish — Sprinkling jokes on a serious script fails. Extracting the structure of material we actually love, and pouring the facts into it, worked on the first read.
07.29The form decided before we did — An internal priority decision sat unratified for a month — then an outside form with one mandatory field made it for us, on the record.
07.28Grep your own slides — A pitch deck is a stack of claims wearing nice typography. Our brain grepped the codebase for every noun on one slide — and an opinion became a finding.
07.27The premise of a question is a claim — “Should I open a PR?” reads as humility. It smuggles in an assertion — one doesn't exist — and nobody audits the premise of a question.
07.26A name that needs a footnote is a bug — “The room, the person, you” read cleanly on the page — then fell apart the first time it was said aloud. Name layers by the question they answer.
07.25A landing page is a list of claims — We shipped an argument page four times in a day — falsified, recency-audited, citation-verified. Persuasive register doesn't exempt a page from epistemic discipline.
07.24When the evidence itself is cached — We chased a bug that didn't exist and nearly shipped a redesign that didn't exist either — both times judging evidence that wasn't the artifact we shipped.
07.23Writing it down is not deciding — Every fact in the record was true, and the record was still wrong: a well-argued candidate looks identical to a decision on the page.
07.22You can't accelerate a track record — Capability responds to effort. A verifiable record accrues at exactly one day per day — and only counts from the day you publish it.
07.21A second conversation is not a second opinion — A returning enthusiast doubles your notes and moves your evidence count by zero. Count distinct people, not conversations.
07.20Not an information supermarket — The most useful thing we did to the product this week was delete four features nobody had built yet. Breadth needs a prosecutor, not just a sponsor.
07.19Winning on average, benched for variance — A model beat the baseline on pooled metrics and still got zero production weight. A promotion gate reads the distribution, not the mean.
07.18Cite the artifact, not the narrator — An agent pasted a link to a pull request it had computed, not observed. Before you state that something exists, read it.
07.17Cheaper isn't isolated — Tiering bulk work to a cheaper model cut our burn and changed nothing about reliability: the cheaper model drinks from the same pool.
07.16Done isn't deployed — A background agent that was never installed produces no errors — it produces nothing, which reads identically to “nothing went wrong.”
07.15Don't let the medic share a bloodstream with the patient — Our self-healing agent needed model capacity to think — and was dispatched exactly when the run had exhausted it. The doctor died of the disease.
07.14Don't ask the agent to grade itself — run a competition — You don't make an agent trustworthy by making it more self-aware. You make it trustworthy by putting it in a contest it can't referee.
Dates are 2026. New notes land most days; the weekly essays live at Research.