Research

Lab notes. Not published benchmarks.

These are working notes from the canary tasks. Dates are 2026. They are honest placeholders until we publish evals.

12 Aug 2026

Economic signal vs noise in Base USDC flow

Transfer count is not flow. On Base USDC, a pool can print an 8× day from two wash rings and searcher inventory while unique-actor taker flow barely moves. This note is a lab sketch of the clustering, circularity, and time/size tests in the first canary task. It is not a published benchmark.

28 Aug 2026

Rubric-gated evals for agents that can sign

If an agent can sign, fluency is the wrong pass condition. Permit2 max-approvals, stale deadlines, and unverified spenders are policy failures even when the model explains itself well. This note records why the signing canary refuses unbounded approvals and what a 0–1 rubric score is for. Lab note, not a leaderboard.

4 Sep 2026

Why generic web data fails the moment money moves

Web text teaches models to talk about crypto. It does not teach them to cluster funders, decode Permit2, or refuse a pause they cannot call. The gap is not more pages. It is expert judgment under an explicit rubric, on a named chain, with a named unit. Base and USDC first, because that is where agents are about to transact.