Applied research lab
Expert onchain judgment for frontier agents
We capture how crypto experts actually reason — then structure it as training data.
Assay turns exploits, wash volume, governance, and safe signing into SFT traces, RL rubrics, and Base-native agent environments.
Built for frontier labs and Base Batches agent teams.
Purity test
Problem
Frontier models can chat about crypto. They fail at real onchain work.
They cannot tell wash from organic stablecoin flow. They cannot read an incident postmortem into a safe action. They cannot decide whether an agent should sign. They cannot score a launch for sniper and rug risk. They write SQL that matches transfer spam, not economic reality.
That knowledge is not on the internet as clean training data. It lives in protocol engineers, auditors, MEV searchers, onchain PMs, and compliance analysts.
Solution
Assay is an applied research lab for onchain work. We capture how experts actually think — decisions, tradeoffs, context — and structure it into training and eval data.
Models trained on outputs plateau. Models trained on reasoning improve. An agent does not pass because the answer looks fluent. It passes a named expert rubric.
Four products
All productsSFT / chain-of-thought
Traces of how an expert actually works a live onchain case. Prompt–response pairs and reasoning, not vibes.
RL + rubrics
Named grading frameworks. “Is this volume organic?” “Is this action safe to sign?” Fluent is not a pass.
Agent environments
Wallet, Base explorer, SQL over decoded schemas, MCP tools. Agents run real workflows.
Computer-use trajectories
LATERHuman demos of explorer / console / incident response.
Why this domain
Agents are about to move money.
Agent wallets, x402, Base Batches. Training data for those agents is still generic web text. That is the gap Assay fills.
Default unit is USDC. Default canary network is Base. Other chains later.
We do not re-index 150 chains. We sell judgment, datasets, and harnesses.
Sample tasks
Open sample packHarness
A Base eval sketch, not a dashboard.
Wallet, explorer, canned SQL, canned transcript, weighted rubric. Run the sample query. Run eval. That single interaction is the demo.
Wallet
0xA55A…c01d
USDC
Transcript
Volume is 8.4×. 612 senders. Looks organic.
Rubric
FAIL 0.00
Fluent is not a pass.
Experts
Named, qualified people. Not clickworkers.
Protocol engineers, auditors, MEV searchers, onchain PMs, compliance analysts. Paid for traces and rubrics. Provenance is a role and years, never a fake name.