Train the agents
that do the work.
You know what good looks like. Sharpen agents against real cases, catch the edge cases, and lift accuracy.
Get paid to make agents better.
How it works.
You are the judgment that makes an agent trustworthy. This is where that judgment gets paid.
What training looks like.
The work is judgment, not code.
Review the
agent's output.
Judge what the agent produced against what a great human would do in the same case, and mark exactly where it is right or wrong.
Label the
edge cases.
The exceptions and gray areas are where agents fail. You know them cold, so you flag them and explain why they are hard.
Write the
evaluation set.
Turn your know-how into gold-standard test cases the agent is measured against, every version from here on.
Raise the
accuracy bar.
Your feedback loops back into the agent, so it gets sharper with every cycle, not just once.
Who this is for.
You do not need to code. You need to know what good looks like.
Agents you could train.
These are the kinds of agents that need your judgment. Open any one to see how it works.
Questions.
How training works, how you are paid, and what we protect.
No. You judge outputs and edge cases; the platform handles the rest.
Apply to train.
Invite-only. We review every application. Takes about two minutes.



