Train the agents
that do the work.

You know what good looks like. Sharpen agents against real cases, catch the edge cases, and lift accuracy.

Get paid to make agents better.

How it works.

You are the judgment that makes an agent trustworthy. This is where that judgment gets paid.

Notes and a laptop during an application review
01

Apply

Tell us the domain and functions you can judge better than almost anyone. We review and, if it is a fit, invite you.

Reviewing work closely on a laptop screen
02

Train & evaluate

Review agent outputs, label the edge cases, write evaluation cases, and give the feedback that makes it sharper.

A calm, finished workspace
03

Get paid

You are paid for the training work. Aonami handles the platform, the contract, and delivery.

What training looks like.

The work is judgment, not code.

Review

Review the
agent's output.

Judge what the agent produced against what a great human would do in the same case, and mark exactly where it is right or wrong.

Label

Label the
edge cases.

The exceptions and gray areas are where agents fail. You know them cold, so you flag them and explain why they are hard.

Evaluate

Write the
evaluation set.

Turn your know-how into gold-standard test cases the agent is measured against, every version from here on.

Iterate

Raise the
accuracy bar.

Your feedback loops back into the agent, so it gets sharper with every cycle, not just once.

Who this is for.

You do not need to code. You need to know what good looks like.

An expert presenting to a room

Domain experts

You know the rules and the exceptions cold, better than any model can guess.

A team inspecting a large site from above

Reviewers & QA leads

You already judge work for a living. Now that judgment trains an agent.

Stacked boxes ready to ship

Ex-operators

You ran the function for years and know exactly what good looks like.

Agents you could train.

These are the kinds of agents that need your judgment. Open any one to see how it works.

Questions.

How training works, how you are paid, and what we protect.

No. You judge outputs and edge cases; the platform handles the rest.

Apply to train.

Invite-only. We review every application. Takes about two minutes.