Outsourcing AI answer review and RLHF data: a small-team guide
AI answer review means people judging model answers against a written rubric: ranking pairs, scoring accuracy and tone, and flagging harm. OpenAI's InstructGPT work used about 40 screened contractors whose judgements agreed about 73% of the time. For a small team, start with a clear rubric, a calibration round, an agreement target and a paid pilot.

What AI answer review covers
| Task | What reviewers do | What you get |
|---|---|---|
| Preference ranking | Compare two or more answers to the same prompt and pick the better one | Ranked pairs for reward models or evaluation |
| Rubric scoring | Score each answer on accuracy, helpfulness, tone and policy | Scores per criterion, with notes |
| Fact checking | Check claims in an answer against sources you allow | Flags with the source that contradicts the claim |
| Safety flags | Mark answers that break your content or safety policy | Labelled examples for policy tuning |
| Ideal answers | Write the answer the model should have given | Reference answers for fine-tuning or tests |
This is how RLHF (reinforcement learning from human feedback) was done in published work. In OpenAI's InstructGPT paper, about 40 contractors chosen by a screening test ranked between 4 and 9 model outputs per prompt. Anthropic's helpful-and-harmless research used human preference comparisons in the same way, updated weekly with new feedback.
Try it: which answer is better?
This is the core of the work. Read the shop's policy and the customer's message, pick the better reply, then compare your choice with the rubric.
Shop policy: late orders get a free express replacement; refunds only if the courier confirms the parcel is lost. Customer: "My order is 5 days late and I need it for Saturday." Reply as the shop.
Tap the answer you think is better.
See the rubric
| Criterion | Better | Why |
|---|---|---|
| Follows policy | B | A promises a refund the policy doesn't allow, because the parcel isn't lost. |
| Meets the real need | B | B checks the delivery date against Saturday and gives a backup plan. |
| Tone | Tie | Both apologise and sound friendly. |
| Clear next step | B | B says exactly what happens, and when. |
| Brevity | A | A is shorter, but brevity doesn't outweigh breaking policy. |
B is better. A sounds kind but breaks the shop's policy, which is exactly the kind of mistake a reviewer is there to catch.
Agreement: the number that tells you the rubric works
When careful reviewers often disagree, check the rubric first: unclear criteria are a common cause. Measure agreement on a shared set of answers before you scale.
- Raw agreement is the share of items where two reviewers chose the same answer. In InstructGPT, training labelers agreed with each other 72.6% of the time, and held-out labelers 77.3%.
- Cohen's kappa adjusts for agreement by chance: 0 means chance level, 1 means perfect. A widely cited guide treats kappa below 0.60 as inadequate and 0.80 to 0.90 as strong.
- With more than two reviewers, or missing ratings, Krippendorff's alpha is the usual measure.
Build a gold set: 50 to 100 answers your own team has scored. New reviewers must match it before they work on live data, and you can re-check against it every week.
Is your project ready to hand off?
Tick what you already have
Tick everything that is true. Your result appears here, and changes as you go.
How to read this
- Write the rubric first: Without a rubric and examples, reviewers will each invent their own, and agreement will be low.
- Nearly ready: Fill the gaps, then run a calibration round on your gold set before live data.
- Ready for a pilot: Run a small paid pilot, measure agreement and your own review time, then scale.
A pilot plan that works for small teams
- Share the rubric, the scored examples and your policies. Our brief builder helps you put it on one page.
- Run a calibration round: reviewers score the gold set, then you go through the disagreements together.
- Update the rubric with every disagreement you resolve.
- Run a paid pilot of a few hundred items, with a share double-reviewed to measure agreement.
- Check a random sample yourself, and note how long your own review takes.
- Scale only when agreement and your review time both hold.
The Partnership on AI makes the same points in its guidance on responsible sourcing of data enrichment: provider selection, a pilot, clear task instructions, fair pay terms and a quality process are the decisions that shape both data quality and workers' conditions.
The rules your data may need to meet
If your model falls under the EU AI Act's high-risk rules, Article 10 requires training, validation and test data to be relevant, representative, and as free of errors and as complete as possible, with checks for bias. The Act applies generally from 2 August 2026, and a 2026 amendment moved the obligations for most high-risk systems to 2 December 2027.
In the US, NIST's AI Risk Management Framework (2023) and its Generative AI Profile (2024) are the common voluntary reference for evaluating AI systems. Either way, a written rubric, an agreement measure and a record of who reviewed what make your data easier to defend.
If prompts or answers contain personal data, share only what the task needs and put a processor agreement in place. Our cost and GDPR guide covers the contracts.
Fair pay is part of quality
Careful review takes time, and pay per task that doesn't allow for it pushes people to rush. Independent ratings of online work platforms have found poor conditions common: in Fairwork's 2022 cloudwork ratings, 12 of 15 platforms met no more than 3 of the 10 fairness thresholds.
Checking AI answers is one of the main kinds of work HerWorkCircle does. The women on your project passed our skill check for careful reading and judgement, they are paid fairly and on time, and our team checks every batch before it reaches you. Send a brief to get a clear price for a pilot.
Questions
What is RLHF data?
Human judgements about model answers, such as which of two answers is better or how an answer scores on a rubric. It is used to train reward models and to evaluate AI systems.
How many reviewers do I need?
Enough that each item in your agreement sample is reviewed twice. Published projects have used dozens of screened reviewers; a small pilot can start with a handful who pass your gold set.
What is a good inter-annotator agreement?
It depends on the task. InstructGPT reported about 73% raw agreement between labelers. For Cohen's kappa, a common guide treats below 0.60 as inadequate and 0.80 or more as strong.
Can I outsource AI evaluation under the EU AI Act?
Yes, but if your system is high-risk, your training and test data must meet Article 10's quality rules, so keep the rubric, the agreement results and a record of the review process.
How do I know the reviewers are careful?
Use a gold set they must match, double-review a share of items, measure agreement, and check a random sample yourself every week.
Sources
- Training language models to follow instructions with human feedback (InstructGPT) Ouyang et al., OpenAI, arXiv (March 2022)
- Training a helpful and harmless assistant with RLHF Bai et al., Anthropic, arXiv (April 2022)
- Interrater reliability: the kappa statistic McHugh, Biochemia Medica (2012)
- Computing Krippendorff's alpha-reliability Krippendorff, University of Pennsylvania
- Regulatory framework for AI (AI Act timeline) European Commission
- AI Act, Article 10: data and data governance AI Act text (artificialintelligenceact.eu)
- AI Risk Management Framework US National Institute of Standards and Technology
- Responsible sourcing of data enrichment services Partnership on AI (June 2021)
- New Fairwork cloudwork rating WZB Berlin Social Science Center (25 August 2022)