Crowdsourcing quality control: proof, gold questions and fair reviews
Crowdsourcing quality control is not about catching cheaters. It is a system: tasks that are hard to misunderstand, proof that is easy to verify, known answers that measure accuracy, and a review routine that scales from 10 submissions to thousands while staying fair to the people who did the work. This guide walks through each layer and how it maps to Tasklify.
The layers of crowdsourcing quality control
Each layer catches problems the previous one missed. Skipping the early layers makes the later ones expensive.
| Layer | What it prevents | Tools on Tasklify |
|---|---|---|
| 1. Task design | Honest mistakes, misunderstandings | Steps, instructions, proof note, fair reward |
| 2. Targeting | Wrong device, wrong market, inexperienced workers | Countries or Tier 1, device, min approval rate, min completed tasks, Plus-only |
| 3. Proof | Skipped work | Screenshot, text, link, username/email, screen recording, file upload |
| 4. Gold and redundancy | Careless or random answers | Known-answer items in your content; multiple workers per item |
| 5. Automated signals | Copied or reused proof | Duplicate-file and duplicate-text flags |
| 6. Review | Everything that got through | Approve, request changes once, reject with reason |
| 7. Disputes | Unfair decisions on either side | Worker disputes within 7 days; Tasklify decides |
Most “bad workers” problems are layer 1 problems. If more than a handful of people get the same thing wrong, the task is unclear. Fix the instructions before you tighten anything else; our guide on how to write a micro task covers that layer in depth.
Proof design
Good proof has three properties:
- It shows the outcome, not the effort. A screenshot of the confirmation screen, not of the home page.
- It contains a unique detail. An order number, the worker’s username on the screen, a completion code, the third word of your welcome email.
- It is cheap to check. You should be able to verify it in seconds. Match usernames or emails against your own database; paste codes into a lookup.
| Task | Weak proof | Strong proof |
|---|---|---|
| Sign-up flow | “Screenshot of the app” | Username you signed up with + screenshot of the welcome screen showing it |
| Survey | “Type DONE” | Unique completion code from the last page |
| Usability test | “Describe your experience” | Unedited screen recording with narration + 3 structured answers |
| Data lookup | “The values” | Values in a fixed format with the source URL for each |
Gold questions and redundancy
Gold questions
A gold question is an item with a known correct answer, hidden among the real work: a product whose price you already know in a data batch, an image whose category is obvious in a tagging task, an instructed-response item in a survey. Gold questions let you measure accuracy per worker without re-doing the work yourself.
- Use one or two per submission; more wastes paid time.
- Make them unambiguous. A gold item that experts disagree on measures nothing.
- Keep them indistinguishable from real items and rotate them over time.
- State in the task that the work contains checks. You do not need to say which items.
Redundancy
For judgement tasks (categorising, rating, transcribing messy text), let two or three workers do the same item and compare. Agreement is strong evidence; disagreement tells you where to look. It multiplies reward cost, so reserve it for data that matters. The data entry guide shows a double-entry setup with costs.
Reviewing at scale
On Tasklify, submissions wait in your review queue for the review period you chose (24 hours, 48 hours, 72 hours or 7 days). Rewards stay held until you decide. A triage routine keeps review fast:
- Flagged first. Open submissions marked as duplicate file or duplicate text.
- Gold failures next. Check any submission that got a gold item wrong.
- Sample the rest. Look closely at a random share; skim the others for the unique detail.
- Approve clear passes promptly. Fast approvals build a reputation that attracts careful workers to your next task.
- Log reasons. Keep a tally of rejection reasons. A reason that repeats is a task-design fix.
Worked example. A 5-minute sign-up feedback task at $0.35 with 400 workers costs $140 in rewards + $16.80 service fee (12%). Suppose 30 submissions are flagged or fail the gold check. Reviewing those plus a 10% sample of the rest means looking closely at about 67 submissions instead of 400, and skimming the others for the unique detail.
With auto-approve on, anything you have not reviewed by the deadline is approved automatically, so schedule review time inside your chosen period.
Request changes before rejecting
If the worker did the task but the proof is incomplete (screenshot cut off, answer in the wrong format), use request changes. It is available once per submission; the worker gets your note and resubmits before a deadline, then you approve or reject. It recovers good work and costs you nothing extra.
Fraud signals
| Signal | What it may mean | What to do |
|---|---|---|
| Duplicate file flag | Same screenshot or recording submitted more than once | Compare with the matching submission; reject copies |
| Duplicate text flag | Copied answers or a shared completion code | Check whether the text is legitimately short (e.g. “yes”) before rejecting |
| Wrong device or language on screen | Targeting misrepresented, or proof from someone else | Reject against the device requirement |
| Signs of editing | Fabricated proof | Reject; report repeated cases to support |
| Missing unique detail | Work not actually done | Request changes if plausible, otherwise reject |
| Impossibly fast completion | Skipped steps | Check gold items and the unique detail |
Tasklify’s rules forbid workers from using bots, emulators, location-masking VPNs, multiple accounts and fabricated or reused proof, and violations can lead to forfeited balances and bans. Your job is to reject the submission with a clear reason; report patterns so the platform can act on the account.
Fair rejections and disputes
The Terms require employers to approve every submission that reasonably satisfies the published instructions and to reject only for a legitimate reason, stated in writing. Fair rejections protect your data quality and your reputation with workers.
Rejection reason templates
- “The screenshot does not show the Welcome screen with your username, as required in step 4.”
- “The completion code does not match any response in our survey.”
- “The answer is identical to another submission; the task requires your own feedback.”
- “The recording starts after step 2; the task asks to record from the beginning.”
- “Two of the three check items are incorrect; the task states the batch contains accuracy checks.”
Every reason should point to a rule that was in the task before the worker reserved it. “Low quality” alone is not a reason.
How disputes work
A worker can dispute a rejection within 7 days. You can respond with evidence, and Tasklify reviews the published instructions and the proof, normally within 48 hours; the decision releases the held reward to the worker or returns it to you. Repeated unjustified rejections can lead to restrictions on an employer account, just as repeated unfounded disputes can for workers. Clear instructions and specific reasons mean disputes are rare and quick. Read more on the safety page.
Frequently asked questions
What is a gold question in crowdsourcing?
A gold question (or gold-standard item) is an item whose correct answer you already know, mixed in with the real work. Comparing workers’ answers on gold items measures their accuracy without checking every submission.
How many submissions should I review manually?
Review everything in a pilot. Once accuracy is stable, review every flagged or gold-failing submission plus a random sample of the rest, and widen the sample if you see problems.
When should I use request changes instead of rejecting?
When the worker clearly did the work but missed something fixable, such as a missing screenshot detail or wrong format. You can request changes once per submission; the worker resubmits before a deadline and you then approve or reject.
What happens if a worker disputes my rejection?
The worker can open a dispute within 7 days of the rejection. Both sides can add evidence, and Tasklify reviews the published instructions and the proof, normally deciding within 48 hours. The decision releases or returns the held reward.
Should I turn on auto-approve?
Turn it on when you are confident in your task design and cannot always review on time; submissions not reviewed within your review period are then approved automatically. Keep it off for new or high-stakes tasks, and review within the period either way.
What are common signs of fraudulent submissions?
Duplicate files or text across submissions (Tasklify flags these), screenshots that do not match the requested device or show edits, proof that is missing the unique detail you asked for, and completion times that are impossibly short.