Crowdsourcing quality control: proof, gold questions and fair reviews

By Tasklify TeamUpdated 9 min read

Crowdsourcing quality control is not about catching cheaters. It is a system: tasks that are hard to misunderstand, proof that is easy to verify, known answers that measure accuracy, and a review routine that scales from 10 submissions to thousands while staying fair to the people who did the work. This guide walks through each layer and how it maps to Tasklify.

The layers of crowdsourcing quality control

Each layer catches problems the previous one missed. Skipping the early layers makes the later ones expensive.

LayerWhat it preventsTools on Tasklify
1. Task designHonest mistakes, misunderstandingsSteps, instructions, proof note, fair reward
2. TargetingWrong device, wrong market, inexperienced workersCountries or Tier 1, device, min approval rate, min completed tasks, Plus-only
3. ProofSkipped workScreenshot, text, link, username/email, screen recording, file upload
4. Gold and redundancyCareless or random answersKnown-answer items in your content; multiple workers per item
5. Automated signalsCopied or reused proofDuplicate-file and duplicate-text flags
6. ReviewEverything that got throughApprove, request changes once, reject with reason
7. DisputesUnfair decisions on either sideWorker disputes within 7 days; Tasklify decides

Most “bad workers” problems are layer 1 problems. If more than a handful of people get the same thing wrong, the task is unclear. Fix the instructions before you tighten anything else; our guide on how to write a micro task covers that layer in depth.

Proof design

Good proof has three properties:

  • It shows the outcome, not the effort. A screenshot of the confirmation screen, not of the home page.
  • It contains a unique detail. An order number, the worker’s username on the screen, a completion code, the third word of your welcome email.
  • It is cheap to check. You should be able to verify it in seconds. Match usernames or emails against your own database; paste codes into a lookup.
TaskWeak proofStrong proof
Sign-up flow“Screenshot of the app”Username you signed up with + screenshot of the welcome screen showing it
Survey“Type DONE”Unique completion code from the last page
Usability test“Describe your experience”Unedited screen recording with narration + 3 structured answers
Data lookup“The values”Values in a fixed format with the source URL for each

Gold questions and redundancy

Gold questions

A gold question is an item with a known correct answer, hidden among the real work: a product whose price you already know in a data batch, an image whose category is obvious in a tagging task, an instructed-response item in a survey. Gold questions let you measure accuracy per worker without re-doing the work yourself.

  • Use one or two per submission; more wastes paid time.
  • Make them unambiguous. A gold item that experts disagree on measures nothing.
  • Keep them indistinguishable from real items and rotate them over time.
  • State in the task that the work contains checks. You do not need to say which items.

Redundancy

For judgement tasks (categorising, rating, transcribing messy text), let two or three workers do the same item and compare. Agreement is strong evidence; disagreement tells you where to look. It multiplies reward cost, so reserve it for data that matters. The data entry guide shows a double-entry setup with costs.

Reviewing at scale

On Tasklify, submissions wait in your review queue for the review period you chose (24 hours, 48 hours, 72 hours or 7 days). Rewards stay held until you decide. A triage routine keeps review fast:

  1. Flagged first. Open submissions marked as duplicate file or duplicate text.
  2. Gold failures next. Check any submission that got a gold item wrong.
  3. Sample the rest. Look closely at a random share; skim the others for the unique detail.
  4. Approve clear passes promptly. Fast approvals build a reputation that attracts careful workers to your next task.
  5. Log reasons. Keep a tally of rejection reasons. A reason that repeats is a task-design fix.

Worked example. A 5-minute sign-up feedback task at $0.35 with 400 workers costs $140 in rewards + $16.80 service fee (12%). Suppose 30 submissions are flagged or fail the gold check. Reviewing those plus a 10% sample of the rest means looking closely at about 67 submissions instead of 400, and skimming the others for the unique detail.

With auto-approve on, anything you have not reviewed by the deadline is approved automatically, so schedule review time inside your chosen period.

Request changes before rejecting

If the worker did the task but the proof is incomplete (screenshot cut off, answer in the wrong format), use request changes. It is available once per submission; the worker gets your note and resubmits before a deadline, then you approve or reject. It recovers good work and costs you nothing extra.

Fraud signals

SignalWhat it may meanWhat to do
Duplicate file flagSame screenshot or recording submitted more than onceCompare with the matching submission; reject copies
Duplicate text flagCopied answers or a shared completion codeCheck whether the text is legitimately short (e.g. “yes”) before rejecting
Wrong device or language on screenTargeting misrepresented, or proof from someone elseReject against the device requirement
Signs of editingFabricated proofReject; report repeated cases to support
Missing unique detailWork not actually doneRequest changes if plausible, otherwise reject
Impossibly fast completionSkipped stepsCheck gold items and the unique detail

Tasklify’s rules forbid workers from using bots, emulators, location-masking VPNs, multiple accounts and fabricated or reused proof, and violations can lead to forfeited balances and bans. Your job is to reject the submission with a clear reason; report patterns so the platform can act on the account.

Quality control also applies to what you ask for. The Acceptable Use Policy prohibits tasks that buy fake reviews, followers, likes or any engagement that deceives third parties or breaks another platform’s rules. No review process makes that kind of task acceptable.

Fair rejections and disputes

The Terms require employers to approve every submission that reasonably satisfies the published instructions and to reject only for a legitimate reason, stated in writing. Fair rejections protect your data quality and your reputation with workers.

Rejection reason templates

  • “The screenshot does not show the Welcome screen with your username, as required in step 4.”
  • “The completion code does not match any response in our survey.”
  • “The answer is identical to another submission; the task requires your own feedback.”
  • “The recording starts after step 2; the task asks to record from the beginning.”
  • “Two of the three check items are incorrect; the task states the batch contains accuracy checks.”

Every reason should point to a rule that was in the task before the worker reserved it. “Low quality” alone is not a reason.

How disputes work

A worker can dispute a rejection within 7 days. You can respond with evidence, and Tasklify reviews the published instructions and the proof, normally within 48 hours; the decision releases the held reward to the worker or returns it to you. Repeated unjustified rejections can lead to restrictions on an employer account, just as repeated unfounded disputes can for workers. Clear instructions and specific reasons mean disputes are rare and quick. Read more on the safety page.

Frequently asked questions

What is a gold question in crowdsourcing?

A gold question (or gold-standard item) is an item whose correct answer you already know, mixed in with the real work. Comparing workers’ answers on gold items measures their accuracy without checking every submission.

How many submissions should I review manually?

Review everything in a pilot. Once accuracy is stable, review every flagged or gold-failing submission plus a random sample of the rest, and widen the sample if you see problems.

When should I use request changes instead of rejecting?

When the worker clearly did the work but missed something fixable, such as a missing screenshot detail or wrong format. You can request changes once per submission; the worker resubmits before a deadline and you then approve or reject.

What happens if a worker disputes my rejection?

The worker can open a dispute within 7 days of the rejection. Both sides can add evidence, and Tasklify reviews the published instructions and the proof, normally deciding within 48 hours. The decision releases or returns the held reward.

Should I turn on auto-approve?

Turn it on when you are confident in your task design and cannot always review on time; submissions not reviewed within your review period are then approved automatically. Keep it off for new or high-stakes tasks, and review within the period either way.

What are common signs of fraudulent submissions?

Duplicate files or text across submissions (Tasklify flags these), screenshots that do not match the requested device or show edits, proof that is missing the unique detail you asked for, and completion times that are impossibly short.