verifier.org

best 8 data labeling and rlhf vendors

ranked on what they'll tell you: what it costs, who actually does the labelling, and what the public record says about how those people are treated.

last reviewed 29 jul 2026 · 8 tools tested ·list curated by Onur Ozcanxin

the short version
best overallSnorkel AIteams with domain-heavy text and documents who would rather write labelling functions than commission a crowd.80/100runner-upArgillateams with their own annotators, or open communities labelling public datasets.78/100best free optionScale AIfrontier-scale programmes and government work, for buyers who have weighed the meta relationship.54/100

seven of these eight publish no price at all. not a per-task rate, not a per-hour rate, not a starting figure — pricing pages that 404, demo forms, or nothing. mercor is the single exception, and it publishes contractor rate bands rather than customer pricing. so the first honest thing to say about this category is that you cannot compare it on cost without eight sales conversations.

the second is that the labelling is done by contractors almost everywhere, and their pay is disclosed almost nowhere. no vendor here publishes a wage floor. what exists instead is reporting and litigation: at scale's remotasks, pay as low as around a cent for short tasks per the business and human rights resource centre; at toloka's crowd, $1 to $6 an hour per worker-review sites; at mercor, published bands of $60 to $120 an hour for credentialed experts. these are different businesses wearing one label.

the legal record is a purchasing consideration, not gossip, and it is unusually active. scale faces two us suits filed in january 2025 — one over wages and misclassification, one over the mental-health toll of content moderation at outlier — and shut its kenya operation abruptly in march 2024 with wages reportedly owed. surge faces a contractor-misclassification class action. mercor faces both a proposed misclassification class action covering some 30,000 contractors and at least six actions arising from a 2026 breach in which roughly four terabytes were exfiltrated including social security numbers, passports, biometric face and voice data and banking details. none of these has produced a finding of liability.

one structural change reshaped the top of the market. meta took a reported 49% stake in scale ai for about $14.3 billion in june 2025 and scale's founder left to run meta's superintelligence lab; google, openai and microsoft were reported to have reduced or ended their scale engagements afterwards over confidentiality and competitive concerns. if you are training a model, the largest vendor in this category is now part-owned by one of your competitors.

advertisement
  1. 1

    Snorkel AI

    reduces how much human labelling you need in the first place, which is the only structural answer to this category's problems.

    80/100

    verdictthe only vendor whose product reduces the amount of contract labour involved rather than organising it — and like everyone here, it won't tell you what that costs.

    best for
    teams with domain-heavy text and documents who would rather write labelling functions than commission a crowd.
    price
    not published
    pricing note
    snorkel.ai/pricing returned 404 and no rate card was found elsewhere; the model is quote-only
    free tier
    no
    published pricing
    none — page 404s
    workforce
    programmatic + expert contractors
    pay disclosed
    no
    documented disputes
    none found
    rlhf
    yes

    snorkel flow's premise is programmatic supervision: you encode labelling rules and weak signals rather than paying people to annotate every example, which cuts manual volume substantially. that is a genuinely different answer to a category where the ethical and cost problems both scale with human hours.

    it also sells expert data-as-a-service, contracting phds and subject-matter specialists directly rather than routing through a crowd platform, positioned for rlhf ranking, code evaluation and creative judgement. we found no litigation or regulatory action against snorkel, which distinguishes it from four vendors here — though absence of found reports is weaker evidence than a clean finding.

    the gaps are the category norm. no pricing at all, no worker locations, no wage floor, and no data-ownership terms we could reach. worth knowing that its $100m series d in may 2025 at a $1.3bn valuation included in-q-tel, the cia's venture arm, among the investors — not a criticism, but a fact some buyers weigh.

    pros
    • +programmatic supervision cuts manual annotation volume
    • +expert contributors contracted directly, not via a crowd
    • +no litigation or regulatory action found
    • +positioned for rlhf ranking and code evaluation
    cons
    • pricing page returns 404 — no rate published
    • no worker locations or wage floor disclosed
    • multimodal breadth beyond text and documents unverified
    • data ownership terms not reachable
  2. 2

    Argilla

    free open-source annotation tooling with no workforce attached — which is both its strength and the reason it can't finish first.

    78/100

    verdictthe only entry with no labour record to examine, because it employs nobody — you supply the people, and their working conditions become your problem rather than a vendor's.

    best for
    teams with their own annotators, or open communities labelling public datasets.
    price
    free
    pricing note
    core argilla is free open-source software; hugging face's enterprise hub adds sso and audit logs at general hub pricing, with no argilla-specific rate
    free tier
    yes
    published pricing
    free, open source
    workforce
    none — you supply it
    pay disclosed
    n/a
    documented disputes
    none — no workforce
    rlhf
    tooling only

    argilla is annotation software rather than an annotation service. hugging face acquired it in june 2024 for a reported $10m, and the core remains free and open source, deployable on your own hugging face space or infrastructure, tightly integrated with the hub.

    that architecture answers the data question cleanly — datasets stay under your control, and there is no vendor in the path reusing them — and it removes the pay-transparency problem by removing the workforce. no contractors, no wage floor to publish, no disputes to report.

    it also cannot do the job the other seven do. somebody still has to label the data, and argilla supplies the interface rather than the hands. if you don't have in-house annotators you will end up hiring a vendor from elsewhere on this page anyway, and inheriting that vendor's labour practices along with its output. rlhf is a supported workflow; the human effort behind it is yours to source.

    pros
    • +free and open source, no software cost
    • +deep integration with the hugging face hub
    • +datasets stay under your control
    • +no contractor labour risk because there is no contractor labour
    cons
    • supplies no workforce at all
    • you must source annotators separately
    • no argilla-specific enterprise pricing published
    • reuse terms for private enterprise deployments unverified
  3. 3

    Labelbox

    annotation platform plus an expert marketplace, with the highest reported pay band of the crowd-model vendors.

    74/100

    verdictthe most complete combination of platform and workforce here, reported to pay contributors well — with an unpaid screening stage that only shows up in worker reviews.

    best for
    teams wanting tooling and on-demand domain experts from one vendor, including medical and legal specialists.
    price
    not published
    pricing note
    labelbox.com/pricing returned 404 during this review; the model is quote-only with a platform fee
    free tier
    no
    published pricing
    none — page 404s
    workforce
    alignerr contractor marketplace
    pay disclosed
    no — $15-60/hr reported
    documented disputes
    none found; reviews only
    rlhf
    yes

    labelbox pairs its annotation platform with alignerr, its expert contractor marketplace, claiming 250,000-plus specialised hours a month spanning entry-level work up to medical and legal domain experts. a third-party review site reports contractor pay of roughly $15 to $60 an hour depending on domain — the highest reported band among the crowd-model vendors here, though labelbox publishes no figure itself.

    the same review site reports unpaid or low-paid multi-hour evaluation stages before contributors reach paid projects. we found no court filing or regulatory action on this, so it stands as aggregated worker anecdote rather than a documented dispute, and we're labelling it that way rather than either amplifying or ignoring it.

    it raised a $110m series e at roughly a $1bn valuation in october 2025 led by softbank vision fund 2 and a16z, and grew from 232 to 551 employees between 2024 and mid-2026. its pricing page returns 404, and its data-processing terms weren't reachable either, so who owns the labelled output is unverified.

    pros
    • +platform and expert marketplace from one vendor
    • +highest reported contractor pay band of the crowd vendors
    • +domain experts including medical and legal
    • +no litigation or regulatory action found
    cons
    • pricing page returns 404
    • unpaid multi-hour evaluation stages reported by contributors
    • no wage floor or worker locations published
    • data ownership terms not reachable
    advertisement
  4. 4

    Toloka

    the largest open crowd here at the lowest reported pay, and a corporate lineage worth tracing before you sign.

    70/100

    verdictgenuine global reach at genuine crowd-labour rates — around $1 to $6 an hour by third-party reports — from a company still majority-owned by nebius.

    best for
    very high-volume microtask labelling across many languages where unit cost dominates.
    price
    not published
    pricing note
    no dedicated pricing page found; customer rates are quote-only and worker pay appears only on third-party gig-review sites
    free tier
    no
    published pricing
    none
    workforce
    open crowd platform
    pay disclosed
    no — $1-6/hr reported
    documented disputes
    none found
    rlhf
    yes

    toloka runs an open crowd platform across a claimed 100-plus countries and 40-plus languages, where workers sign up directly and complete microtasks. for sheer breadth and unit economics at volume, nothing else here matches it, and it reportedly serves amazon, microsoft and anthropic.

    worker pay is where the model shows. third-party review sites report roughly $1 to $6 an hour for casual tasks, rarely above $8 even for experienced workers, with typical monthly earnings from $20 to $400 depending on intensity. toloka publishes no wage floor and no country-level pay breakdown, so those figures are worker-reported rather than disclosed. we found no litigation specific to toloka.

    the ownership is easy to miss. toloka was spun out of yandex and became part of nasdaq-listed nebius group; a $72m investment led by bezos expeditions in may 2025 ended nebius's majority voting control while nebius retained a significant majority economic stake. for buyers with jurisdictional sensitivities, that history is worth knowing before it comes up in a security review.

    pros
    • +100+ countries and 40+ languages
    • +the deepest crowd for high-volume microtasks
    • +reportedly serves amazon, microsoft and anthropic
    • +no litigation found specific to toloka
    cons
    • reported crowd pay of roughly $1-6 an hour
    • no published wage floor or pay breakdown
    • no customer pricing published anywhere
    • still majority economically owned by nebius group
  5. 5

    Surge AI

    bootstrapped, profitable and serving the frontier labs — facing a class action over how it classifies the people doing the work.

    68/100

    verdictthe quality reputation in this category and no venture capital behind it — with a misclassification class action that goes to the heart of how the work is organised.

    best for
    frontier-lab-grade rlhf and reasoning data where quality outweighs procurement transparency.
    price
    not published
    pricing note
    no customer rate card; the homepage cites a single unrepresentative figure for one benchmark role rather than a price list
    free tier
    no
    published pricing
    none
    workforce
    ~50,000 expert contractors
    pay disclosed
    no — 30-40c/min reported
    documented disputes
    misclassification class action
    rlhf
    yes — core focus

    surge built to a reported $1.2bn revenue run-rate without raising venture capital, which in a category defined by capital intensity is unusual and gives its founder full control. it focuses on rlhf, rl environments and language annotation for frontier labs, reportedly including openai, anthropic, meta and microsoft, working through around 50,000 expert contractors alongside roughly 130 employees.

    contractor pay is reported by third parties at 30 to 40 cents per working minute — roughly $18 to $24 an hour if the work is continuous — which is well above crowd rates and below mercor's published expert bands. surge publishes no wage floor and no country breakdown.

    a class action alleges surge misclassified data annotators as independent contractors, denying benefits, with claims of unpaid training time and time limits tight enough to reduce effective pay. the case is pending and no liability has been determined. worth noting the shape of the allegation matters more than the vendor: the same claim is live against mercor and was made against scale.

    pros
    • +bootstrapped and profitable, no external investors
    • +reported contractor rates well above crowd platforms
    • +focused specifically on rlhf and rl environments
    • +reportedly serves openai, anthropic, meta and microsoft
    cons
    • contractor-misclassification class action pending
    • no customer pricing published at all
    • no wage floor or worker locations disclosed
    • modality coverage beyond text unverified
  6. 6

    Invisible Technologies

    human-plus-automation services at enterprise scale, with worker complaints that echo the category's pattern.

    64/100

    verdicta genuine enterprise operation with a dedicated rlhf practice — and worker accounts describing the misclassification and monitoring pattern seen elsewhere here.

    best for
    enterprises outsourcing blended human-and-automation work rather than pure annotation.
    price
    not published
    pricing note
    no public rate card; enterprise quote-only with 'services-to-software' positioning
    free tier
    no
    published pricing
    none
    workforce
    ~24,000 vetted contractors
    pay disclosed
    no
    documented disputes
    worker reviews only
    rlhf
    yes — dedicated practice

    invisible blends human workers with automation across several products, with meridial described as a network of 24,000-plus vetted experts and a dedicated ai training and rlhf service line. it raised $100m in september 2025 at a reported valuation above $2bn, and cites soc 2, hipaa and gdpr compliance.

    worker reviews on indeed and glassdoor describe being treated as employees while classified as contractors, delayed payments, and time-tracking software not disclosed until after a contract is signed. these are aggregated anecdotes rather than court filings — we found no litigation or regulatory finding — and we report them as what they are, while noting the pattern matches allegations that have been formally pleaded against other vendors here.

    one practical warning for anyone researching this: there is a name collision. a separate, smaller company also called invisible, working on ai-agent infrastructure, was acquired by perplexity in august 2025. several aggregators conflate the two, which would tell you this vendor had been acquired when it had not.

    pros
    • +dedicated ai training and rlhf service line
    • +soc 2, hipaa and gdpr compliance cited
    • +blends human work with automation at enterprise scale
    • +raised $100m at a reported $2bn+ valuation
    cons
    • worker reviews describe misclassification-style treatment
    • monitoring software reportedly undisclosed until after signing
    • no pricing, wage floor or worker locations published
    • name collision with an unrelated acquired company
  7. 7

    Mercor

    the only vendor that publishes what it pays, and the only one that lost its contractors' passports and biometrics.

    60/100

    verdictthe best pay transparency in this category by a wide margin, and the worst documented harm to the people it pays.

    best for
    sourcing credentialed experts — physicians, lawyers, senior engineers — at rates you can see before engaging.
    price
    $60-120/hr contractor bands
    pricing note
    publishes indicative contractor hourly rates by role rather than customer pricing; reporting cites over $2m paid daily to contractors and 60-70% of top-line revenue passed through
    free tier
    no
    published pricing
    contractor bands only
    workforce
    expert contractor marketplace
    pay disclosed
    yes — $60-120/hr bands
    documented disputes
    breach + misclassification suits
    rlhf
    implied, not branded

    mercor is the only vendor here that publishes rates at all: bands of roughly $60 to $120 an hour for specialised roles, with reporting citing over $2m paid daily across its network and 60 to 70% of top-line revenue passed to contractors. after seven vendors publishing nothing, that deserves credit, and it is why the entry isn't last.

    the security record is the reason it is seventh. a 2026 supply-chain compromise of litellm, attributed to the lapsus$ group, exfiltrated roughly four terabytes of data reportedly including social security numbers, dates of birth, passports, biometric face and voice data and banking details. at least six to seven class actions followed, including gill v. mercor in the northern district of california filed 1 april 2026. a separate suit, cox v. mercor io corporation filed october 2025, alleges undisclosed monitoring software installed on contractors' personal computers without reimbursement.

    a proposed federal class action in texas additionally alleges independent-contractor misclassification on behalf of around 30,000 experts including physicians, attorneys, bankers and engineers. all of these are pending and no liability has been determined. mercor was reported in july 2026 to be raising at a $20bn valuation.

    pros
    • +publishes contractor rate bands, unique in this category
    • +reportedly passes 60-70% of revenue to contractors
    • +credentialed experts including physicians and attorneys
    • +over $2m reportedly paid to contractors daily
    cons
    • 2026 breach reportedly exposed ssns, passports and biometrics
    • at least six class actions arising from that breach
    • separate suit alleges undisclosed monitoring on personal devices
    • misclassification class action covering ~30,000 contractors
  8. 8

    Scale AI

    the biggest vendor in the category, part-owned by one of your competitors, with the longest labour record to read.

    54/100

    verdictunmatched capacity and capital, ranked last here because this page measures transparency and labour record — and on both, scale has the most on the record.

    best for
    frontier-scale programmes and government work, for buyers who have weighed the meta relationship.
    price
    not published
    pricing note
    no per-task, per-hour or per-seat rate published; a self-serve teaser offers the first 1,000 labelling units and 10,000 images free, then enterprise sales
    free tier
    yes
    published pricing
    none beyond a free teaser
    workforce
    remotasks crowd + outlier experts
    pay disclosed
    no — ~1c/task reported
    documented disputes
    two pending us suits
    rlhf
    yes — core offering

    on capability scale is the leader: the deepest capacity for frontier training data and rlhf, the most capital, and the most government contract wins including a department of defense engagement reported above $300m. a ranking built on throughput would put it first, and this one is not.

    the ownership changed the market. meta took a reported 49% stake for about $14.3bn in june 2025, valuing scale at $29bn, and founder alexandr wang left to lead meta's superintelligence lab. google, openai and microsoft were reported to have reduced or ended engagements afterwards over confidentiality and competitive concerns. scale still operates separately, but a large minority owner is now a direct rival to several of its customers.

    the labour record is the longest here. remotasks operations in the philippines have been documented by the business and human rights resource centre with pay as low as around a cent for short tasks and reports of withheld payments; the kenya operation shut abruptly in march 2024 with wages reportedly owed and the country's data labelers association weighing action. two us suits were filed in january 2025 — one alleging wage violations and misclassification with effective pay near $15 an hour, one over the mental-health toll of content-moderation work at outlier. all pending, no liability determined. scale cut around 200 employees and 500 contractors in july 2025 while pivoting toward defence.

    pros
    • +deepest capacity for frontier training data and rlhf
    • +most capital and the strongest government contract position
    • +covers text, image, video, audio and code
    • +free tier of 1,000 labelling units and 10,000 images
    cons
    • meta holds a reported 49% stake, creating a competitive conflict
    • major labs reportedly reduced engagements after that deal
    • kenya operation closed in 2024 with wages reportedly owed
    • two pending us suits over wages and content-moderation harm

how this ranking was made

we rank on transparency and labour record rather than on capability, and that choice changes the order substantially — the largest and best-capitalised vendor finishes last. a buyer choosing purely on throughput would rank this differently, and we think the things we measure are the things a public ranking can actually establish.

every claim about pay, disputes or corporate events is attributed. court filings and company announcements are treated as primary; named publications as secondary; worker-review sites like glassdoor and indeed as aggregated anecdote, labelled as such and never presented as findings. where only a review site reports something — labelbox's unpaid evaluation stages, invisible's monitoring software — the entry says that is where it comes from.

pending litigation is described as pending. we state what a complaint alleges, who filed it and when, and we say plainly that no liability has been determined. these are live matters concerning real people and real companies, and the point is to tell a buyer what exists on the record, not to reach a verdict.

argilla is included but is a different kind of product: annotation software you run yourself, with no workforce attached. it therefore has no pay or labour record to report, and it also cannot answer the question this category exists to answer — someone still has to do the labelling, and that responsibility stays with you.

data ownership could not be verified for most of these vendors. terms and data-processing pages returned 404 at scale, labelbox and snorkel, and were not locatable at several others, so who owns the labelled output and whether the vendor may reuse it is marked unverified rather than assumed favourable.

our general methodology and disclosures →
was this useful?