Blog · Getting hired

How to get data annotation jobs (and why applications stall)

Published Updated 14 min read

How to get data annotation jobs is mostly not a résumé secret. It is three things: landing on a real requisition you are actually eligible for, passing a fussy test that most people rush, then keeping your quality score high enough in the first fortnight that the queue stays open. This page is the long version of that process, and it answers the searches that sit next to it: how to find a data annotation job, where to find annotation job listings, apply for data annotation jobs, and why you cannot find work as an AI annotator. The apply surface is the live jobs feed.

The five-step loop that actually works

There is no hidden channel where the good annotation work is posted. The process is public, dull, and repeatable, and almost every stalled application skipped one of five steps. Run the loop once per vendor, then run it again for the next one.

  1. Build a shortlist you are genuinely eligible for. Read the location string, the language requirement, and the task type before the pay band. A Remote (US) card is closed to you outside the United States regardless of how well you write, and a card asking for native Tagalog is closed to you if your Tagalog is conversational. Aim for five to eight cards you could defend applying to, not forty you cannot.
  2. Apply on the employer domain, and keep your own record. Open the role page here for the stable URL, the company name, and the listed band, then use Apply to reach the employer form. Save the date, the vendor, the project name, and the email address you used. Vendors run several brands and portals, and applicants who lose track of which login belongs to which program lose real hours later.
  3. Prepare for the assessment before you open it. Assessments usually have one attempt and a clock that starts when the page loads. Block ninety uninterrupted minutes, read the whole guideline first, keep it open in a second window, and write down the three or four rules the project clearly cares most about. Opening the test to see what it looks like is how most first attempts are wasted.
  4. Write justifications a reviewer could audit. Every judgment needs a reason that points at a specific rule, a specific line in the item, and the consequence. Generic praise such as saying one answer is more helpful scores as noise. Do not paste chatbot prose either: assessors screen for it, and a fluent paragraph with no concrete reference reads as a red flag rather than a strong answer.
  5. Protect the first two weeks, then add a second vendor. Early samples decide how much volume the platform releases to you. Read every QA note, redo flagged items the way the reviewer described rather than the way you would prefer, and ask in the project channel instead of guessing twice. Once your queue is healthy, start a second vendor assessment, because the cure for a paused project is another project.

Notice what is not in the loop. There is no networking stage, no referral trick, and no portfolio review. Vendors are hiring at volume against a rubric, so they replaced the human screen with a test that scales. That is good news if you are a careful reader with no formal background, and bad news if you were hoping a strong CV would carry you past the part where you have to justify a judgment in writing.

The loop is also not linear in practice. Most people who end up with steady hours have three or four applications at different stages at once: one in assessment, one waiting on onboarding, one live and producing, one in a pool waiting for a project to start. If you treat the process as a single application you either wait or panic. Treat it as a pipeline and the pauses stop being emergencies.

Where to find annotation job listings, and where not to

Specialist boards beat generic job search for this niche, and the reason is structural rather than promotional. Annotation work is almost entirely remote contractor work posted by a small number of vendors, so a general aggregator has to guess at a category that does not exist in its taxonomy. City queries are the clearest symptom. Search a metro name plus data annotation and you will mostly get national remote posts with a city glued on for local search traffic. We refuse to clone those doorway pages, which is why you will not find a page here for every city in the country. Use the remote filter instead: remote data annotation jobs.

If your question is really “what companies are hiring for data annotation jobs right now,” the answer is the companies index. On this board the recurring names are Scale AI, Surge AI, Mercor, Micro1, and HireCade. The phrase “data annotation tech companies hiring” usually means that same list plus confusion with the DataAnnotation.tech platform, which is a separate contractor site with a similar name. The difference is untangled in DataAnnotation careers versus this board.

Boards, career pages, and the searches in between

People search for a job board looking for AI evaluator, annotation, or RLHF roles because those three titles rarely sit together anywhere else. That is the product here: one feed with RLHF jobs, LLM evaluator jobs, and AI trainer jobs alongside classic labeling. Employer career pages still matter, because a vendor sometimes opens a program before it reaches any board. Use both. We are not going to pretend a directory replaces the source.

  • Worth your time. A specialist feed you can filter by location and task, the career pages of the five or six vendors that dominate this market, and the project channels you gain access to once you are inside a platform.
  • Occasionally useful. General aggregators, if you search the vendor name rather than the job title and then verify the posting on the company domain before applying.
  • Usually a waste. City-plus-keyword searches, content-farm listicles of “top 50 annotation sites,” and reposted cards with no company name.
  • Never. Recruiters who only exist inside a messaging app, ads that promise a rate with no test, and anything asking you to pay for training, equipment, or a seat.

One honest limit: we will not fabricate requisitions for companies we do not carry. If you searched for a specific agency or brand and this board has nothing, that means we have nothing, not that the role is hidden. How we decide what goes on a card is described in how listings work.

How to triage a listing in sixty seconds

Most wasted applications are wasted before the writing starts. Before you open an apply form, read the card in a fixed order and decide whether you are eligible, whether the pay format is workable, and whether the task type matches something you can actually do well. Sixty seconds of this saves the hour you would otherwise spend on a test you were never allowed to pass.

Read this lineWhat it really tells youAct on it by
Location stringWhether you are eligible at all. Remote (US) means physically in the United States, not willing to work US hoursClosing the tab if it does not match your country, or checking the worldwide guide for global equivalents
Language requirementWhether you are the quality bar or a liability. Native or near-native means your phrasing becomes the dataset standardApplying in the language you write best, not the one with more cards
Task typeHow you will spend six hours: drawing, tagging, rating, or writing. Titles are inconsistent, task descriptions are notMatching it to work you have done before, even in another context
Pay formatHourly means paid idle time and stricter time tracking. Per-task means your reading speed is your wageEstimating minutes per item before you accept a per-task rate
Listed bandRoughly where the role sits: entry review near $12 an hour, general labeling around $18 to $30, specialist and RLHF higherComparing it against the salary guide instead of against your hopes
Assessment mentionWhether the vendor is serious. A real program tests you before it pays youTreating any test-free offer with a high rate as a scam until proven otherwise
Company name and apply URLWhether there is a requisition behind the card at allVerifying the domain matches the named company before entering any personal data
Bands reflect what cards on this board have listed, not offers or averages. Check the live feed for current numbers.

The pay format line deserves more attention than it usually gets. Two dollars an item is excellent at three minutes an item and poor at twelve, and the difference is decided by the guideline rather than by your typing speed. If a card lists a per-task rate, time yourself on the practice items and convert it into an hourly number before you commit. The full conversion method is in the data annotator salary guide, and location patterns are covered in data annotation jobs worldwide.

What to put in the application itself

Vendor application forms are short and mostly administrative. They ask where you live, what languages you work in, what hardware you have, what domains you can evaluate, and how many hours you want. There is usually one free-text box, and that box is the only place a human reads you before the test. Write it like a rubric-following adult rather than a job seeker performing enthusiasm.

  • Lead with eligibility, not motivation. Country, time zone, working languages and level, and whether you have a desktop machine with a stable connection. The screener is checking boxes before they read anything else.
  • Name the domains you can genuinely judge. Nursing documentation, contract law, secondary school maths, Python, or a second language you actually write in. Two real domains beat a list of ten you half know.
  • Show that you can follow someone else’s definition. Any past work where you applied a written standard counts: medical coding, copyediting to a house style, teaching to a marking scheme, QA testing, compliance review, translation to a client glossary.
  • State your availability honestly. If you can only do fifteen hours a week, say fifteen. Overpromising gets you onto a project with volume expectations you will miss in week three.
  • Keep the writing sample plain. Short sentences, concrete nouns, no adjectives doing the work of evidence. If they ask why you want the role, answer in three sentences and move on.

What to leave out: an objective statement, a paragraph about your passion for artificial intelligence, and any claim to a credential you cannot document. Specialist queues verify. Faking a code, medical, or legal background gets accounts removed once your output reaches a reviewer who has the credential for real, and removal is permanent in a way a rejection is not.

If you have never done this work in any form, do the reading before the applying. The on-ramp is how to become a data annotator, and the free AI data annotation foundations course walks through reading a guideline the way a reviewer does. To see where you currently stand, run the skills assessment first.

The assessment, in detail

The assessment is the hiring decision. Everything before it is filtering and everything after it is administration. It usually looks like a guideline document of anywhere from ten to fifty pages, a set of practice items with feedback, then a scored set of twenty to sixty items with a time limit and no feedback at all. Some of those scored items are gold standard items with a known correct answer, seeded to measure you against the people who wrote the rubric.

Three habits separate people who pass from people who do not. First, they read the entire guideline before item one, including the edge case appendix everyone else skips. Second, they slow down on the first five items, because the calibration you set early tends to persist across the whole set and a systematic misreading costs far more than a single wrong answer. Third, they write justifications that a reviewer could check, rather than justifications that merely sound confident.

A weak justification and a strong one

Suppose the item asks you to pick between two model answers about mixing an over-the-counter painkiller with a prescription, and the guideline has a rule saying medical claims need a hedge and a referral to a professional. A weak justification looks like this:

“Response B is better. It is more accurate and more helpful, and it reads better overall.”

That sentence tells a reviewer nothing. It could be pasted under any item in the set. It does not name the rule, the offending text, or the consequence, so it cannot be audited and it cannot be used to train anything. A strong justification for the same item looks like this:

“Response B wins on the medical claims rule. Response A says the combination is safe at any dose, which is an unhedged clinical claim and the guideline lists that as a hard fail even when the underlying information is correct. Response B gives the same dosing detail, marks it as general information, and tells the user to confirm with a pharmacist. B loses slightly on concision, but concision is a tiebreak criterion and the safety rule is not.”

The second version names the rule, quotes the specific failure, weighs the tradeoff, and states which criterion outranks which. It is also obviously written by a person who read the document. Do not outsource this to a chatbot: assessors screen for model prose, and a fluent generic paragraph with no reference to the actual item is now a stronger negative signal than clumsy writing. The longer treatment is in how to pass annotation assessments, with timed practice in the assessment course.

Practical mechanics matter too. Assume one attempt. Assume the timer starts on page load. Do it on a desktop with the guideline open in a second window, not on a phone between other tasks. If the platform allows a flag or comment field for ambiguous items, use it: a well-reasoned flag on a genuinely underspecified item reads better than a confident guess, because guideline authors know which of their items are ambiguous.

Your first two weeks and how QA decides your volume

Passing does not mean the queue opens. Most vendors release a small calibration batch first, sample it heavily, and then decide how much real work you see. That decision is often made by a reviewer who spends four minutes on your output and never speaks to you. Volume arrives as a number changing on a dashboard, and so does its disappearance.

What a reviewer actually flags is narrower than new annotators expect. They are not grading your intelligence. They are grading agreement, measured with inter-annotator agreement and gold items, and the flags cluster into a short list.

  • Rule misses. You applied a defensible standard that is not the standard in the document. This is the most common flag and the easiest to fix, because the reviewer usually names the rule.
  • Drift. Your item four and your item four hundred disagree with each other. Reviewers spot this in a batch even when each item looks fine alone.
  • Thin justifications. The label is right and the explanation is a sentence of filler. On writing and preference projects the explanation is the deliverable, so a thin one is a defect even with a correct choice.
  • Over-flagging. Marking everything ambiguous to avoid being wrong. It reads as an inability to make the call, and it costs the project more time than a wrong answer would.
  • Pace mismatch. Far faster than the batch median usually means skimming. Far slower usually means you are rewriting the guideline in your head instead of applying it.

Treat the first QA note as an instruction rather than an opinion. Redo the flagged items the way the reviewer described even where you still think you were right, then raise the disagreement separately in the project channel with a specific example. Annotators who argue in the rework itself tend to see their volume quietly taper. Annotators who ask one precise clarifying question, in public, where other people can see the answer, tend to be remembered when a better project opens.

A realistic timeline to your first payment

The single most common source of frustration is a timeline mismatch. People apply on Monday expecting to be working on Wednesday. The honest shape of it looks more like the table below, and any stage can stall if a project is between batches.

StageTypical elapsed timeWhat is happeningWhat you should be doing
Application submittedDay 0An automated eligibility screen on location, language, and hardwareApplying to your next shortlisted card rather than waiting
Invitation to assessA few days to two weeksThe vendor batches invitations to match project need, not to match your urgencyReading the guideline the day the invite arrives, not the day it expires
Assessment windowUsually a week to complete, one sitting to doOne timed attempt, scored against gold items and reviewer judgmentBooking ninety uninterrupted minutes on a desktop
ResultDays to a couple of weeksScoring is often manual on writing and preference tasksStarting a second vendor rather than refreshing an inbox
Onboarding and paperworkA few daysContractor agreement, tax form, identity check, sometimes a background checkHaving documents and payment details ready in advance
Calibration batchFirst week of workA small sampled batch that decides your volume tierWorking slowly and reading every QA note in full
Real volumeWeek two to week fourThe queue opens if your sampled quality holdsTracking your effective hourly rate from the first day
First paymentWeekly payout or net-30 invoice, depending on the vendorPayment terms are set by the vendor, not the boardReading the payment terms before you count on the money
Ranges describe patterns applicants commonly report across vendors, not a guarantee from any company.

Two to six weeks from application to first paid task is a reasonable expectation. If you need income next week, this is the wrong plan for next week, and no amount of application volume compresses the vendor side of the timeline.

Ready to apply?

The shortlist is the part you control. Filter the live feed by location and task type, apply on the employer site from the role page, then practise a rubric before your assessment window opens.

Eight reasons you cannot find work as an AI annotator, and the fix

When someone says they have applied everywhere and heard nothing, the diagnosis is almost always one of eight specific things. None of them are about talent. All of them are about mechanics.

What went wrongHow it shows upThe fix
Failed gold questions and never reappliedA rejection after an assessment, then no further attempts, often after the project updated its guidelineCheck the stated cooldown, practise on a rubric, then apply to a different project or vendor
Applied to Remote (US) from another countrySilence, or a rejection minutes after applying, with no assessment invitationFilter to global cards and read the worldwide guide before applying
Hunted expired aggregator postsApply links that 404, forms that never respond, no named company anywhere on the pageApply from a live feed and verify the domain matches the company
Needed full-time hours on day oneQuitting after week one because the calibration batch was smallTreat the ramp as normal and keep other income while the queue opens
Beginner expectations on a specialist rubricRepeated assessment failures on code, medical, legal, or expert evaluation tasksStart on general review and multilingual tracks, then specialise once you have a QA record
One vendor, one application, then waitingWeeks of silence with nothing else in the pipelineKeep three applications at different stages at all times
Justifications that read like a chatbotPassing the label choices but failing the written componentReference the rule and the specific text in the item, in your own plain words
Applying in a language you only speak conversationallyPassing screening, then failing on phrasing and idiom in the writing sectionsApply in the language you write best, and use multilingual cards where you are the scarce supply

There is a ninth pattern that is really a category error. Language data annotation jobs, Fil-English annotator posts, proofreading queues, and AI content writer roles are still annotation work: you are following a language rubric and being sampled on it. They are not a separate industry with separate rules, and treating them as unrelated means missing the tracks where a strong bilingual writer is genuinely scarce. If you want a different sector entirely, we will not invent requisitions for companies we do not list, and any board that does invent them is wasting your assessment attempts.

If the honest answer is that you are aiming too senior, restart at beginner and entry-level data annotation jobs. If you want the market overview before you pick a track, read data annotation jobs in 2026.

Where a CS or specialist background fits

A good place to do AI annotation jobs with a computer science background is not the cheapest global content queue. It is code and reasoning evaluation, STEM grading, agent trajectory review, and the specialist tracks that need someone who can read a failing function and say precisely why the model was wrong rather than that it felt wrong. Those queues sit at the top of the general band because the supply of people who can do them is thin.

The same logic applies to other credentials. A working nurse, a qualified accountant, a practising lawyer, or a translator with a real professional register is scarce in this labor pool in a way that a fast clicker is not. Scarcity is the mechanism that moves rates, not seniority in the conventional sense, which is why an evaluator with a niche skill can list well above a general annotator with more years on a platform.

Set expectations honestly, though. You will still take an unpaid test. You will still be quality-sampled. Hours still fluctuate, and the contract is still a contractor agreement with no notice period. If you were aiming at a staff engineering role, this market will feel like a detour, because it is one. It is a reasonable detour if you want flexible hours, exposure to how frontier models are evaluated, or a route into evaluation design, and a poor one if you want a title and a salary.

The usual doors for technical applicants are Surge AI annotation jobs and Scale AI annotation jobs, with a direct comparison of the two if you can only prepare for one assessment. To practise the exact format, the code and reasoning data course and the RLHF and model evaluation course cover the two most common technical rubrics.

What to do after a rejection

Rejection emails in this market are short and mostly uninformative, which makes people read meaning into them that is not there. It helps to know what the standard phrasings usually mean in practice.

  • “We are not moving forward at this time.” Usually a score below the project threshold. It is a project decision, not a permanent platform ban, and a different project at the same vendor may still be open to you.
  • “We will keep your profile on file.” Often literal. Vendors maintain pools of passed and near-miss candidates and pull from them when a batch lands. It is not a reason to wait, but it is a reason to keep your profile and availability current.
  • “This project is now closed.” The batch filled or the lab paused it. This one is not about you at all, and reapplying to a different program is the correct response.
  • Silence for four weeks or more. Treat it as a no and stop counting on it. Chasing an unanswered application costs more attention than starting a new one.

Then do the diagnostic honestly. Was it eligibility, in which case no amount of preparation would have helped and you should re-filter your shortlist. Was it the rubric, in which case reread the guideline you were given and find the rule you misapplied. Was it the writing, in which case rewrite three of your own justifications to name a rule and a specific line, and compare them to what you originally submitted. Was it timing, in which case the fix is more applications in the pipeline rather than better answers.

Then reapply deliberately. Change the vendor or change the project, respect any stated cooldown, and do not resubmit the same application the following week hoping for a different reviewer. Between attempts, practise against a real rubric rather than reading advice about practising: the course track is built for that, and the glossary covers the vocabulary that appears in guidelines without explanation. If a listing on this board looks wrong or dead, tell us on the contact page and we will pull it.

FAQ: how to find, apply for, and actually get data annotation jobs

How do I find a data annotation job?

Use a specialist board rather than a generic city search, then apply on the employer domain. Open the live jobs feed on DataAnnotationJobs.org, filter to roles that match your location and skill level, read the task description on the card, and use Apply to reach the vendor form. Generic aggregator listings for this niche are frequently expired, duplicated, or lead-capture pages with no real requisition behind them.

Why can I not find work as an AI annotator?

The usual causes are specific and fixable: a failed or rushed assessment you never retook, applications to US-only cards from another country, chasing expired aggregator posts instead of live apply links, expecting full-time hours immediately, or applying to specialist rubrics with beginner experience. Passing the test and holding a quality score is the job. Volume of applications is not what moves the outcome here.

How long does it take to get hired as a data annotator?

Plan for two to six weeks from application to first paid task, and longer if a project is between batches. A typical sequence is a few days to hear back, an assessment window of a week or two, then onboarding paperwork and identity checks, then a small calibration batch before real volume. Some vendors keep passed candidates in a pool until a project needs people.

Where should I apply for data annotation jobs with no experience?

Start with general content review, policy queues, and multilingual labeling rather than specialist tracks. Those cards on this board have listed roughly $12 to $20 an hour and screen on a written test instead of a resume. Skip autonomous-vehicle QA, code reasoning, and expert-domain evaluation for a first application. The beginner guide on this site lists which tracks realistically take first-time applicants.

Do I need a resume or a cover letter?

Most vendors screen with an assessment, so a resume mainly matters for the eligibility and specialist checks: language, location, credential, and equipment. Keep it short and factual, name the domains you can genuinely evaluate, and never claim a code, medical, or legal background you do not have. Fabricated credentials get accounts removed once your work reaches a specialist reviewer.

Can I reapply after failing an assessment?

Often yes, though the waiting period varies and some vendors allow only one attempt per project. Check the rejection email for a stated cooldown, then apply to a different project at the same vendor or a different vendor entirely. Before you spend another attempt, work through a practice rubric so the second try is a change in method rather than a change in luck.

How many vendors should I apply to?

Two or three live qualifications is the practical target, added one at a time rather than all at once. Projects pause without notice, so a single vendor means a single point of failure for your income. Applying everywhere simultaneously is the opposite mistake: you spread attention thin, rush assessments, and burn one-shot attempts across several platforms in the same week.

Where should someone with a CS degree do AI annotation jobs?

Code, reasoning, and STEM evaluation tracks, not the cheapest global review queue. Those projects need people who can read a failing function, judge a proof step, or explain why a model answer is subtly wrong, and they list at the top of the general band. You still take an unpaid test and you still get quality-sampled like everyone else.

Are language, proofreading, and AI content writer roles real annotation jobs?

Yes. If you are following a written rubric to label, rate, correct, or rewrite text, that is annotation work regardless of the job title on the card. Fil-English annotator posts, proofreading queues, and AI content writer roles all sit in the same market and use the same quality machinery: guidelines, gold items, sampled review, and volume that depends on your score.

For the market overview, start at data annotation jobs. If you are searching around childcare rather than around a career move, the same method adapted for that is in how to find work from home jobs as a mom. When you are ready to shortlist, the live feed is open roles.

Read next

More from the blog