Blog · Companies

Scale AI annotation jobs: Scale AI data annotation, AV, and polygons

Published Updated 12 min read

Scale AI data annotation jobs and Scale AI annotation jobs are the same hiring funnel with two search phrasings. This page is the worker-side map: what Scale actually posts, how automotive and polygon work fits, what the tooling feels like at hour four, where language and evaluation projects sit, and how all of that differs from Surge. The live list is open data annotation jobs, the company hub is Scale AI jobs, and the official site is scale.com. We are not Scale.

What Scale AI annotation jobs usually are

Scale is the large data platform: perception for vehicles and robotics, government programs, and language data for models. That breadth is the single most useful fact about applying there. A Scale AI data annotation job might be remote vision-and-language labeling at a mid hourly band, or senior on-site autonomous-vehicle quality assurance at a much higher listed rate, and those two cards have almost nothing in common except the company name. The title and the task description on the card matter more than the brand.

That is different from a specialist vendor where every queue looks roughly the same. It means the question “what is it like to work for Scale” has no single answer, and it means two people can both be correct while describing completely different experiences: one drawing cuboids in a point cloud with a stopwatch running, the other reading two model answers and writing a justification. Browse the current rows on the jobs feed and read them as separate jobs that happen to share an employer.

Worker-facing brands under the Scale umbrella have changed names over the years. If an old forum post says Remotasks or Outlier, that is history, not a second careers site you must find. Apply through the URL we publish for that opening, and be suspicious of any site claiming to be a Scale worker portal that wants a fee, a deposit, or your banking details before there is a signed agreement with a named entity. Our side of the arrangement is explained in how our listings work.

The program mix and where you fit

The tracks below are the shapes we see, drawn from cards this board has carried plus what Scale describes publicly. Not all of them appear as open contractor listings at any given moment, and the mix moves with client demand.

TrackWhat the work isUsual shape
Perception and sensor dataBoxes, polygons, cuboids, and object tracking across camera and point-cloud framesRemote or on-site depending on the data, with spatial skill weighted heavily
Autonomous vehicle QAReviewing and correcting other people’s perception labels against a driving guidelineSenior, slower, sometimes on-site, the highest listed band we have carried
Language and NLP dataEntities, classification, translation review, and multilingual labelingUsually remote, often the widest door for a first Scale project
Evaluation and preference dataRating and comparing model output against a rubric, sometimes writing a better answerRemote, writing-heavy, overlaps what people call general AI annotator work
Public-sector programsWork Scale describes publicly around government and defense-adjacent AIRarely visible as an open contractor card, and the source of on-site rules when it is
Track descriptions come from listings on this board and public company positioning, not from an internal Scale org chart.

Picking a track honestly is the highest-leverage thing you do before applying. Perception and language assessments test unrelated skills, and a strong writer who applies to a cuboid track fails on tooling precision while concluding, wrongly, that Scale rejected them as a candidate. If you are unsure which side of that line you sit on, the free skills assessment gives you a rough read in a few minutes, and AI annotation jobs explains the four hiring tracks in the market generally.

Automotive annotation and a polygon annotation career

Searches for an AI annotator for automotive jobs and for a polygon annotation career are pointing at the same thing: perception data. Draw a tight polygon around a pedestrian, cuboid a vehicle, track it through occlusion, mark weather and lighting. That is slower, more spatial, and closer to bounding box work than to writing a chatbot critique, and it rewards a completely different temperament.

The craft is worth describing concretely, because job ads never do. Tight does not mean small: it means the boundary follows the silhouette rather than a convenient hull, so a cyclist gets the gap between the arm and the torso, and a parked car gets the wheel wells rather than a rectangle glued to the tarmac. The shadow is not part of the object. The wing mirror usually is. A pole crossing a pedestrian splits the visible region into two pieces, and the guideline, not your instinct, decides whether that is one instance or two.

What tight polygon and cuboid work demands

  • Occlusion and extent. When the back half of a truck is hidden, you still estimate the full cuboid. Reviewers check whether your estimate is plausible against the visible wheels and the road plane, and a cuboid that shrinks as the object hides is a classic flag.
  • Orientation. A cuboid with the right footprint and the wrong heading is wrong. Yaw errors on parked vehicles sitting at an angle to the curb are one of the most common corrections in review.
  • Track identity. An object that leaves the frame and comes back should keep the same identity if the guideline says so. Identity swaps between two similar cars passing each other are expensive downstream and easy to create by accident.
  • Truncation at the frame edge. Half a bus at the image border has its own rule, and it is usually different from the rule for an occluded object in the middle of the scene.
  • Attributes. Night, rain, glare, wet road, and heavy-versus-light occlusion tags exist so a model can be evaluated by condition. They feel like busywork and they are graded like labels.
  • Ambiguous classes. Is a delivery robot a vehicle. Is a trailer separate from the tractor. Is a group of twelve people one crowd region or twelve instances. The document decided months ago, and consistency with that decision is the whole job.

What the annotation tooling feels like at hour four

The tooling shapes the day as much as the rules do. You work in a browser with hotkeys you will eventually use without thinking, interpolation between keyframes that saves an hour and quietly introduces drift in the middle frames if you never check them, and point-cloud panes beside camera views where a sparse object at forty metres is a dozen returns you have to interpret. Six hours of that without your boundaries loosening is the actual skill. People who cannot sit with it discover the fact in week two, not in the assessment.

Some of that queue is remote. The highest listed specialist quality assurance role we have carried for Scale was on-site in San Francisco, not a worldwide laptop gig. If you specifically need data annotation jobs worldwide, do not assume every Scale AI annotation job is global. Autonomous vehicle annotation jobs also tend to sit at the senior end, so the realistic route in is production labeling first and review later.

Language, RLHF, and general AI annotator work at Scale

Scale also runs language and evaluation projects. Those overlap what people call general AI annotator jobs or AI data annotator jobs: rate answers, follow a rubric, survive QA. The daily texture is the opposite of perception work. Nothing is pixel-precise and everything is arguable, so the graded artefact becomes your written justification rather than your boundary.

Switching between perception and language work

If you have only ever done vision work, the adjustment is real. You will be asked to say why one answer beats another in a way a stranger could audit, and “it reads better” is not an answer. If you have only ever done language work, the adjustment in the other direction is also real, because a guideline that felt like a suggestion in a preference task becomes a measurement standard when the output is a polygon. Both jobs are scored with inter-annotator agreement against a gold standard set, which is the thread connecting them.

For a writing-heavy preference queue specifically, Surge AI annotation jobs may fit better. For a mixed platform where vision and language sit side by side, Scale is the default search. If you want the concept before the hiring language, start at what RLHF is and the RLHF jobs feed, or the narrower LLM evaluator jobs list.

Remote versus on-site, and what that means for eligibility

Searchers looking for Scale AI jobs remote are asking a fair question with an annoying answer: it depends on the data. Language and evaluation projects travel well, so those cards are usually remote. Some perception and public-sector data does not leave a controlled environment, which is why a senior quality assurance role can be on-site even at a company built on distributed work.

Location lineWhat it usually meansWhat to verify
Remote (Global)Open beyond the US, subject to whatever the individual project requiresThat a payout method you can actually receive is available where you live
Remote (US)Closed to you unless you are physically in the United StatesRight to work and the tax form the vendor expects at onboarding
On-siteThe data or the program requires a controlled location, historically San Francisco on the cards we carriedWhether relocation is actually on the table before you invest in the process
Hybrid or unspecifiedThe requirement exists and the card is vague about itThe eligibility question on the employer form, which is the authority
Patterns observed on listings, not Scale policy. The employer application is the only reliable source on eligibility for a given project.

Read the location line before the pay line, every time. A US-locked program carrying data residency or public-sector requirements often lists higher than a global sibling running a similar rubric, and that gap is compliance pricing rather than a comment on your ability. There is no version of a strong application that overcomes a geography mismatch, so filter first and apply second. The narrower feed is remote data annotation jobs.

How hiring works, step by step

Assessment first, hours later. That order is the thing most people get backwards, and it explains why passing feels anticlimactic: you have joined a qualified pool rather than started a job.

  1. Read the location line before the title. Scale cards on this board have ranged from remote labeling to senior on-site quality assurance in San Francisco. Decide whether you are eligible for the location before you get attached to the rate, because no writing sample fixes a geography mismatch.
  2. Match the track to a skill you can prove. Perception work rewards spatial patience with polygons and cuboids. Language and evaluation work rewards precise writing. Applying to the wrong track is the quiet reason a strong candidate fails an assessment that had nothing to do with their actual strength.
  3. Apply through the employer link on the card. Open the role page on DataAnnotationJobs.org for the stable URL, the company name, and the listed band, then follow the apply link into Scale’s own hiring or contractor flow. Old worker-facing brand names from forum threads are not a second door to hunt for.
  4. Treat the assessment or calibration task as production. Read the whole guideline before your first item and assume gold questions are mixed into the set. On perception tasks the graded thing is boundary precision and label consistency, not how many frames you got through.
  5. Protect your quality score through the first batches. Early samples decide how much volume you see later. Redo flagged items exactly as the reviewer described, and raise ambiguous cases in the project channel rather than guessing twice and dragging your agreement rate down.
  6. Hold a second vendor qualification. Perception and language programs both end when a client changes direction. Pass an assessment somewhere else while your Scale queue is healthy, so a paused project turns a zero week into a slow week.

Step four deserves extra attention on perception tracks, because the graded qualities are unintuitive. Reviewers look at boundary precision, label consistency across similar objects, correct use of attributes, and whether you followed the occlusion rule rather than your eye. Throughput is a later conversation. On language tracks the same step means quoting the guideline in your reasoning instead of asserting a preference. The longer version is how to pass annotation assessments, with timed practice in the assessment course and the fundamentals in the free AI data annotation foundations course.

Ready to apply?

Scale cards move on and off the board as programs open and close, and a remote language listing and an on-site QA listing are different jobs. Filter the live feed, then apply on the employer site from the role page.

Pay bands by track

Pay bands for Scale cards on this board have included roughly $18 to $30 an hour for remote vision and language work, and much higher for senior autonomous-vehicle quality assurance. Always trust the live card over any number in an article, including this one.

TrackWhat cards have listedWhy the number sits there
Remote vision and languageRoughly $18 to $30 an hourThe widest applicant pool and the most transferable skills
Senior on-site AV QAMaterially higher than the remote band, the top Scale card we have carriedSeniority, an on-site constraint, and responsibility for other people’s output
Evaluation and preference workClusters with the remote band, project by projectWriting quality is the product, so the rate follows the rubric depth
Entry content reviewNearer the $12 floor we see across the whole boardLeast scarce skill, highest churn, shortest onboarding
Bands reflect what job cards on this board have listed, not offers, averages, or a salary survey. Check the live feed for current numbers.

Three things move a rate more than seniority does: scarcity of the skill, location restrictions, and how long a single item takes. On per-item perception projects that last point is the one that bites, because a rate that looks generous per frame can be poor once you include the frames where interpolation drifted and you fixed twelve of them by hand. Track your effective hourly number through the first week. The cross-vendor picture is in the data annotator salary guide, and hours stay variable on contractor work whatever the band says.

Scale AI versus Surge AI, and when to pick which

The short version: Scale is breadth, Surge is a writing filter. Scale gives you more ways in, including tracks where spatial skill counts for more than prose. Surge concentrates on language, reasoning, code, and preference quality, so its assessment leans hard on how you justify a judgment.

If youLeanBecause
Have spatial patience and a real machineScalePerception and polygon work exists at volume and rewards precision over prose
Write precisely and enjoy defending a callSurgePreference and reasoning queues grade the written justification itself
Want the widest set of possible first projectsScaleOne employer, several unrelated tracks, so a rejection on one is not a rejection overall
Hold a professional credential to sellMercor or Micro1Marketplace matching pays for scarcity rather than throughput
Just want a straightforward remote contractHireCadeContract AI training and evaluation staffing with a simple route in

The side-by-side detail lives in Scale AI versus Surge AI, and the current roster is the companies index with hubs for Surge AI, Mercor, Micro1, and HireCade. If you searched Meta AI annotator roles, Wolfestone, Prodigy, or other names we do not currently list, we will not invent those requisitions. Use the companies index for who is on the board today.

What people get wrong about applying to Scale

The recurring mistakes are specific, and most of them happen before anyone grades your work.

  • Treating the brand as one job. Scale contractor jobs span perception, language, and evaluation. Applying to whatever is newest instead of whatever matches your skill is how good candidates collect unrelated rejections.
  • Assuming every role is remote. The senior quality assurance card we carried was on-site. Reading the pay line first and the location line second wastes more time than any other single error here.
  • Hunting for a legacy worker portal. Old brand names in forum threads send people looking for a door that is not there, and straight into whatever lookalike site ranks for it.
  • Racing the clock on a perception assessment. Boundary precision and consistent labels are what get scored. Frames per hour matters once you are in production, not before.
  • Trusting interpolation. Keyframing an object across forty frames and never checking the middle is the fastest way to deliver drift that a reviewer will catch and you will redo unpaid.
  • Skimming the guideline. Occlusion, truncation, and class edge cases are all decided in the document. Guessing them consistently is still worse than reading them once.
  • Expecting steady hours. Programs pause when a client changes direction, so a single vendor is a single point of failure for your income rather than a job.

If you have applied and heard nothing, the diagnosis is usually one of a short list of fixable things in how to get data annotation jobs. If the specialist perception cards are out of reach for now, start with beginner and entry-level roles, build a quality record, and read how to become a data annotator for the full on-ramp. For the market view across every vendor, the long read is data annotation jobs in 2026.

FAQ: Scale AI annotation jobs, automotive work, and polygon careers

How do I get Scale AI data annotation jobs?

Open the Scale listings on this board, apply through the employer link into Scale’s own career or contractor flow, then pass their assessment and hold your quality score once work starts. DataAnnotationJobs.org does not hire for Scale and cannot pass you through an assessment. We publish the opening, the listed band, and the location line, and Scale runs the test, the contract, and the payments.

Does Scale AI hire for automotive and polygon annotation?

Scale’s public mix includes vision and autonomous-vehicle data, which is where bounding boxes, polygons, cuboids, and specialist quality assurance show up. Not every Scale AI annotation job is automotive, though. Language, evaluation, and general annotation projects sit alongside perception work, so read the task description on the card rather than assuming the brand means vehicles.

Are Scale AI annotation jobs remote?

Many are, but not all. Remote vision and language cards are the common shape on this board, while the highest listed specialist quality assurance role we carried for Scale was on-site in San Francisco rather than a worldwide laptop job. Some perception data cannot leave a controlled environment, so treat every location string as a real constraint.

How much do Scale AI annotation jobs pay?

Remote vision and language cards on this board have listed roughly $18 to $30 an hour, and senior on-site autonomous-vehicle quality assurance has listed much higher. Entry-level content review across the market floors nearer $12. Those are numbers printed on job cards rather than a salary survey, and hours on contractor work are rarely guaranteed.

What is Remotasks or Outlier, and should I apply there?

Worker-facing brand names under the Scale umbrella have changed over the years, so an old forum post naming one of them is history rather than a second careers site you need to track down. Apply through the URL published on the opening you are actually looking at. If a site claims to be a Scale worker portal and asks you for money, it is not one.

Is polygon annotation a real career?

It is a real entry point and a narrow ceiling on its own. People who stay in perception data usually move from producing labels to reviewing them, then into quality management or guideline work, because that is where judgment beats throughput. Treating a polygon annotation career as a step into reviewer and quality roles is the version that compounds.

Do you need experience to apply for Scale contractor jobs?

General annotation and review queues screen with a test rather than a resume, so no degree is required. Specialist tracks are stricter: sensor fusion, code, and expert-domain projects genuinely check the skill. Spatial patience and a real computer matter more than credentials for perception work, since three-dimensional and video tools are not usable on a phone.

Does Scale AI hire outside the United States?

Remote global cards do appear, but Scale runs programs with data residency and public-sector constraints, so some queues are US-only or on-site by design. Read each location line rather than generalising from one listing, and check that a payout method you can use is available in your country before you invest an assessment attempt.

The official Scale site is scale.com, the roles we track sit on the Scale AI company page, and the vocabulary you will meet in a perception guideline is in the glossary. A Scale AI annotation job is filled by Scale, not by us.

Read next

More from the blog