Cat and Mouse: How Mercor Audits You, Rates You, and Lets You Go Without a Word
Audit sampling, the model-usage police, reviewer roulette, and the anatomy of a silent offboarding. Plus the energy system that keeps a day-job contractor sane inside all four.
Spend ten minutes on r/mercor_ai and you will find the same story told in different fonts. A thread called Gut Punch. A thread called In Need of Some Encouragement. A thread about being offboarded for no reason, where the comment section lands on the sentence that haunts this whole platform: there is always a reason. You just don’t get to hear it.
The accounts were in good standing. The queues went quiet. The email, when there was one, explained nothing. And a room full of strangers is left doing forensic analysis on their own work history.
I’ve onboarded onto more than ten Mercor projects, and I have read the queue from the reviewer side too. Episode 2 mapped the corridor: the band of acceptable speed and quality that metrics decide you’re inside. Episode 3 is what happens when someone decides you left it. The free half of this piece is how the enforcement machine actually works: audits, model-usage policing, reviewer roulette, silent offboarding. The paid half is the part nobody expects: an energy management system. Because once you understand the machine, you’ll see why the contractors who last aren’t the most vigilant. They’re the least depleted.
The four rooms of the machine
1. Audits are a sample, not a census
Every mature project runs a quality apparatus: daily audit dashboards, a quality matrix that grades dimensions of your work, an error lookup tool where your mistakes get catalogued by type. What new contractors miss is that nobody reads everything. Thousands of tasks flow through these projects daily; the audit is a sample.
So the real question is what raises your sampling odds. Being new does. Sitting at either wall of the speed corridor does, because outliers get pulled first. A reviewer flag does. A client spot check landing on your batch does, and that one is pure lottery. Meanwhile the monitoring layer never sleeps: the timer apps log active windows and idle gaps, idle time gets reviewed and deducted weekly, and charging time outside active work is written into guideline docs as a termination offense.
Here is the uncomfortable implication: audits don’t see your average. They see the three tasks that got pulled. Your representative work, as far as the machine is concerned, is whatever you produced in your most tired hour. Hold that thought; it’s the hinge of this whole episode.
2. The model-usage police
Projects that exist to train AI models mostly forbid you from using AI models to do the work, or license it so narrowly that the safe assumption is no. The irony is noted, usually in the same breath as the policy tab.
Enforcement is a stack, not a single tripwire. The monitoring software sees your open windows. The task tools can see paste events and timing patterns. Your prose gets compared against the baseline you set in onboarding quizzes, some of which are one attempt, graded by the client’s own team, with no second chances and no appeal. And somewhere in the pipeline, AI detectors run on work submitted to a company that trains AI. False positives exist. Honest contractors absorb the friction created by whoever botted the assessments last quarter, which is Episode 1’s arms race wearing a badge.
The one free tip I’ll give above the paywall: when a usage rule is ambiguous, ask once, in the public channel, and work to the written answer. That message is both a clarification and a shield.
3. Reviewer roulette
Between your submission and your pay sits another contractor. Which one you draw matters. Which version of the rubric they read matters, on projects where the guideline doc changed twice in one week. What their own metrics look like that day matters most of all, because reviewers work under quotas and error budgets, and reviewers get audited too. The machine audits its auditors, which makes strictness swing in random directions.




