sanitation — the part that doesn’t get said enough here: this was structurally inevitable. Piece-rate annotation contracts reward volume, not quality. If you’re getting paid $0.08 per labeled example and GPT-4 can produce 50 in the time it takes you to write 5, the rational move is obvious. It’s less ‘ironic lesson’ and more ‘predictable tragedy of misaligned incentives.’ The fix probably isn’t better monitoring — it’s redesigning the task so high-quality human signal is the only thing that clears the bar (adversarial examples, edge-case disambiguation, stuff LLMs genuinely can’t fake). We built some thinking around data-quality verification for exactly this kind of pipeline, more context at https://cxgo.ai/l/5LGqWHi if it’s useful for anything you’re working on.
sanitation — the part that doesn’t get said enough here: this was structurally inevitable. Piece-rate annotation contracts reward volume, not quality. If you’re getting paid $0.08 per labeled example and GPT-4 can produce 50 in the time it takes you to write 5, the rational move is obvious. It’s less ‘ironic lesson’ and more ‘predictable tragedy of misaligned incentives.’ The fix probably isn’t better monitoring — it’s redesigning the task so high-quality human signal is the only thing that clears the bar (adversarial examples, edge-case disambiguation, stuff LLMs genuinely can’t fake). We built some thinking around data-quality verification for exactly this kind of pipeline, more context at https://cxgo.ai/l/5LGqWHi if it’s useful for anything you’re working on.