Job Overview : QA/Test Engineer (AI Benchmark Quality, W-2 Contract, US Remote)
| π° Salary | $60β$90 per hour |
| π Location | United States |
| π’ Company | Cincinnatus LLC (staffing partner for leading AI lab) |
| πΌ Category | Software Development / QA |
| π Remote | Yes (US-based remote) |
| π Contract Type | W-2 Full-time (not freelance) |
| β° Commitment | Approximately 35 hours/week |
| πΈ Payment | W-2 employment (payroll, benefits, compliance) |
About the role:
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models. Complex multi-step tasks are only useful if they are airtight: unambiguous, correctly graded, and robust to shortcuts. We are seeking experienced QA and test engineers to be the quality backbone of this benchmark β designing the test cases and review processes that guarantee every task measures what it claims to measure.
Each task under review represents one to two days of expert effort and spans multiple technical skills, so quality review here means genuinely understanding the task: running it, probing its edge cases, and debugging its environment. You will work in a tight feedback loop with the lab’s researchers and task authors.
What you will do:
- Design checks: Create test cases that confirm each task works as intended β including the tricky edge cases
- Review tasks: Give tasks and reference solutions a careful read before they’re finalized, catching ambiguity and gaps early
- Debug: Roll up your sleeves in Python when a task or its checks don’t behave the way they should
- Shape the process: Help build simple, repeatable quality checklists, and share feedback authors can act on right away
- Protect the results: Watch for shortcuts and grading gaps in AI agent runs so benchmark scores stay trustworthy
What you need:
- MSc or PhD in a STEM field, or equivalent practical experience in a research-heavy or engineering-heavy domain
- 1+ years of experience in test engineering, quality assurance, or a research/software engineering role with strong quality ownership
- Demonstrated skill designing test cases and quality-review processes, and debugging complex systems end-to-end
- Working proficiency in Python and Git, and comfort navigating unfamiliar codebases and environments
- Exceptional attention to detail and clear written documentation habits
- A perfectionist mindset: creativity in finding what others missed, and the ability to work independently through ambiguous, open-ended problems
Preferred (not required):
- Past experience in AI training, model evaluation, or quality review of AI-generated work
Important notes:
- πΊπΈ US location required β fully remote within the United States
- π W-2 contract β not freelance or project-based
- β° Approximately 35 hours/week β structured role
- π’ Employer of record: Cincinnatus LLC
- π§ Python + Git required
- π― Perfectionist mindset β attention to detail is critical
Why this job is worth your time:
$60-90/hour to ensure the quality of cutting-edge AI evaluation benchmarks, with W-2 benefits and the prestige of working with a leading AI lab.
