Job Details : Software Engineering Expert (W-2 Contract, US Remote)
| π° Salary | $60 β $90 per hour |
| π Location | United States |
| π’ Company | Cincinnatus LLC (staffing partner for leading AI lab) |
| πΌ Category | Software Engineering / AI Research |
| π Remote | Yes (US-based remote) |
| π Contract Type | W-2 Full-time (not freelance) |
| β° Commitment | Approximately 35 hours/week |
| πΈ Payment | W-2 employment (payroll, benefits, compliance) |
About the role:
A leading AI lab is building the next generation of agentic evaluation benchmarks for frontier models and needs experienced software engineers to act as task curators. You will design, implement, and review complex, multi-step engineering tasks that simulate real-world challenges research engineers faceβrealistically hard problems that today’s best AI coding agents cannot yet solve reliably. You’ll work in a tight feedback loop with the lab’s researchers, verifying exactly where and why frontier models fail on your tasks.
What you will do:
- Design realistic, multi-step software engineering challenges that push the limits of AI coding agents
- Build reference solutions in Python with setup and checks for verifiable answers
- Use AI coding assistants in your workflow and observe where they help and where they fall short
- Review tasks from fellow experts and provide feedback on clarity, correctness, and difficulty
- Analyze how AI agents attempted your tasks and help researchers understand failure points
What you need:
- MSc or PhD in Computer Science or another STEM field, or equivalent practical experience in a research-heavy domain
- 1+ years of experience in research, research-engineering, or software engineering
- Strong hands-on Python scripting and debugging skills with clean-code habits
- Everyday fluency with version control (Git), IDEs, and standard software development workflows
- A perfectionist mindset: high attention to detail, creativity in task design, and strong written communication
- Ability to reliably engage for approximately 35 hours per week
Preferred (not required):
- Experience with AI coding assistants, prompt engineering, or agent workflows
- Past experience in AI training, model evaluation, or benchmark/task authoring
Important notes:
- πΊπΈ US location required β fully remote within the United States
- π W-2 contract β not freelance or project-based
- β° Approximately 35 hours/week β structured role
- π’ Employer of record: Cincinnatus LLC
- π§ Python + Git required
- π― Research engineering background β ideal for those with a perfectionist mindset
Why this job is worth your time:
$60-90/hour to design cutting-edge AI evaluation benchmarks, with W-2 benefits and the prestige of working with a leading AI lab.
