Senior Software Engineer - AI Evaluation / Coding Agents Python / TypeScript / JavaScript / Golang
Location: USA Approved Regions
Job Type: Contract Part-Time
Work Arrangement: Remote
Contract Duration: 3 Months
About the Role
We're looking for experienced, hands-on software engineers to help evaluate and improve AI coding models.
Rather than primarily building production applications, you'll work with coding agents across real-world repositories and assess the quality of their work. You'll review generated code and agent behavior, determine whether solutions are technically correct, identify failure modes, and create the evaluation signals and feedback used to improve model performance.
Think of the coding agent as another engineer whose work you're reviewing: Can it understand the task? Did it choose the right approach? Is the resulting code correct, robust, and maintainable? Can you explain precisely where it succeeded or failed?
What You'll Do
What We're Looking For
Experience with LLM evaluation, coding agents, RLHF, preference data, rubric design, or post-training is a plus, but not required.
Engagement Details
Evaluation Process
The practical exercise focuses on your ability to review and evaluate AI-generated code, not competitive programming or algorithm puzzles.