PlatformShanghaiFull-timePosted
Evaluation & Self-Improvement Engineer
Own the method and implementation behind night school exams, judge scoring and coaching, so every improvement is measured against a baseline you can trust.
Apply for this rolePut the role title in the subject line and attach your CV.
What you will do
- Design profession-specific eval sets and rubrics, and maintain the judge model prompts and calibration process.
- Build and improve coaching: generate candidates on a snapshot, re-score, and hand only improvements that beat the baseline to the user for approval.
- Set up regression gates and drift monitoring to answer whether it is getting worse over time.
- Turn eval results into reports and diff explanations that users can read.
What we look for
- Real experience evaluating LLMs or building datasets; you know how judge models self-confirm and how to prevent it.
- Fluent Python, able to read and change backend code; enough statistics to tell a real gain from noise.
- A point of view on bounded self-improvement, and the will to turn literature findings into engineering constraints.
- Clear writing, for teammates and for users.
Nice to have
- Published work or open-source eval tooling.
- Reinforcement learning, program synthesis or automatic prompt optimization.
How to apply
Send your CV and a short note on how an agent should avoid getting worse as it changes to the careers address, with the role title as the subject. We reply within two working days.