createrole

PlatformShanghaiFull-timePosted

Evaluation & Self-Improvement Engineer

Own the method and implementation behind night school exams, judge scoring and coaching, so every improvement is measured against a baseline you can trust.

Apply for this rolePut the role title in the subject line and attach your CV.

What you will do

  • Design profession-specific eval sets and rubrics, and maintain the judge model prompts and calibration process.
  • Build and improve coaching: generate candidates on a snapshot, re-score, and hand only improvements that beat the baseline to the user for approval.
  • Set up regression gates and drift monitoring to answer whether it is getting worse over time.
  • Turn eval results into reports and diff explanations that users can read.

What we look for

  • Real experience evaluating LLMs or building datasets; you know how judge models self-confirm and how to prevent it.
  • Fluent Python, able to read and change backend code; enough statistics to tell a real gain from noise.
  • A point of view on bounded self-improvement, and the will to turn literature findings into engineering constraints.
  • Clear writing, for teammates and for users.

Nice to have

  • Published work or open-source eval tooling.
  • Reinforcement learning, program synthesis or automatic prompt optimization.

How to apply

Send your CV and a short note on how an agent should avoid getting worse as it changes to the careers address, with the role title as the subject. We reply within two working days.

Evaluation & Self-Improvement Engineer

Put the role title in the subject line and attach your CV.

Apply for this role