AI eval & model-quality engineer jobs

Model-quality roles — eval harnesses, LLM-as-judge, observability (LangSmith, Arize).

228 open roles

Open roles
228
Median salary
$249k

55% publish a range

Remote
12%

25% hybrid

Top skill
Eval harnesses

in 74% of roles

Hiring most: OpenAI 20 · Anthropic 11 · LangChain 9 · Anaplan 6 · Drata 6 · Okta 6

Common stack: Eval harnesses 169 · Python 150 · Orchestration 82 · RAG 78 · Go 64 · MCP 63 · Tool use 62 · AWS 61 · TypeScript 56 · LangSmith 55

← All jobs RSS feed

Filter these 228 roles Filter 228