AI eval & model-quality engineer jobs

Model-quality roles — eval harnesses, LLM-as-judge, observability (LangSmith, Arize).

432 open roles

Open roles
432
Median salary
$238k

median of 245 priced roles, USD

Remote
13%

plus 22% hybrid

Top skill
Eval harnesses

in 79% of roles

Hiring most:LangChain19Anthropic14OpenAI14ServiceNow12Mercor9Cisco8

Common stack:Eval harnesses340Python261Orchestration175Go157Tool use135AWS114RAG107MCP100TypeScript100Kubernetes92

← All jobsRSS feed