← All jobs

AI Quality Engineer

Xapien

Salary
Not published
Location
Xapien London
Work type
Hybrid
Posted
today

Apply on company site (opens in new tab)

Eval harnessesLLM-as-judgeCI/CDPython

About Us

Every business needs to know who they’re really working with whether it’s suppliers, investors, partners, or third parties. At Xapien, we’re reinventing how organisations do that, combining speed, scale and accuracy with cutting-edge AI.

Since 2018, we’ve grown from a deep-tech startup to a global player in AI-driven due diligence and risk intelligence. 2024 was a landmark year as we closed a $10M Series A, earned recognition in the Chartis RiskTech100® and Everest Group’s Leading 50™, and expanded our products, markets and customer base.

2025 was even bigger: With new regulations, rising compliance pressure and growing reputation risks, organisations everywhere are demanding smarter, faster ways to work and they’re turning to us. Customers worldwide from global law firms and private banks to universities and nonprofits rely on Xapien to turn days of manual research into trusted insights, delivered in minutes.

Our momentum so far

Customers in a diverse range of industries, with particular growth in wealth management, financial services and supply chain onboarding.

Demand accelerating beyond our UK headquarters, across multiple continents with new customers in the Middle East, Asia, and Oceania.

Why this is an exciting time to join

This isn’t just another job role; it’s an opportunity to shape the future of the due diligence industry with a market-leading product trusted by global organisations. Whether you’re in marketing, communications, sales, or finance, you’ll play a critical role in driving growth, building credibility, and defining how we connect our product to a rapidly evolving market.

We’re scaling fast with more customers, more releases, and bigger regulatory and market challenges. Expectations are rising, due diligence now means real-time insight, delivered efficiently and with impact. You won’t just be executing campaigns, closing deals, or managing numbers, you’ll be influencing strategy, optimising processes, and helping position a brand that’s setting the new standard for trust and transparency.

For people who love learning, innovating, and making an impact that matters now is the moment to be a part of Xapien.

The Role

You'll work across our evaluation systems and test infrastructure — the machinery that lets us ship non-deterministic AI output with confidence. Evaluate, automate, dig, improve.

  • Enhance and run our layered eval harnesses — automated checks, LLM-as-judge scoring, rubrics, targeted human review
  • Maintain versioned golden datasets so evaluation stays reproducible and auditable
  • Track the signals that matter for investigative output: groundedness, hallucination rate, entity resolution, source quality
  • Run and improve a tiered CI model built on Python, Playwright and pytest — six authenticated personas, three browsers, real multi-tenant coverage
  • Dig into flaky tests and wobbling metrics and fix them at the root

We use Claude Code by default — selector sync, failure triage and eval scoring are AI-assisted end to end. The craft is knowing when a green tick is lying.

What we're looking for

Two doors here. You're a proven senior who can do most of this today — or you have the foundations and the drive, and we'll grow you into the rest. Both doors are real.

  • Strong test-automation engineering. Python (or comparable) with a modern stack — Playwright, pytest or similar — and fluent in CI/CD. Tests as infrastructure, not scripts.
  • Analytically sceptical and numerate. You question baselines and never take a green tick at face value.
  • You make complex things simple. The same wobbling metric explained to an engineer, a PM and a customer-facing lead — and understood by all three.
  • A collaborator. Quality cuts across squads you don't manage — you build trust with the engineers whose work you're testing.
  • A fast learner with proof. You've gone novice-to-competent in something genuinely hard, under real delivery pressure.
  • Genuinely interested in AI quality. You can articulate why testing non-deterministic output is a different — and more interesting — problem.
  • (Senior door) You've shipped LLM evals — rubric design, LLM-as-judge, regression harnesses — and raised the quality bar across teams that didn't report to you.
  • Good to work with. Small team — it matters.

Here’s our promise to you:

  • We are going to work with you - to build a rewarding and fulfilling career with the opportunities, challenges and resources you need to do you your best work.
  • We succeed together- you will own a meaningful part of the business through our employee shares & equity programme.
  • Private health insurance to keep you in tip-top condition.
  • Life Insurance – let's hope nobody ever needs this!
  • Unlimited holidays – yes, it really is uncapped – take the time you need, when you need it.
  • Everyone is learning, developing and challenging themselves so we have a £1k professional development fund per year if you want to learn new skills, even new things outside of work.
  • Most importantly of all, we will work with you, to help you realise your fullest potential, always.

Apply on company site (opens in new tab)