Senior ML Researcher
- Salary
- Not published
- Location
- European Union
- Work type
- On-site
- Level
- Senior
- Posted
- today
- Verified live
- today
Apply on company site (opens in new tab)
Filed underLLM EngineerAI AgentsEvals & QualityInference / ServingFine-tuningCore ML
703 of 1517 AI Agents roles on this board publish pay; their median is $220k.
About Toloka
At Toloka AI we create data that powers leading GenAI models and innovations. We work with frontier labs, big tech, renowned AI startups, enterprises and non-profit research organizations worldwide. We use a combination of Experts + Crowd + Tech Platform to teach AI models to reason and evaluate their efficacy and safety. We have experts in more than 50 different domains—from doctors and lawyers to physicists and engineers—and boast one of the most diverse global crowds, representing ove r 100 countries and speaking 40+ languages. We are a well-funded startup with an enviable portfolio of clients including Anthropic, Amazon, Microsoft, Poolside, Recraft, and Shopify.
Recently, we secured strategic investment led by Bezos Expeditions and Nebius Group with participation from Mikhail Parakhin, CTO of Shopify and board advisor to leading GenAI companies, who now serves as our Chairman of the Board. Our remote-first team is globally distributed around the world: USA, UK, the Netherlands, Serbia, and more.
About the Team
We are the ML team inside Toloka — we build the machine-learning products that power the platform itself, so every project running on Toloka is faster, cheaper, and more reliable.
A few examples of what we own:
- LLM QA — the core technology behind Toloka's automated quality-check mechanism. Every annotation flowing through Self-Service is reviewed by an LLM agent we design, train, and operate.
- Model distillation and fine-tuning — adapting frontier and open-source models to Toloka's tasks to hit the right quality at the right cost.
- Evaluation, benchmarking, cost modeling, and model selection across providers.
We own the full chain. The same team designs the ML solution, ships it to production, keeps it running 24/7, analyzes the results coming back from real projects, and feeds that signal into the next iteration. No hand-off between research, engineering, and operations — it's all us.
About the Position
You will own Toloka’s end-to-end fine-tuning, RL, and evaluation stack, bridging applied research and product engineering. In this role, you will spearhead greenfield post-training initiatives (such as GRPO and reward modeling), transform complex ML experiments into scalable platform features, and occasionally author technical write-ups on your findings for the AI community.
What you’ll do
- Own end-to-end fine-tuning pipelines: data prep, SFT/LoRA training, distillation from frontier models to smaller ones, evaluation, and serving handoff — both as self-serve platform capabilities and in hands-on client engagements.
- Extend our post-training stack beyond SFT into RL (RFT/GRPO-style methods, reward modeling, LLM-judge-based rewards) and help design the user-facing RL flow on the platform — this part is greenfield.
- Build and calibrate evaluation harnesses: LLM-as-judge setups calibrated against human labels, golden datasets, regression evals for optimization runs.
- Improve the platform's guiding agent: prompt and tool design, eval-driven improvement loops, stress-testing scenarios and fixing what breaks.
- Run experiments for client projects (e.g. prompt compression vs. fine-tuning trade-off studies) and turn the results into repeatable platform features.
- Work closely with platform engineers on the SDK/API surface so that training and eval jobs are callable from a developer's existing workflow.
What we're looking for
- 4+ years in ML engineering or applied research, with at least 1–2 years hands-on with LLMs in production or research settings.
- Practical experience fine-tuning open-weight models (LoRA/full FT), including data curation and knowing when fine-tuning is the wrong answer.
- Solid grasp of LLM evaluation: building evals from scratch, LLM-as-judge pitfalls, calibration against human judgments.
- Strong Python engineering: you write code others can run, not just notebooks; comfortable with the training/inference stack (PyTorch, HF ecosystem, vLLM or similar).
- Product mindset: you'll often be the ML person closest to a client problem, so you need to reason about what's worth building, not just what's possible.
- Comfortable with ambiguity — priorities shift as we learn from pilots.
- Language: Fluency in English (B2 or above).
Nice to have
- Hands-on RL for LLMs: GRPO/PPO-style post-training, reward modeling, RLHF/RLAIF pipelines.
- Prompt optimization frameworks (DSPy/GEPA or similar) or prompt-compression research (gisting).
- Experience with distillation and quantization for cost/latency optimization.
- Experience building agentic systems (tool use, multi-step workflows) or shipping ML features in a self-serve product.
What we can offer
- You will be part of an international, dynamic environment that drives innovation and sets new standards in the AI and technology sector.
- Competitive compensation package including base salary, bonus, and ESOP.
- Paid PTO and benefits will vary depending on location.
- We offer a full remote or hybrid model (if you are based in NL or Serbia).
- IT setup and home office allowances.
Equal Opportunity Employer:
Toloka is committed to providing equal opportunity and fostering an inclusive environment. We welcome applications from all qualified individuals and do not discriminate on the basis of race, religion, color, national origin, sex, sexual orientation, gender identity, age, marital status, veteran status, disability, or any other characteristic protected by applicable law. Selection decisions are made based on qualifications, merit, and business need.
[Important Notice] Scam Alert Regarding Fake Job Postings
It has come to our attention that an individual or group is fraudulently impersonating Toloka to post fake jobs and solicit personal information from applicants. Please be aware:
- Official Communication: Our recruiting team will only contact you from an official " toloka.ai " email address. We will NEVER use Gmail, Yahoo, Tolokainc, toloka.inc, or other personal or seemingly business email accounts.
- Our Process: We will never ask for your bank account details, credit card number, or any fees as part of the application or interview process.
- Official Listings: All legitimate job openings are posted on our official careers page: https://toloka.ai/careers#job-list
What to do: If you see a suspicious job posting or have been contacted by someone you suspect is a scammer, please do not provide any personal information. Instead, report the incident to us directly at security@toloka.ai and report the profile/post to LinkedIn.We are taking this matter very seriously and are working with the appropriate parties to resolve it.
Thank you for your vigilance!
To learn how we collect, use, disclose, and store personal data, check out our Privacy Notice.