← All jobs

AI Ops Engineer

Salary
Not published
Location
Remote, United States
Work type
Remote
Posted
today

Apply on company site (opens in new tab)

LangChainRAGEmbeddingsTool useMulti-agentOrchestrationMCPGCPPythonGo

Technology is our how. And people are our why. For over two decades, we have been harnessing technology to drive meaningful change.

By combining world-class engineering, industry expertise and a people-centric mindset, we consult and partner with leading brands from various industries to create dynamic platforms and intelligent digital experiences that drive innovation and transform businesses.

From prototype to real-world impact - be part of a global shift by doing work that matters.

We are looking for an AI Automation Engineer – Level II to join an enterprise AI initiative focused on transforming Infrastructure & Operations (I&O) through Generative AI, intelligent automation, and agentic systems.

This is a hands-on engineering role for someone who combines a strong foundation in backend/software engineering and systems integration with practical experience building solutions using LLMs, agentic AI, RAG, and modern AI development tools.

You will help build an Enterprise Operational AI Platform that provides engineers with a unified, intelligent interface to operational data and capabilities across platforms such as ServiceNow, Dynatrace, Zabbix, and Google Cloud Platform (GCP). The platform will evolve from read-only AI-assisted workflows for incident investigation and root cause analysis toward increasingly autonomous operational capabilities.

What You Will Do

  • Design, build, and enhance LLM-powered and agentic AI solutions for enterprise Infrastructure & Operations use cases.
  • Develop and integrate domain-specific AI agents that collaborate to answer questions, investigate operational issues, and execute defined workflows.
  • Build Model Context Protocol (MCP) integrations and tool-calling capabilities that securely connect AI agents with enterprise platforms and APIs.
  • Develop backend services and integrations primarily using Python and REST APIs.
  • Integrate the AI platform with infrastructure and operational systems such as ServiceNow/CMDB, Dynatrace, Zabbix, and GCP.
  • Build pipelines and services that ingest, normalize, enrich, and contextualize structured and unstructured operational data for AI consumption.
  • Implement Retrieval-Augmented Generation (RAG) and grounding strategies that provide LLMs with accurate enterprise context.
  • Develop initial read-only AI workflows supporting incident triage, incident management, root cause analysis (RCA), infrastructure discovery, and Help Desk automation.
  • Progressively extend workflows toward controlled automation, including change-window support, maintenance suppression, proactive outage prevention, and coordinated self-healing.
  • Apply appropriate guardrails, validation, access controls, and human-in-the-loop patterns as AI workflows move from recommendations toward autonomous actions.
  • Use modern AI-assisted engineering tools, such as Devin, Windsurf, or comparable platforms, to accelerate software development and automation.
  • Implement AI observability and evaluation capabilities to measure response quality, confidence, token consumption, reliability, latency, and operational outcomes such as MTTR.
  • Collaborate with AI architects, platform engineers, observability teams, IT operations, enterprise search, data, and security teams to deliver production-ready solutions.
  • Help improve knowledge quality and cross-validation mechanisms to reduce LLM hallucinations and ensure responses are grounded in authoritative enterprise data.
  • Contribute to engineering standards and reusable patterns for deploying secure, scalable, observable, and maintainable enterprise AI systems.
  • 3–5 years of overall relevant engineering experience, combining modern AI engineering with a strong software, integration, data, infrastructure, or AIOps foundation.
  • Approximately 1–2 years of hands-on experience with Generative AI/LLMs, including building applications or workflows using modern LLM platforms.
  • Approximately 2–3 years of foundational engineering experience in one or more areas such as backend software development, Python engineering, API integration, data engineering, cloud engineering, automation, or AIOps.
  • Strong programming skills in Python, including experience developing production-quality backend services and automation.
  • Strong experience designing, building, and consuming RESTful APIs and integrating multiple enterprise systems.
  • Practical knowledge of LLMs, prompt engineering, context management, embeddings, vector retrieval, and Retrieval-Augmented Generation (RAG).
  • Hands-on experience with agentic or multi-agent AI frameworks, such as LangChain/LangGraph, AutoGen, CrewAI, or comparable technologies.
  • Experience with or a strong understanding of Model Context Protocol (MCP), function/tool calling, agent registries, and AI orchestration patterns.
  • Experience with modern AI coding assistants or autonomous development tools, such as Devin, Windsurf, or comparable solutions.
  • Familiarity with enterprise IT and infrastructure platforms, ideally including one or more of ServiceNow/CMDB, Dynatrace, Zabbix, and GCP.
  • Understanding of ITSM, incident management, observability, monitoring, infrastructure telemetry, or AIOps concepts.
  • Experience working with both structured and unstructured data and preparing enterprise information for AI consumption.
  • Understanding of AI safety, data governance, security, access control, grounding, hallucination mitigation, and responsible AI principles.
  • Experience designing or operating production systems where reliability, scalability, observability, and maintainability are important.
  • Strong systems-thinking and problem-solving skills, with the ability to understand complex enterprise environments and translate operational requirements into practical technical solutions.
  • Ability to collaborate effectively with architects, software engineers, infrastructure teams, IT operations, security, and other technical stakeholders.
  • Comfortable working iteratively, delivering measurable value through a phased approach from read-only AI assistance to controlled automation and ultimately agentic execution.

Discover some of the global benefits that empower our people to become the best version of themselves:

  • Finance: Competitive salary package, share plan, company performance bonuses, value-based recognition awards, referral bonus;
  • Career Development: Career coaching, global career opportunities, non-linear career paths, internal development programmes for management and technical leadership;
  • Learning Opportunities: Complex projects, rotations, internal tech communities, training, certifications, coaching, online learning platforms subscriptions, pass-it-on sessions, workshops, conferences;
  • Work-Life Balance: Hybrid work and flexible working hours, employee assistance programme;
  • Health: Global internal wellbeing programme, access to wellbeing apps;
  • Community: Global internal tech communities, hobby clubs and interest groups, inclusion and diversity programmes, events and celebrations.

Additional Employee Requirements

  • Participation in both internal meetings and external meetings via video calls, as necessary.
  • Ability to go into corporate or client offices to work onsite, as necessary.
  • Prolonged periods of remaining stationary at a desk and working on a computer, as necessary.
  • Ability to bend, kneel, crouch, and reach overhead, as necessary.
  • Hand-eye coordination necessary to operate computers and various pieces of office equipment, as necessary.
  • Vision abilities including close vision, toleration of fluorescent lighting, and adjusting focus, as necessary.
  • For positions that require business travel and/or event attendance, ability to lift 25 lbs, as necessary.
  • For positions that require business travel and/or event attendance, a valid driver’s license and acceptable driving record are required, as driving is an essential job function.
  • If requested, reasonable accommodations will be made to enable employees requiring accommodations to perform the essential functions of their jobs, absent undue hardship.

USA Benefits (Full time roles only, does not apply to contractor positions)

  • Robust healthcare and benefits including Medical, Dental, vision, Disability coverage, and various other benefit options
  • Flexible Spending Accounts (Medical, Transit, and Dependent Care)
  • Employer Paid Life Insurance and AD&D Coverages
  • Health Savings account paired with our low-cost High Deductible Medical Plan
  • 401(k) Safe Harbor Retirement plan with employer match with immediately vest

At Endava, we’re committed to creating an open, inclusive, and respectful environment where everyone feels safe, valued, and empowered to be their best. We welcome applications from people of all backgrounds, experiences, and perspectives—because we know that inclusive teams help us deliver smarter, more innovative solutions for our customers. Hiring decisions are based on merit, skills, qualifications, and potential. If you need adjustments or support during the recruitment process, please let us know.

Apply on company site (opens in new tab)