Senior ML / Evaluation Engineer

Remote in Portugal·Added today

EPAM

5 open roles

Overview

Job details

  • Fully remote

    Per the ad.

Pay & benefits

  • Competitive compensation depending on experience and skills
  • Variety of projects within one company
  • Being a part of a project following engineering excellence standards
  • Individual career path and professional growth opportunities
  • Internal events and communities
  • Flexible work hours

Requirements

  • 5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform development
  • Hands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)
  • Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assurance
  • Proficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)
  • Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoring

The role

We're looking for a Senior ML / Evaluation Engineer to join our team in Portugal in a fully remote working mode. In this role, you will own the design and implementation of advanced evaluation frameworks for an Enterprise Agent Development Platform-a production-grade, cloud-native ecosystem enabling scalable, secure AI agent deployment. You will create evaluation strategies that combine LLM-as-judge grading with deterministic checks, define enterprise evaluation standards, and implement CI/CD deployment gates to enforce quality metrics prior to release. This position requires strong expertise in ML system testing, evaluation design, and integration into automated pipelines for agentic environments.

Responsibilities
  • Design and implement multi-layer evaluation frameworks for agentic workflows and AI-driven applications
  • Build LLM-as-judge evaluators leveraging AWS AgentCore built-in modules and custom logic for correctness and helpfulness checks
  • Develop deterministic evaluators as AWS Lambda functions for rule-based validation
  • Define enterprise evaluation standards, including mandatory dimensions, scoring criteria, and pass/fail thresholds
  • Implement CI/CD deployment gates using on-demand evaluation modes to enforce quality in automated pipelines
  • Enable online evaluation in production by integrating sampling-based evaluation strategies and PII detection guardrails
  • Incorporate observability signals (OpenTelemetry spans) from AWS AgentCore into grading frameworks for trace-level assessment
  • Generate metrics, logs, and dashboards from evaluation outcomes via CloudWatch or equivalent monitoring platforms
  • Collaborate with platform, orchestration, and DevOps teams to maintain evaluation reliability and scalability
Requirements
  • 5+ years of experience in ML engineering, AI evaluation frameworks, or AI platform development
  • Hands-on expertise designing LLM evaluation frameworks (LLM-as-judge and deterministic graders)
  • Practical experience implementing CI/CD deployment gates for ML model or AI agent quality assurance
  • Proficiency in Python for building evaluation logic (deterministic Lambda-based evaluators)
  • Strong understanding of advanced validation dimensions, including multi-turn context integrity and workflow-level scoring
Nice to have
  • Familiarity with AWS AgentCore Evaluations API (CreateEvaluation, GetEvaluationResult)
  • Exposure to AWS Bedrock Guardrails for compliance and sensitive data validation
  • Experience integrating evaluation metrics into AWS CloudWatch for monitoring and alerting
  • Knowledge of OTel instrumentation and trace ingestion for quality scoring inputs
Benefits
  • Competitive compensation depending on experience and skills
  • Variety of projects within one company
  • Being a part of a project following engineering excellence standards
  • Individual career path and professional growth opportunities
  • Internal events and communities
  • Flexible work hours

Hiring process

3 steps3 interviews
  1. Apply

  2. Conversation with Talent Acquisition

    With a Talent Acquisition Specialist

    An introductory conversation with a Talent Acquisition Specialist.

  3. Technical interview

    Video call

    An online technical interview.

  4. Interview with the hiring manager

    With a hiring manager

    A general interview with the hiring manager.

  5. Offer

  • The block is a template on EPAM location pages; the Spain page itself does not show it.

About EPAM

EPAM Systems is a global leader in digital platform engineering and software development services, headquartered in Newtown, Pennsylvania, USA. The company partners with a wide range of industries, including finance, healthcare, retail, and technology, to deliver innovative solutions in cloud, data, artificial intelligence, and enterprise platforms. With over 60,000 employees worldwide, EPAM is recognized as one of the largest IT services companies in the world, consistently ranked among the top providers in the industry.

In Spain, EPAM has established a significant presence with offices in Madrid and Barcelona, serving as key hubs for its European operations. The company is known for its collaborative and innovative work culture, offering opportunities for international professionals to work on cutting-edge projects with global teams. EPAM's Spain offices are particularly active in areas like SAP, cloud engineering, data engineering, and capital markets, making it an attractive destination for skilled IT professionals looking to relocate to Spain.

Industry
IT Services
Founded
1993
Employees
50,000–100,000
Headquarters
Newtown, USA
Website
epam.com
All open roles at EPAM
WhatsApp