US$65
Machine Learning Engineer – ML Evaluation & Experiment Design
Remote in Portugal·Part-time·Added 13 days ago
Overview
Job details
US$65 an hour
Gross, as stated in the ad.
Fixed-term contract
Part-time hours.
Fully remote
Per the ad.
Requirements
Have 3+ years of experience
Mid-level role.
The role
Anyone AI is recruiting experienced Machine Learning Engineers for a specialized project focused on reviewing and evaluating machine learning challenges used in AI model training and evaluation.
The work involves analyzing ML experiments, datasets, metrics, and pipelines to determine whether challenges are technically sound, reproducible, appropriately difficult, and genuinely require strong machine learning reasoning.
What you'll do
You'll review ML challenges involving:
Experiment design and model selection
Small and synthetic datasets
Data quality and preprocessing
Distribution shift and data contamination
Label noise and feature leakage
Model evaluation and metric selection
Hyperparameter tuning
Train / validation / test methodology
Reproducibility and deterministic pipelines
Statistical significance of model improvements
A key part of the role is determining whether a challenge actually rewards good ML reasoning, rather than simply being solvable through brute-force model selection or large hyperparameter searches.
Reviewing ML challenges and determining whether they are well designed and technically solvable
Evaluating whether datasets contain meaningful and learnable signals
Identifying unintended shortcuts or artifacts in synthetic datasets
Determining whether tasks require genuine diagnosis of the underlying ML problem
Reviewing evaluation metrics and improvement thresholds
Detecting metric gaming, data leakage, and evaluation flaws
Verifying reproducibility across the complete data → model → evaluation pipeline
Assessing whether challenge difficulty is appropriately calibrated
Providing clear recommendations for improving, recalibrating, or excluding problematic tasks
How you'll work
Work Type: Remote
Engagement: Part-time, project-based consulting
Focus: Applied machine learning, experiment design, data quality, and model evaluation
This role is a strong fit for ML engineers who enjoy debugging experiments, understanding why models succeed or fail, identifying problems in datasets and evaluation pipelines, and designing rigorous machine learning experiments.
Requirements
What we're looking for
3+ years of hands-on applied machine learning experience
Strong experience with:
ML experiment design
Model selection
Hyperparameter tuning
Model evaluation
Data preprocessing and validation
Strong understanding of train, validation, and test splits
Ability to identify:
Data leakage
Label noise
Distribution shift
Spurious correlations
Feature leakage
Data contamination
Experience evaluating whether performance improvements are statistically meaningful rather than random fluctuations
Strong understanding of ML evaluation metrics and when different metrics are appropriate
Experience debugging ML workloads across CPU and GPU environments
Ability to analyze technical problems and provide clear written feedback
Nice to have
Experience creating or participating in Kaggle, DrivenData, or similar ML competitions
Experience designing benchmark datasets or ML challenges
Background in data-centric AI or dataset quality
Experience with synthetic data generation and validation
Familiarity with statistical testing, confidence intervals, and effect sizes
Experience with ML evaluation pipelines, RLHF, or AI model evaluation
Experience developing ML curricula or technical assessments
Understanding of common ML failure modes such as:
Shortcut learning
Spurious correlations
Goodhart's Law
Simpson's paradox
Metric gaming
About Anyone AI
Anyone AI is an EdTech company that develops AI-powered educational tools and platforms. Based in San Francisco, the company focuses on leveraging artificial intelligence to enhance learning experiences, offering products that range from AI tutoring systems to adaptive learning technologies. While the company is relatively young, it has established a presence in the AI education space, working with experts across various domains to train and refine its AI models.
Anyone AI has a significant remote-first presence in Spain, hiring experts in fields like mathematics, physics, biology, and software engineering to contribute to AI training and development. The company offers fully remote positions within Spain, making it an attractive option for international professionals seeking flexible work arrangements. With a focus on AI model training and development, Anyone AI provides opportunities for experts to apply their knowledge in a cutting-edge field, though its brand recognition is limited outside of niche AI and EdTech circles.
- Industry
- EdTech
- Employees
- 11–50
- Headquarters
- San Francisco, USA
- Website
- anyoneai.com
More jobs like this
or browse Machine learning & AI·Mid-level·Remote·Education & EdTech·Artificial Intelligence·EdTech



