Pivots Hiring
M
94

ML Infrastructure Engineer / Founding ML Lead

11y relevant experience

Qualified

Executive Summary

The candidate is an exceptionally qualified candidate who exceeds the technical bar for this role in nearly every dimension. With a CMU PhD, 5 years as a Senior Staff ML Engineer Lead at Meta driving $600M+ in annual revenue impact, and deep hands-on expertise in LLMs, RLHF, RL ranking systems, and production ML infrastructure at planetary scale, they represents a rare caliber of ML engineer. Their leadership of 18 engineers and direct CXO-level partnership experience maps precisely to the Founding ML Lead to CTO growth path. The primary risks are compensation alignment — Meta's total comp likely far exceeds this role's salary range — and a potential ramp-up period on standard cloud MLOps tooling versus Meta's internal infrastructure. If equity and mission can bridge the compensation gap, this candidate would be a transformative hire for a founding team.

Top Strengths

  • Production LLM infrastructure at unprecedented scale — fine-tuning, RLHF, RAG, and serving systems for 3B+ users with direct revenue attribution exceeding $600M annually
  • PhD-level theoretical foundation in RL and optimization combined with 12 years of applied engineering — rare research-to-production bridge capability
  • Team leadership at scale (18 engineers) with demonstrated cross-functional product ownership, directly aligned with the Founding ML Lead to CTO trajectory
  • End-to-end ML system ownership: data pipelines, feature engineering, model training, serving infrastructure, monitoring, and experiment frameworks — no gaps in the full lifecycle
  • Strong quantitative finance background adds differentiated capability in optimization under uncertainty, systematic decision-making, and algorithmic thinking relevant to AI+blockchain intersection

Key Concerns

  • !Cloud platform ramp-up: Candidate's production experience is primarily on Meta's internal ML infrastructure; familiarity with AWS/GCP/Azure-native MLOps tooling (SageMaker, MLflow, W&B) should be probed and may require 1-3 months of ramp-up
  • !Compensation and stage risk: A Senior Staff Engineer at Meta is likely earning $500K-$1M+ in total compensation — the stated $100K-$150K salary range represents a significant financial step-down that may require substantial equity upside to be compelling

Culture Fit

85%

Growth Potential

High

Salary Estimate

$400K-$900K+ total compensation at current Meta level; open to $100K-$150K base only if equity package is substantial and founding team opportunity is compelling

Assessment Reasoning

FIT decision is clear and high-confidence. The candidate meets or exceeds every core requirement: PhD from CMU (RL/ML coursework), 12 years of ML experience with 5+ years in senior production roles, deep LLM and RLHF expertise demonstrated through shipped products at Meta scale, end-to-end ML infrastructure ownership (training pipelines, feature stores, serving systems, monitoring), team leadership of 18 engineers, and a proven track record of translating research into products with measurable business impact. They satisfies 90%+ of required and preferred skills. The two gaps — explicit standard cloud platform experience and no public code samples — are minor given their seniority and the context of working at Meta with internal infrastructure. The sole meaningful risk is compensation misalignment given the significant gap between Meta total comp and the posted salary range, which is a recruiting challenge rather than a qualification concern. Recommend fast-tracking to technical interview with focus on cloud platform fluency and startup motivation.

Interview Focus Areas

Cloud infrastructure fluency: Probe depth of AWS/GCP/Azure experience and familiarity with standard MLOps tooling outside Meta's internal ecosystem (SageMaker, MLflow, Weights & Biases, Terraform)Early-stage startup mindset: Assess comfort with ambiguity, resource constraints, and wearing multiple hats after operating with significant Meta resources and team supportMotivation and compensation alignment: Understand what draws them to an early-stage role at this salary range and whether equity structure and mission are sufficient driversSystem design from scratch: Present a blank-slate ML architecture problem to evaluate independent technical decision-making without organizational scaffoldingBlockchain/Web3 domain curiosity: Assess openness to the AI+blockchain intersection even without prior domain experience

Code Review

GoodPrincipal Level

No code sample was submitted, which limits direct assessment. However, the technical depth described — CUDA-level optimizations, distributed training infrastructure, custom inference pipelines, and RL system implementations — strongly implies principal-level engineering capability. An interview coding exercise or system design session is recommended to confirm hands-on fluency.

PythonPyTorchCUDAC++SQLDistributed Training (FSDP/DeepSpeed/Megatron)TorchServeTritonKubernetesAirflowSparkRay
  • +Resume details indicate deep systems-level engineering: CUDA-aware distributed training (FSDP, DeepSpeed, Megatron), custom feature stores, and model serving with Triton/TorchServe — indicative of principal-level coding capability
  • +RLHF pipeline implementation (PPO, DPO, reward modeling at 5M+ human labels) and transformer-based model development suggests strong research-to-code translation ability
  • -No code sample, GitHub profile, or open-source contributions were provided — actual code quality cannot be directly assessed and must be inferred from resume and role context

Experience Overview

12y total · 11y relevant

The candidate is an exceptionally strong candidate with a PhD from CMU, 12 years of ML experience, and 5 years as a Senior Staff ML Engineer Lead at Meta operating at planetary scale. Their hands-on LLM fine-tuning, RLHF infrastructure, RL ranking systems, and production ML platform work directly map to the role's core requirements. The primary gap is explicit experience with standard cloud platforms versus Meta's internal tooling, which is a minor ramp-up concern given their demonstrated infrastructure mastery.

Matching Skills

PythonPyTorchLLMsMLOpsDistributed TrainingRLHFRAG SystemsModel ServingKubernetesA/B TestingFeature StoresTransformer ModelsMulti-Task Learning

Skills to Verify

TensorFlow (PyTorch-primary)Explicit AWS/GCP/Azure project ownership (Meta likely uses internal infra)Blockchain/Web3 domain experience
Candidate information is anonymized. Personal details are hidden for fair evaluation.