Pivots Hiring
A
62

AI Research Engineer (Early Career)

1y relevant experience

Under Review

Executive Summary

The candidate is a technically credible candidate whose GenAI project work directly maps to the role's core stack — LangGraph, hybrid RAG, multi-agent orchestration, and adversarial evaluation — and whose production ML engineering background (fraud systems, MLflow, Airflow, Docker at scale) signals they can write code that ships, not just notebook experiments. The primary tension is that their 5 years of professional experience places them well above the entry-level framing, which creates dual risks: compensation misalignment with the $45k-$70k range, and uncertainty about whether this role represents genuine growth or a step backward for them. Their LLM experience, while technically sophisticated, is entirely project-based as of the application date, which is a gap for a role emphasizing production ownership. If compensation and motivation can be validated, they represents a high-upside borderline candidate who may accelerate faster than a typical entry-level hire — worth a screening conversation to resolve the key uncertainties before advancing.

Top Strengths

  • Directly relevant GenAI project work: LangGraph multi-agent orchestration with hybrid RAG, external data integration via MCP, and adversarial evaluation suite — strong technical alignment with the role's core stack
  • Production ML engineering discipline established through fraud detection and financial modeling at scale — MLflow, Airflow, Docker, PySpark signal beyond-notebook engineering habits
  • Demonstrates learning velocity: pivoted from traditional DS/fraud analytics into GenAI tooling within a structured program (EPAM x Tech Orda) and immediately applied it in a hackathon context
  • International exposure (exchange student in Germany, work experience in Serbia) and multilingual background suggest adaptability for a remote EU/US timezone role
  • Evaluation-first mindset evident in both fraud model monitoring (KL/JS divergence) and LLM test suite development — a discipline the job description explicitly values

Key Concerns

  • !Experience level mismatch: 5 years of professional DS experience positions them above the 0-2 year entry framing, which may create compensation tension given the $45k-$70k range and could signal they views this as a lateral step-down rather than a growth role
  • !LLM/GenAI experience is entirely project-based with no professional production deployment track record — the role requires owning pieces of the stack from prototype to deployment, and this gap would need rapid validation

Culture Fit

65%

Growth Potential

High

Salary Estimate

$55,000 - $80,000 (likely above posted range given 5 years of professional experience, even with limited production LLM background)

Assessment Reasoning

The candidate is scored BORDERLINE at 62 primarily because the technical skill alignment on the GenAI stack is genuine and the project work is meaningfully deeper than typical entry-level candidates, but two structural concerns prevent a FIT decision. First, their 5 years of professional experience significantly exceeds the 0-2 year entry-level framing, creating likely compensation misalignment with the posted $45k-$70k range and raising questions about motivation fit. Second, all LLM/GenAI work is project and capstone-based with no professional production deployment experience, which is the central requirement of the role. Missing required skills (vector databases, Rust/Java) and absence of a code sample further reduce confidence. A screening call focused on compensation expectations, motivation, and a live technical discussion of the capstone project would determine whether they clears the bar — if they's genuinely willing to work within the entry-level scope and salary range, their technical ceiling makes them a stronger-than-average borderline candidate.

Interview Focus Areas

Deep technical walkthrough of the Multi-Agent FinRisk Assistant: architecture decisions, tradeoffs in the hybrid RAG design, how the evaluation suite was designed and what failure modes it coversProduction engineering mindset: probe on observability, latency/cost tradeoffs, deployment and versioning practices — distinguish what they's done in production vs. in notebooks or capstone contextsMotivation and compensation alignment: understand why they's applying to an entry-level role given 5 years of experience, and whether the salary band and scope are genuinely acceptableAsync remote work style: given the remote-first, fast-moving team description, assess written communication habits, self-direction, and experience working across timezones

Code Review

FairMid Level

No code example was provided with the application, which is a meaningful gap for a role that explicitly distinguishes production-quality code from notebook-style work. The GitHub handle is referenced on the resume but no direct link or sample was submitted, making direct code quality assessment impossible. The project descriptions suggest architectural awareness, but without review of actual code the assessment remains inferred.

PythonLangGraphClaude APIDockerMLflowPySparkSQL
  • +GitHub profile referenced on resume (cpipi handle) and capstone project linked — indicates willingness to share work publicly
  • +Multi-agent project description implies modular, testable design with evaluation rigor (29/29 test suite)
  • -No code sample submitted — cannot directly assess code quality, style, readability, or production-readiness standards

Experience Overview

5y total · 1y relevant

The candidate has a solid ML engineering foundation built across banking and fintech, with genuine hands-on GenAI project work that directly maps to the role's LangGraph, RAG, and agentic systems requirements. However, their LLM experience is exclusively project/capstone-level as of mid-2025, and their overall professional history at 5 years significantly exceeds the entry-level framing of the role. The core technical signals are promising, but the experience level mismatch and absence of production LLM deployments are the primary evaluation risks.

Matching Skills

PythonLLMRAGLangGraphAI Agentsprompt engineering

Skills to Verify

RustJavavector databasesproduction LLM deployment experience
Candidate information is anonymized. Personal details are hidden for fair evaluation.