Pivots Hiring
A
74

AI Research Engineer (Early Career)

2y relevant experience

Qualified

Executive Summary

The candidate is an unusually strong entry-level candidate whose profile significantly exceeds typical recent-graduate expectations. They combine published multi-agent LLM research, hands-on production RAG implementation, CTO-level product ownership, and fluency in Python, Java, and Rust — achieving near-complete alignment with the role's technical requirements. The primary risk factors are the absence of any submitted code sample (which must be addressed in technical screening) and the complexity of their current commitments (PhD + CTO + multiple roles). Their academic and entrepreneurial ambition is evident and their cover letter's stated desire to learn from experts in a structured AI environment is culturally credible given their pioneering role in previous companies. Subject to a strong technical interview, they are a FIT for this role and likely one of the stronger candidates at this experience level.

Top Strengths

  • Published peer-reviewed research on multi-agent LLM architectures with LoRA/QLoRA fine-tuning and LLM-as-a-judge evaluation — directly maps to the role
  • Implemented RAG in a production recruiting platform, with pgvector and Neo4j knowledge graphs — practical, real-world retrieval experience
  • Full-stack product ownership as CTO of Sila — demonstrates ability to operate independently and ship end-to-end, critical for a small remote team
  • Strong multi-language engineering profile (Python, Java, Rust) matching all preferred language requirements
  • PhD candidacy combined with 2+ years of concurrent commercial engineering work signals both intellectual depth and practical execution ability

Key Concerns

  • !No code sample or accessible GitHub repository submitted — production code quality remains unverified and must be assessed via technical screening
  • !Multiple overlapping concurrent roles raise questions about depth of focus and availability — needs clarification on current commitments and hours

Culture Fit

72%

Growth Potential

High

Salary Estimate

$45,000 - $60,000 (likely toward lower-mid range given Serbia location and early-career stage, but CTO/founder experience and research publications may push expectations higher)

Assessment Reasoning

The candidate is assessed as FIT (score 74) primarily because they meets or exceeds the core technical requirements: Python proficiency, LLM/RAG/agent hands-on experience, LangChain/LangGraph exposure, vector database usage (pgvector, Neo4j), Java and Rust skills, and published research directly on multi-agent LLM systems. Their production experience (deployed Sila SaaS, RAG in Gramian recruiting platform, GCP CI/CD at Xenon Seven) demonstrates they write real code for real systems rather than notebook-only work — directly addressing the role's explicit concern about that. Their PhD candidacy and multi-publication research record add intellectual credibility beyond what the salary range typically attracts. The score does not reach the 80s due to two meaningful gaps: (1) no code sample was provided, creating genuine uncertainty about code quality that must be resolved in technical screening, and (2) their numerous concurrent commitments require clarification to confirm they can engage fully with this role. These are screening risks rather than disqualifying factors. A technical interview with a code exercise focused on LangGraph agentic workflow design and RAG pipeline architecture is strongly recommended before a hiring decision is finalized.

Interview Focus Areas

Technical deep-dive on RAG pipeline design — ask them to walk through the Gramian RAG implementation: chunking strategy, embedding model choice, retrieval quality evaluation, and latency tradeoffsMulti-agent architecture trade-offs — probe depth of understanding from their published paper: how would they design an agentic workflow in LangGraph for a real marketing/SEO task?Production engineering practices — assess observability, versioning, cost/latency awareness, and how they handle failure modes in deployed ML systemsConcurrent commitments clarification — understand current bandwidth given CTO role at Sila, PhD candidacy, and teaching assistant position

Code Review

FairMid Level

No code was submitted for direct review, which is a meaningful gap for an engineering role. The resume narrative strongly implies production engineering competence across multiple stacks and real deployed systems, but this cannot be validated without seeing actual code. The score reflects the uncertainty inherent in this absence rather than a negative assessment of the candidate's abilities. A technical interview or take-home exercise is essential to close this gap.

PythonFastAPIJavaSpring BootDockerPostgreSQLpgvectorAngularNeo4jLangChainLangGraph
  • +Resume describes production-grade implementations (FastAPI microservices, CI/CD on GCP, pdflatex pipeline, SSE, message queues) suggesting solid engineering practices
  • +Evidence of integration testing discipline mentioned across multiple roles
  • -No code sample or GitHub profile was submitted — assessment is entirely inferred from resume descriptions and cannot be verified
  • -Resume references a GitHub URL (github.com/Novke) but it was not provided for review, preventing any direct code quality evaluation

Experience Overview

3y total · 2y relevant

The candidate presents a notably strong profile for an entry-level AI Research Engineer, with published multi-agent LLM research, hands-on RAG implementation in production, and CTO-level ownership of an AI SaaS product. Their stack alignment (Python, LangChain/LangGraph, vector databases, Java, Rust) is excellent for the role requirements. The primary gap is lack of visible code output (no GitHub, no code sample), which limits confidence in assessing production code quality.

Matching Skills

PythonLLMRAGAI AgentsJavaRustLangChainLangGraphvector databasesembeddingsmulti-agent systemsprompt engineeringproduction-quality code

Skills to Verify

explicit fine-tuning deployment in productionformal LLM evaluation frameworks beyond research context
Candidate information is anonymized. Personal details are hidden for fair evaluation.