← All Positions
Forward Deployed Machine Learning Engineer
Job Summary
We are seeking a Forward Deployed Machine Learning Engineer to join an early-stage AI infrastructure company as the first dedicated MLE focused on benchmarks and evaluations. You will partner closely with the GM, researchers, and early customers to establish the technical foundation for AI model evaluation capabilities. This high-impact role combines backend infrastructure development with customer-facing engineering in a fast-paced startup environment backed by leading venture investors.
Essential Functions
- Partner with the GM and early customers to define, design, and build benchmarks and evaluations across different domains and modalities
- Build and own backend infrastructure including data pipelines, execution environments, storage, and orchestration
- Stand up sandboxed environments for agentic evaluations requiring tools, code execution, or multi-step tasks
- Own the engineering portion of customer engagements from start to finish
- Identify repeatable evaluation patterns and infrastructure gaps to drive future product opportunities
- Scope and execute ambiguous technical problems independently from feasibility through delivery
Required Qualifications
- 4+ years of engineering experience with hands-on machine learning model evaluation work
- Prior ownership of backend and infrastructure including data pipelines and execution environments
- 3-8 years of experience in machine learning engineering, evaluation systems, and end-to-end ownership in fast-moving environments
- High tolerance for ambiguity and strong bias to action in dynamic startup settings
- Strong written communication skills for customer-facing and cross-functional collaboration
- Customer-facing engineering experience managing enterprise customers
- Experience building data pipelines for large-scale data
- Degree in Computer Science, Physics, or related technical field
Preferred Qualifications
- Experience building benchmarks, evaluations, or human data pipelines for large language models
- Background at early-stage B2B startups, forward-deployed engineering roles, or high-ownership positions at larger organizations
- Proficiency in ML evaluation frameworks and benchmark design such as LLM-as-judge approaches
- Experience interfacing directly with AI researchers and foundation model labs
- Published work or open-source contributions in ML evaluations or benchmarks
Technical Skills
- Python
- LLMs
- Data Pipelines
- Agentic Systems
- Code Execution Sandboxes
- RL Environments
- Backend Infrastructure
- Orchestration
- Benchmark Design
- ML Evaluation Frameworks
Education & Certifications
Bachelor's degree or higher in Computer Science, Physics, or a related technical field is required.