← All Positions

Software Engineer - RL Environments

Innovative AI research organization San Francisco, CA Direct Hire Other $180,000 - $220,000

Job Summary

Design datasets and evaluation rubrics that shape how cutting-edge AI models learn and improve. Work closely with research teams at leading AI labs, running rapid experiments to diagnose failure modes and refine metrics for model progress. Your work directly impacts large-scale model training runs across diverse domains.

Essential Functions

  • Design data slices and explore data shapes that expose meaningful model failure modes across domains like finance, code, and enterprise workflows
  • Build and refine evaluation rubrics and reward signals for RLHF and RLVR training pipelines
  • Model annotator behavior and run experiments to improve different model capabilities
  • Develop quantitative frameworks for measuring dataset quality, diversity, and downstream impact on model alignment and capability
  • Create and manage both real-world and synthetic data pipelines
  • Partner with lab research teams to translate their training objectives into concrete data and evaluation specifications

Required Qualifications

  • 1-4 years of experience as a software engineer
  • Strong backend development focus with fast coding abilities
  • Experience creating benchmarks, with supervised fine-tuning (SFT) or reinforcement learning (RL)
  • CS or similar degree from a top university
  • Depth in Python, TypeScript, and other fullstack or backend languages
  • Bias for action and execution, willing to tackle difficult and tedious work

Preferred Qualifications

  • Experience at fast-growing startups or creating complex simulations
  • Background at top VC-backed startups, quant firms, hedge funds, or similar
  • Track record of side projects, published papers, or AI research/products
  • Previous experience as a founder or early-stage startup engineer
  • Explicit interest or experience in reinforcement learning

Technical Skills

  • Python
  • TypeScript
  • Backend development
  • Reinforcement Learning (RL)
  • Supervised Fine-Tuning (SFT)
  • RLHF (Reinforcement Learning from Human Feedback)
  • RLVR pipelines
  • Data pipeline development
  • Quantitative framework development
  • Benchmark creation

Education & Certifications

CS or similar degree from a top university.

Compensation & Benefits

Competitive salary range of $180,000 - $220,000 based on experience, plus comprehensive benefits package.

Apply for This Position

📄 Drag and drop or browse PDF, DOC, or DOCX (10 MB max)