Data Engineer
Job Summary
We are seeking an experienced Data Engineer to own the caselaw and docket data layer for an AI-powered litigation platform. You will design and maintain production pipelines that ingest, normalize, and enrich large volumes of legal data from court systems and published opinions, transforming messy unstructured inputs into reliable structured data that powers downstream AI agents and judicial intelligence tools.
This high-ownership role sits at the foundation of the product, directly impacting research, drafting, and simulation capabilities used by trial attorneys. Prior experience with legal data is highly valued as it will accelerate your ability to deliver impact in this fast-moving environment.
Essential Functions
- Own the caselaw and docket data layer, building and maintaining scalable production pipelines that ingest data from PACER, NYSCEF, state court systems, and published opinions across millions of records
- Transform messy semi-structured inputs including PDFs, scanned documents, XML, and HTML into clean, queryable data structures
- Design and implement LLM-assisted extraction workflows using models such as Claude, Gemini, or OpenAI to convert unstructured legal text into structured outputs
- Collaborate closely with full-stack engineers to deliver high-quality data feeds for judicial behavioral intelligence and AI strategy layers
- Perform statistical analyses, optimize SQL queries, and apply RAG and semantic search techniques to enhance data utility
- Manage and optimize the AWS data infrastructure including S3 and RDS while maintaining SOC 2 compliance standards
- Operate production data pipelines with a focus on reliability, scalability, and data correctness
Required Qualifications
- 4-8 years of experience building and operating production data pipelines at scale
- Strong proficiency in Python and SQL for data transformation and analysis
- Hands-on experience with relational databases such as PostgreSQL and document stores like MongoDB
- Practical experience working with unstructured data formats including PDF parsing, XML, and HTML scraping
- Familiarity with cloud data platforms, specifically AWS S3 and RDS
- Ability to work independently in a high-ownership, ambiguous startup environment
- Strong product sense and willingness to contribute across the stack when needed
Preferred Qualifications
- Prior exposure to legal data sources such as court dockets, caselaw, or regulatory filings
- Experience implementing LLM-powered extraction or NLP workflows
- Background with semantic search tools, vector databases, or RAG pipelines
- Knowledge of data transformation tools such as dbt
- Comfort operating in a seed-stage, founder-led environment with high autonomy
Technical Skills
- Python
- SQL
- PostgreSQL
- MongoDB
- AWS S3
- AWS RDS
- LLM APIs (Claude, Gemini, OpenAI)
- PDF parsing and document extraction
- XML and HTML scraping
- RAG pipelines
- Semantic search (Pinecone, Voyage AI)
- dbt
Education & Certifications
Bachelor's degree in Computer Science, Engineering, or a related technical field is preferred. Relevant work experience and demonstrated impact in data engineering roles will be considered in lieu of formal education.
Compensation & Benefits
Base salary range of $180,000 - $220,000, plus 0.15 percent to 0.25 percent equity. Role is full-time and fully remote within the US, with a strong preference for candidates based in or near New York City who can maintain four or more hours of daily Eastern Time overlap. Visa transfers are supported.