Senior Software Engineer, AI Benchmarking
Job Description
REQUIREMENTS
- 5+ years of experience in software engineering (industry or open-source)
- Fluency in Python
- Proficiency with Docker and AWS (particularly ECS)
- Experience building software tools that use or evaluate LLMs or AI agents
- Self-directed and comfortable working on small, fast-moving teams
- Commitment to good coding practices and code quality
- Alignment with SecureBio’s mission of preventing catastrophic pandemics
Preferred
- Experience with AI evaluations (agent systems, scorer/grader design, trajectory analysis, jailbreaking, or red-teaming)
- Experience with Terraform, Tofu, or AWS CDK
- Experience with modern, component-based web frontends (React with TypeScript)
- Experience with Inspect AI, OpenAI-compatible APIs, Anthropic, or Together
- Experience with agent frameworks like Codex or Claude Code
RESPONSIBILITIES
- Perform pre-release and post-release assessments of frontier and open-source AI models using SecureBio’s suite of biosecurity evaluations
- Build adapters for models, identify and troubleshoot unexpected agent behaviors, and ensure methodological rigor
- Build, scale, maintain, and continuously improve SecureBio’s suite of biosecurity evaluations
- Improve evaluations to use state-of-the-art agent frameworks, jailbreaking methods, and elicitation strategies
- Develop and improve internal tools that enable researchers to run capability assessments at scale
- Contribute to and maintain cloud infrastructure to run evaluations and analyses
- Contribute analyses to the public dashboard of trends in AI biology capabilities
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers