Senior Software Engineer, AI Benchmarking

Application ends: October 19, 2026

Job Description

REQUIREMENTS

  • 5+ years of experience in software engineering (industry or open-source)
  • Fluency in Python
  • Proficiency with Docker and AWS (particularly ECS)
  • Experience building software tools that use or evaluate LLMs or AI agents
  • Self-directed and comfortable working on small, fast-moving teams
  • Commitment to good coding practices and code quality
  • Alignment with SecureBio’s mission of preventing catastrophic pandemics

Preferred

  • Experience with AI evaluations (agent systems, scorer/grader design, trajectory analysis, jailbreaking, or red-teaming)
  • Experience with Terraform, Tofu, or AWS CDK
  • Experience with modern, component-based web frontends (React with TypeScript)
  • Experience with Inspect AI, OpenAI-compatible APIs, Anthropic, or Together
  • Experience with agent frameworks like Codex or Claude Code

RESPONSIBILITIES

  • Perform pre-release and post-release assessments of frontier and open-source AI models using SecureBio’s suite of biosecurity evaluations
  • Build adapters for models, identify and troubleshoot unexpected agent behaviors, and ensure methodological rigor
  • Build, scale, maintain, and continuously improve SecureBio’s suite of biosecurity evaluations
  • Improve evaluations to use state-of-the-art agent frameworks, jailbreaking methods, and elicitation strategies
  • Develop and improve internal tools that enable researchers to run capability assessments at scale
  • Contribute to and maintain cloud infrastructure to run evaluations and analyses
  • Contribute analyses to the public dashboard of trends in AI biology capabilities

Are you interested in this position?


Apply by clicking on the “Apply Now” button below!

#CrossChannelJobs #JobSearch
#CareerOpportunities #HiringNow
#Employment #JobOpenings
#JobSeekers