Computational Scientist, Modelling & Inference
- Company
- Claryx
- Location
- New York City, NY
- Work type
- Full Time
- Posted
- 2026-10-07
Job description
THE ROLE
We’re hiring a Computational Scientist to take over Claryx’s inference platform: the models that turn sequence data from patient infections and environmental samples into a defensible claim about what came from where, and when.
The list below is in priority order for the first year, with a specific focus on modelling and inference. The platform and security posture underneath it come with the role, but Drata supports them, and we don’t expect you to arrive with deep expertise. We’re early-stage, so whoever takes this role will set technical direction rather than follow it and use existing agentic coding tools to maximize output.
KEY RESPONSIBILITIES
Phylogenetic inference. Time-resolved trees, structured coalescent and birth–death models, ancestral state reconstruction, recombination-aware alignment, across clinical and environmental samples together.
Bioinformatics pipelines. Running and hardening our metagenomic and comparative genomics pipelines, from QC through strain-level comparison between clinical isolates and environmental samples.
Transmission attribution. Predict transmission pathways under incomplete sampling over a facility represented as a network of reservoirs.
Metagenomic inference. Strain-level deconvolution of mixed environmental populations.
Cloud infrastructure. AWS environment. Solidify infrastructure-as-code, proper CI/CD, predictable costs.
Compliance. Deepen HIPAA/SOC 2 controls to maintain compliance and remain audit-ready.
Results communication. Move reporting off static outputs onto an interactive dashboard.
QUALIFICATIONS
PhD in mathematical biology. Population genetics, molecular evolution, biostatistics, genomic epidemiology.
Mathematics. Deep knowledge of Likelihood and Bayesian inference, MCMC, continuous-time Markov chains, coalescent and birth-death processes, identifiability, model selection, etc.
Machine and deep learning. Understand backpropagation and optimization. Simulation-based inference and posterior estimation. GNNs and other network representations are a plus.
Phylogenetics. BEAST, IQ-TREE or equivalents. Transmission-tree tools are a plus.
AWS, Linux/bash, Git. Nextflow preferred for workflows. AWS in production – IAM, S3, Batch, Terraform.
Agentic coding tools. Structuring problems, managing context, reading generated code critically.