Research Platform Engineer
- Company
- Tower Research Capital
- Location
- New York, NY
- Work type
- Full Time
- Posted
- 2026-10-05
Job description
Summary
As part of Tower Research's Core Engineering team, you will develop a multi-tenant research compute platform capable of dynamically orchestrating large scale ML workloads across a hybrid compute infrastructure of GPUs and CPUs. Your primary mission is to work closely with Quant Researchers, Portfolio Managers, and Infrastructure teams, to build a unified, elastic, and multi-tenant research substrate.
Key Responsibilities
Elevating Researcher Experience: Design intuitive platform abstractions, APIs, scheduler wrappers, and a durable job control plane so researchers can seamlessly launch simulations, distributed training runs, and complex research pipelines without incurring infrastructure overhead
Architecting Multi-Tenant Scheduling & Isolation: Build a multi-tenant compute substrate combining HPC-grade batch scheduling (gang scheduling, topology awareness, fair share) with cloud-native flexibility, enforcing strict tenant isolation across trading desks
Enabling Multi-Datacenter & Multi-Cloud Compute Portability: Establish infrastructure agnostic execution abstractions that enable compute workloads to run transparently across multiple on-premise datacenters or burst into external cloud providers
Standardizing Workflow Orchestration: Evaluate, select, and integrate production-grade job graph orchestration engines to automate multi-stage feature, training, and backtesting pipelines
Building Fault-Tolerant Research Pipelines: Implement automated failure detection, retry-on-fault mechanisms, and high-performance checkpointing to ensure long-running distributed jobs survive hardware degradation and faults without manual intervention
Optimizing Data-Paths: Profile and eliminate I/O bottlenecks, ensuring distributed ML workloads align with high-speed network fabrics and high-performance storage
Delivering Observability & Cost Transparency: Implement comprehensive telemetry to track compute utilization, queue pressure, and GPU/CPU cost attribution, giving Portfolio Managers and Senior Management clear visibility into resource consumption and ROI
Technical Requirements
HPC Administration: Deep experience with HPC job schedulers (e.g., Slurm), including gang scheduling, fair-share priority trees, topology-aware node allocation, and containerized HPC execution
Kubernetes-native Engineering: Advanced understanding of Kubernetes architecture, CRDs, HPC focused operators, admission controllers and GPU-native schedulers
Distributed ML Computing Frameworks: Familiarity with distributed computing frameworks (e.g., Ray) and deep learning frameworks (e.g., PyTorch, JAX) with multi-node scaling primitives
Workflow Orchestration Expertise: Hands-on experience evaluating, architecting, and operating job graph orchestration frameworks
Multi-Platform Architecture: Experience designing vendor-agnostic infrastructure layers, cloud-bursting strategies, and compute execution spanning multiple on-premise datacenters and public cloud environments
Storage & Network Performance: Familiarity with high-performance storage solutions and high-speed network fabrics for large scale research workloads
Compute Observability: Proven track record building cluster-wide telemetry and cost-attribution platforms for multi-tenant environments
Architectural Leadership Mindset: A strong platform engineering mindset focused on reducing friction for researchers while maintaining rigorous operational efficiency, cost transparency, and system scalability