Lead Distributed Systems Engineer
- Company
- Cape
- Location
- New York, NY
- Work type
- Full Time · Hybrid
- Posted
- 2026-09-10
Job description
The Team
At Cape, we are the architects of a privacy-centric movement that is just getting started. We are relentless builders, constantly innovating at the edge of what's possible in telecommunications. We operate on a foundation of high trust and high expectations. Our structure is flat, and collaboration matters more than hierarchy. As a member of our team, you will collaborate with world-class engineers, architects, and visionaries, and work across organizational lines to solve "impossible" problems and deliver mission-critical results for our users every single day.
The Role
Cape runs private cellular networks that operate disconnected for weeks at a time. Getting subscriber state where it needs to be means delivering it to nodes we can't reach, over links we don't control, onto hardware we can't trust.
You'll own the architecture and implementation of an edge data distribution system for private 5G bubbles, and lead the team that builds it. The core challenge is synchronizing data from the cloud to the edge reliably, securely, and efficiently. Because we are early in this product's lifecycle and iterating directly with customers, you will have significant autonomy to set the technical direction and shape the product definition.
This is a large and high-complexity problem space with a variety of challenges spanning multiple disciplines including distributed systems, security and building out the services that power end user device connectivity.
What You'll Work On
Architecture and technical direction. The change propagation pipeline, the reconciliation model for nodes returning after long periods offline, the consistency and conflict-resolution semantics, the trust boundaries. You define the guarantees this system commits to, and you're accountable for whether they hold when a node comes back after six weeks with divergent state.
Hands-on engineering. You lead the team and hold a high bar for execution, and you're most likely also rolling up your sleeves to build the hardest parts of the system yourself.
Technical roadmap and prioritization. Maintain a clear view of what the deployments in front of us need while making room for the investments that don't show up in any single customer's requirements but determine what this system can do two years from now.
Team. Recruit, develop, and retain strong distributed systems engineers. Set clear technical standards. Give people the context and the room to do good work.
Cross-functional partnership. Be a credible, honest technical counterpart to product, our forward deployed engineers, and customers. Translate field constraints into what they actually cost. Bring solutions, not just constraints.
Solving customer problems. You'll work directly with customers to understand their mission, ship something, watch how it holds up in their environment, and use what you learn to shape the next version. That includes knowing when the thing being asked for isn't the thing that's needed, and being able to educate the customer on the technical details of the solution.
What We Value
Reliability & Durability: Ensure data arrives exactly where it needs to be every time, leveraging an event-driven architecture so
Offline Resiliency: Build services that remain functional and consistent when offline for extended periods of time
Security in Contested Environments: Design systems under the explicit assumption that edge nodes cannot be inherently trusted and may become compromised at any point and design systems with this principle in mind
Connectivity: Build and scale the core telephony services that provide data connectivity to user devices at the edge and in the cloud
Experience
Distributed systems depth: You've designed and operated distributed systems in production and have real scar tissue around replication, consistency, and partition-related failure modes. You can reason about a system's guarantees precisely rather than by analogy, and you know the difference between a system that works and a system that provably works every time.
Event-driven architecture: Experience with Kafka, SQS, MQTT, RabbitMQ, Flink, or Spark in production. You understand delivery semantics, ordering, and replay well enough to reason about them under adverse conditions rather than happy-path ones.
Correctness under pressure: You've worked in environments where losing data was not an acceptable outcome, and you've been in situations where a customer needed something shipped and the honest answer was that the design wasn't ready. You have a way of thinking about that tradeoff and you can defend your decisions clearly without being defensive about them.
Engineering leadership: You have led engineers as a manager, a tech lead, or both. You can point to specific people you've grown and specific technical bets you owned end to end.
Communication that moves things forward: You can both explain a consistency guarantee to a customer who wants a feature now and an operational reality to an engineer who's optimizing for the wrong constraint.
Experience that’s helpful (though not required)
Building software that runs on customer hardware, on-prem, air-gapped, or in similar disconnected environments
Backend systems serving mobile devices or edge compute at scale
Logical replication, CDC pipelines, or multi-region and active-active database architectures
Radios, telephony, sim cards and modems or an appetite to learn about them
Working with government or defense customers