Network Engineer
- Location
- Toronto, CA
- Work type
- Full Time
- Posted
- 2026-07-28
Job description
Who You Are
Experienced in designing, deploying, and operating large-scale data center networks using modern leaf-spine architectures with strong Layer 2/Layer 3 networking expertise (BGP, EVPN-VXLAN, OSPF, TCP/IP).
Hands-on experience with enterprise networking platforms including Arista (EOS), Dell (SONiC), Cisco (IOS-XE), Fortinet firewalls, and lossless Ethernet fabrics for RDMA (RoCEv2, PFC, ECN, DCQCN).
Proven ability to troubleshoot complex production network issues across switches, routers, firewalls, and distributed infrastructure with a focus on reliability and performance.
Skilled in network automation using Python, Ansible, Terraform, or similar tools to improve operational efficiency and scalability.
What We Need
Design, deploy, and manage Tenstorrent’s global data center network infrastructure supporting AI accelerator clusters and production environments.
Configure, maintain, and troubleshoot switches, routers, firewalls, and high-speed network fabrics to ensure reliable, high-performance operations.
Collaborate with hardware, platform, DevOps, and infrastructure teams to support new cluster deployments, network expansions, and performance optimization.
Build scalable, secure, and highly available network architectures while driving automation, operational best practices, and continuous improvements in network provisioning and management.
What You Will Learn
Gain hands-on experience building and operating networking infrastructure for next-generation AI accelerator systems and large-scale compute clusters.
Develop expertise in high-performance data center networking technologies, including AI fabrics, high-speed Ethernet, and cloud connectivity.
Collaborate with world-class hardware, ASIC, firmware, and software engineers to solve complex infrastructure challenges and shape the networking architecture powering Tenstorrent’s AI platforms.
Build scalable automation, monitoring, and operational processes that improve the reliability, performance, and scalability of global AI infrastructure.