Cerebras Systems Inc.
Jobs at Cerebras Systems Inc.
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Recently posted jobs
Artificial Intelligence • Hardware • Software • Semiconductor
Lead architecture and evolution of enterprise, data center, and cloud networks for hyperscale AI; design secure, high-performance fabrics (spine/leaf, RDMA); build AI-agent review frameworks; implement segmented zero-trust network designs; set resiliency, observability, and capacity standards; mentor global engineers and guide cross-functional initiatives.
Artificial Intelligence • Hardware • Software • Semiconductor
Join IT & Security to secure and scale enterprise IT, cloud, network, and infrastructure; build automation and tooling; support detection, response, and vulnerability management; improve identity, endpoint, and systems security; and partner cross-functionally to develop processes that enable secure, reliable, and scalable AI workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Artificial Intelligence • Hardware • Software • Semiconductor
Execute hardware bring-up, validation, and telemetry monitoring for Cerebras AI clusters in data centers. Perform power-on sequencing, first-line troubleshooting, log collection, incident support under senior guidance, and contribute feedback to tooling and documentation while learning system architecture and networking fundamentals.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement system-level debugging, validation, and observability platforms. Build automated anomaly collection/analysis, visualization and root-cause tools, failure classification and monitoring frameworks. Extend compilers, runtimes and programming interfaces for profiling and instrumentation, improve bring-up and low-level debug workflows, lead cross-functional initiatives, support incident response, and establish debuggability and reliability best practices.
Artificial Intelligence • Hardware • Software • Semiconductor
Design, implement, optimize, and validate high-performance ML and linear algebra kernels for Cerebras hardware. Develop low-level assembly and CSL routines, use mathematical performance models, create unit/system tests, and collaborate with chip and system architects to maximize compute utilization and scale kernels for state-of-the-art AI/HPC workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Design and implement high-performance distributed runtime components for large-scale training and inference. Optimize data and communication pipelines, enable scalable multi-node execution, collaborate with ML and compiler teams, diagnose performance issues via profiling, and contribute to system architecture and roadmap for cutting-edge AI workloads.
Artificial Intelligence • Hardware • Software • Semiconductor
Integrate, validate, and productionize cross-stack inference features across AI frameworks, runtime, compiler, kernels, distributed systems, and hardware. Drive zero-to-one projects, debug system-wide failures, manage accelerated timelines, and improve automation, diagnostics, and repeatable integration practices while collaborating across software and hardware teams.
Artificial Intelligence • Hardware • Software • Semiconductor
Lead a hands-on engineering team to improve kernel-centric reliability of large AI compute clusters. Own technical vision, build diagnostic and debug tooling, collaborate with SW and HW teams to reduce downtime, speed failure analysis, and mentor engineers to deliver scalable, reliable production systems.
Artificial Intelligence • Hardware • Software • Semiconductor
Implement and scale LLM training, fine-tuning, and post-training techniques (RL-based). Build evaluation and data pipelines, debug ML stack issues, optimize training/inference workflows, and ship maintainable ML infrastructure code.
Artificial Intelligence • Hardware • Software • Semiconductor
Drive end-to-end ML model inference performance: build kernel- and system-level performance models, optimize kernel microcode and compiler algorithms, debug runtime performance on system and cluster, and develop tooling to visualize and analyze performance data from the Wafer Scale Engine and compute cluster.
Artificial Intelligence • Hardware • Software • Semiconductor
Build, productionize, and optimize a GPU-based inference stack combining GPU prefill with Cerebras decode. Implement and operate model-serving APIs, vLLM/PyTorch/ROCm runtimes, deployment and reliability practices, performance profiling and optimization, cross-layer debugging, numerical validation, and benchmarking/infrastructure for production inference at scale.
