Clarifai Logo

Clarifai

Senior Site Reliability Engineer

Posted 9 Days Ago
Be an Early Applicant
Canada
Senior level
Canada
Senior level
As a Senior Site Reliability Engineer, you will ensure the high availability of Clarifai's core services, monitor system performance, optimize reliability, develop resources for Kubernetes, and design scalable infrastructure solutions. You will collaborate with teams to resolve engineering challenges in both cloud and on-premise environments.
The summary above was generated by AI

Senior Site Reliability EngineerAbout the Company

Clarifai is a leading, full-lifecycle deep-learning AI platform for computer vision, natural language processing, LLM and audio recognition. We help organizations transform unstructured images, video, text, and audio data into structured data at a significantly faster and more accurate rate than humans would be able to do on their own.  Founded in 2013 by Matt Zeiler, Ph.D. Clarifai has been a market leader in AI since winning the top five places in image classification at the 2013 ImageNet Challenge. Clarifai continues to grow with employees remotely based throughout the United States, Canada, Argentina, India and Estonia. 

We have raised $100M in funding to date, with $60M coming from our most recent Series C, and are backed by industry leaders like Menlo Ventures, Union Square Ventures, Lux Capital, New Enterprise Associates, LDV Capital, Corazon Capital, Google Ventures, NVIDIA, Qualcomm and Osage.

Clarifai is proud to be an equal opportunity workplace dedicated to pursuing, hiring, and retaining a diverse workforce.

Your Impact

Clarifai’s platform is a kubernetes-native distributed system that requires the orchestration of many components. Efficiently serving and training large neural networks presents unique design and infrastructure challenges. 

You will be critical to solving these challenges both in the context of the cloud and in on premise environments. Additionally, you will be responsible for our broader cloud infrastructure and development tools and environments.

The Opportunity

  • Ensure the smooth operation and high availability of Clarifai's core services
  • Monitor system performance, identify bottlenecks, and implement optimizations to enhance reliability and efficiency
  • Develop Kubernetes resources and custom tooling for seamless cloud and on-premise deployments
  • Design and implement scalable, secure, and cost-effective infrastructure solutions.
  • Partner with teams across the organization to identify & solve engineering challenges

Requirements

  • BS/BA in Computer Science or related degree
  • Good knowledge of cloud providers (AWS, GCP or similar)
  • Expertise with Kubernetes (EKS, GKE, self-hosted) and Infrastructure as Code using Terraform, Helm
  • Solid understanding of web and networking (HTTP, TLS, DNS, Certificates, etc)
  • Experience with CI/CD pipelines using tools such as GitHub Actions, ArgoCD, and Atlantis
  • Strong interpersonal skills working with teams across different time zones and regions

Great to Have

  • Knowledge of basic Microservice Architecture principles
  • Familiarity with security best practices for cloud-based systems.
  • Experience with relational databases, message queues, key value stores
  • Experience writing python, golang, or any other popular programming language
  • Familiarity with any RPC framework
  • Experience developing & building custom Kubernetes operators

Top Skills

Go
Python

Similar Jobs

Be an Early Applicant
9 Days Ago
Vancouver, BC, CAN
Hybrid
6,500 Employees
Senior level
6,500 Employees
Senior level
Gaming • Information Technology • Mobile • Software
The Senior Site Reliability Engineer will support infrastructure, monitoring, and tooling needs. Responsibilities include developing automated scalable cloud infrastructure, monitoring systems, diagnosing technical issues, and enhancing CI/CD pipelines. The SRE will collaborate with engineers to maintain highly available services across various technologies, while participating in an on-call rotation for live service issues.
Be an Early Applicant
5 Hours Ago
Toronto, ON, CAN
20,000 Employees
Senior level
20,000 Employees
Senior level
Food • Retail • Agriculture • Manufacturing
The Sr Engineering Manager, SRE & Observability will lead the design, implementation, and monitoring of secure, fault-tolerant SRE and Observability infrastructure. Responsibilities include developing strategies, collaborating with teams, mentoring engineers, and driving operational excellence through advanced monitoring and automation techniques.
Be an Early Applicant
2 Days Ago
Waterloo, ON, CAN
Hybrid
1,902 Employees
Senior level
1,902 Employees
Senior level
Fintech • Software
As a Senior Site Reliability Engineer at Carta, you will build and scale internal platforms, design monitoring systems, and collaborate with software engineers to ensure application reliability and performance. You will also drive improvements in global infrastructure systems.

What you need to know about the Calgary Tech Scene

Employees can spend up to one-third of their life at work, so choosing the right company is crucial, not just for the job itself but for the company culture as well. While startups often offer dynamic culture and growth opportunities, large corporations provide benefits like career development and networking, especially appealing to recent graduates. Fortunately, Calgary stands out as a hub for both, recognized as one of Startup Genome's Top 100 Emerging Ecosystems, while also playing host to a number of multinational enterprises. In Calgary, job seekers can find a wide range of opportunities.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account