Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in Calgary
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead design, automation, and maintenance of cloud-based database infrastructure (primarily SQL Server and MySQL). Improve reliability with monitoring, HA/DR, automation, troubleshooting, on-call support, and mentoring of junior engineers while collaborating across teams.
Top Skills:
AuroraAWSBashFailover ClusteringMySQLNew RelicOrchestratorPmmPythonRdsRubySQL ServerVividcortex
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
The Site Reliability Engineer will manage cloud infrastructure, automate tasks, collaborate in agile teams, and ensure service reliability and quality.
Top Skills:
Aurora MysqlAWSAzureBashDockerGCPGoKubernetesPostgresPythonRubyTerraform
Big Data • Information Technology • Software • Analytics • Energy
Manage and scale Enverus' global AWS infrastructure, automate deployments and CI/CD, ensure high uptime, collaborate with developers to enable zero-downtime releases, participate in on-call rotations, and improve operational practices.
Top Skills:
AWSAzureC#Ci/CdCloudFormationGoKubernetesLinuxPythonTerraformWindows
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Lead high-impact reliability projects to improve resiliency, scalability, and deployment safety for thousands of services. Build secure configuration and secrets systems, improve canary-based release systems, partner with critical teams to reduce operational toil, drive observability and reliability best practices, participate in on-call rotations, and communicate architecture decisions to stakeholders.
Top Skills:
AWSAzureDatadogGCPGenerative AiGoKibanaRubyTerraform
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Reposted 5 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Fintech • Mobile • Payments • Financial Services
Design and build a centralized reliability platform for production systems, integrating distributed-systems engineering with AI-assisted tools. Implement AI agents for incident triage, log/trace summarization, and developer-facing APIs. Own projects end-to-end and collaborate with product, infra, data, and SRE teams to iterate and improve system health and debuggability.
Top Skills:
ClaudeCursorDistributed SystemsGithub CopilotLlmsPython
Payments • Software
Own and scale payments partner integrations and reporting; automate manual processes; build tooling and anomaly detection; collaborate with engineers, accounting, and partners to trace and reconcile large-scale money flows; troubleshoot and implement code changes.
Top Skills:
GitJavaRubySQL
13 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Maintain and improve reliability, scalability, and automation for user-facing production systems. Build infrastructure tooling, operate Kubernetes-based services, write IaC, participate in on-call and incident response, and advance observability and runbooks to reduce toil and improve platform reliability.
Top Skills:
AWSCi/CdGCPGitopsGoInfrastructure As Code (Iac)KubernetesKubernetes Operators/ControllersLoggingMetricsRubySlos/SlisTerraform
Information Technology • Software • Database • Automation
Owner of on-prem reliability and escalations: reproduce and resolve L2/L3 issues across heterogeneous Kubernetes environments, build diagnostics and automation, improve CI and e2e test stability, establish performance baselines, harden install/upgrade flows, and write tooling in Python/Go/Rust to reduce repeat incidents.
Top Skills:
BenchmarkingCiCi/CdContainersE2E TestingGoHealth ChecksHelmInstallersIntegration TestingKubernetesLoad GenerationLogsMetricsNetworkingObservabilityPackagingProfilingPythonRbacRustStorageSupport BundlesTraces
Insurance
As a Reliability Engineer, you'll design, implement, and maintain AWS cloud environments, ensuring systems' reliability and performance, while enhancing monitoring and incident response capabilities.
Top Skills:
AWSNginxPythonUnixWindows
Artificial Intelligence • Software • Generative AI
The Founding Platform & Reliability Engineer will design and operate reliable, scalable infrastructure for an AI storytelling platform, involving hands-on implementation and strategic decision-making.
Top Skills:
AmplitudeAWSCloud RunFirebaseGCPModalNext.JsNode.jsPythonReactRedisSentryTypescriptUpstash
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure, Terraform IaC, Kubernetes/EKS clusters, CI/CD, and observability for a platform serving millions. Lead incident response, improve developer tooling, and set reliability standards across engineering teams.
Top Skills:
AWSCdnCircleCICloud NetworkingCronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Information Technology • Software • Database
As a Site Reliability Engineer, you will design, implement, and maintain scalable infrastructure, ensure system reliability, automate processes, and collaborate with engineering teams.
Top Skills:
DockerElk StackGoGrafanaJavaKubernetesNode.jsPrometheusPulumiPythonRubyTerraform
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Lead reliability and implementation for Kong's Managed Gateways: design and operate multi-cloud, Kubernetes-based systems, own incident response and SLOs, automate CI/CD and IaC, mentor SREs, and drive enterprise customer onboarding and technical implementations.
Top Skills:
AnsibleAWSAzureCassandraCi/CdDatadogElkGCPGoGrafanaIstioKubernetesLinkerdPostgresPrometheusTerraform
Information Technology • Software
The Reliability Engineer will develop and implement reliability test plans for IVD medical devices, conduct analyses, and lead cross-functional projects while ensuring compliance with regulatory standards.
Top Skills:
JmpMatlabMinitabPythonR
Aerospace
Lead reliability strategy and environmental validation for satellite user terminal hardware. Design and run HALT/HASS/ESS tests, oversee weatherproofing and thermal/vibration testing, perform failure analysis (X-ray, microscopy, cross-sectioning), own DFMEA and physics-of-failure models to predict MTBF and warranty risk, collaborate with electrical and mechanical teams and external labs to drive hardware improvements and regulatory compliance.
Top Skills:
CfdCross-SectioningDfmeaElectronic Thermal Cycle SystemsEnvironmental ChambersEssFeaHaltHassIp67JmpMicroscopyMinitabPower Delivery Network (Pdn) AnalysisReliasoftSalt Fog Corrosion TestingUv Exposure TestingVibration/Shaker TablesWeibull Life-Data AnalysisX-Ray Inspection
Software
The Senior Site Reliability Engineer will lead service onboarding, maintain SLAs/SLOs, design secure infrastructure, automate operational tasks, and respond to incidents while ensuring system reliability and performance.
Top Skills:
AWSCloudFormationElk StackGoGrafanaHadoopKubernetesPythonTerraform
Artificial Intelligence • Other • Sales • Software
Design and advance core infrastructure for engineering, ensure Kubernetes reliability, automate operations, and support AI infrastructure.
Top Skills:
AWSAzureCi/CdCloudFormationGitopsGoGCPHelmKubernetesKustomizePostgresPythonTerraform
Fintech • Payments
The Senior Staff SRE leads reliability engineering initiatives, drives operational excellence, mentors staff, and influences architecture to enhance system reliability and performance.
Top Skills:
Ai/MlAWSAzureDockerElk StackGCPGrafanaKubernetesMySQLNoSQLPostgresSplunk
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
The role involves operating and scaling Kong's SaaS platform, building automated infrastructure, optimizing multi-region data layers, enhancing observability, and ensuring reliability across services.
Top Skills:
ArgocdAWSAzureBashClickhouseDatadogDruidGCPGoGrafanaHelmKubernetesPostgresPrometheusPythonRedisTerraformTerragruntThanos
Database • Analytics
The Senior Site Reliability Engineer will ensure reliability and scalability of cloud infrastructure, enhance incident management, and optimize operational efficiencies through collaboration with various teams.
Top Skills:
AnsibleAWSAzureDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Artificial Intelligence • Information Technology • Software
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
Top Skills:
Anthropic ApiAWSDistributed TracingGCPGoGoogle Llm ApisKubernetesOpenai ApiPostgresRedisRustTerraform
Transportation
Design and develop Waabi's observability stack, optimize performance, build automation tooling, and support application requirements while leading projects and mentoring teams.
Top Skills:
AWSC/C++DockerGoGrafanaJavaKubernetesOpentelemetryPythonRust
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
Tech Jobs & Startup Jobs in Calgary
Remote Jobs in Calgary
Hybrid Jobs in Calgary
Account Executive Jobs in Calgary
Account Manager Jobs in Calgary
Accounting Jobs in Calgary
AI Jobs in Calgary
Analyst Jobs in Calgary
Analytics Jobs in Calgary
Automation Engineer Jobs in Calgary
AWS Jobs in Calgary
Azure Jobs in Calgary
Business Analyst Jobs in Calgary
Business Development Jobs in Calgary
Cloud Jobs in Calgary
Communications Jobs in Calgary
Content Writer Jobs in Calgary
Controller Jobs in Calgary
Copywriting Jobs in Calgary
Customer Service Jobs in Calgary
Customer Service Manager Jobs in Calgary
Cyber Security Jobs in Calgary
Data Analyst Jobs in Calgary
Data Engineer Jobs in Calgary
Data Jobs in Calgary
Data Science Jobs in Calgary
Database Administrator Jobs in Calgary
Design Jobs in Calgary
DevOps Jobs in Calgary
Engineering Jobs in Calgary
Engineering Manager Jobs in Calgary
Executive Assistant Jobs in Calgary
Finance Jobs in Calgary
Finance Manager Jobs in Calgary
Financial Analyst Jobs in Calgary
Front End Developer Jobs in Calgary
Full Stack Developer Jobs in Calgary
Graphic Design Jobs in Calgary
HR Jobs in Calgary
HR Manager Jobs in Calgary
IT Jobs in Calgary
IT Support Jobs in Calgary
Java Developer Jobs in Calgary
Legal Counsel Jobs in Calgary
Legal Jobs in Calgary
Linux Jobs in Calgary
Machine Learning Jobs in Calgary
Marketing Jobs in Calgary
Marketing Manager Jobs in Calgary
NET Jobs in Calgary
Network Engineer Jobs in Calgary
Operations Jobs in Calgary
Operations Manager Jobs in Calgary
Outside Sales Jobs in Calgary
Payroll Jobs in Calgary
Product Manager Jobs in Calgary
Product Owner Jobs in Calgary
Program Manager Jobs in Calgary
Project Engineer Jobs in Calgary
Project Manager Jobs in Calgary
Python Developer Jobs in Calgary
Quality Assurance Jobs in Calgary
Quality Engineer Jobs in Calgary
Recruiter Jobs in Calgary
Reliability Engineer Jobs in Calgary
Research Jobs in Calgary
Sales Jobs in Calgary
Sales Manager Jobs in Calgary
Sales Rep Jobs in Calgary
SEO Jobs in Calgary
Software Engineer Jobs in Calgary
Software Testing Jobs in Calgary
Staff Accountant Jobs in Calgary
Talent Acquisition Jobs in Calgary
Tax Jobs in Calgary
Technical Support Jobs in Calgary
UX Designer Jobs in Calgary
Web Developer Jobs in Calgary
Writing Jobs in Calgary
All Filters
Total selected ()
No Results
No Results



.png)












.png)


















