Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in Calgary
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead design, automation, and maintenance of cloud-based database infrastructure (primarily SQL Server and MySQL). Improve reliability with monitoring, HA/DR, automation, troubleshooting, on-call support, and mentoring of junior engineers while collaborating across teams.
Top Skills:
AuroraAWSBashFailover ClusteringMySQLNew RelicOrchestratorPmmPythonRdsRubySQL ServerVividcortex
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
The Site Reliability Engineer will manage cloud infrastructure, automate tasks, collaborate in agile teams, and ensure service reliability and quality.
Top Skills:
Aurora MysqlAWSAzureBashDockerGCPGoKubernetesPostgresPythonRubyTerraform
Big Data • Information Technology • Software • Analytics • Energy
Manage and scale Enverus' global AWS infrastructure, automate deployments and CI/CD, ensure high uptime, collaborate with developers to enable zero-downtime releases, participate in on-call rotations, and improve operational practices.
Top Skills:
AWSAzureC#Ci/CdCloudFormationGoKubernetesLinuxPythonTerraformWindows
Reposted 8 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Maintain and improve reliability, scalability, and automation for user-facing production systems. Build infrastructure tooling, operate Kubernetes-based services, write IaC, participate in on-call and incident response, and advance observability and runbooks to reduce toil and improve platform reliability.
Top Skills:
AWSCi/CdGCPGitopsGoInfrastructure As Code (Iac)KubernetesKubernetes Operators/ControllersLoggingMetricsRubySlos/SlisTerraform
9 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Fintech • Mobile • Payments • Financial Services
Design and build a next-generation reliability platform for production systems, lead delivery of quarterly goals, collaborate across product/design/analytics, support operations and on-call, set code and design standards, mentor engineers, and integrate AI/LLM-assisted tooling to improve service health and debugging.
Top Skills:
Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
News + Entertainment
Build and maintain automated test suites across mobile and backend, integrate quality gates into CI/CD, monitor production systems, participate in incident response and root-cause fixes, and embed quality practices across the engineering team using AI-assisted tools.
Top Skills:
Ai ToolsAi-Assisted Qa ToolsAppiumCi/CdError-BudgetEspressoObservability/Monitoring ToolsPlaywrightSloXctest
Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency
The Senior Site Reliability Engineer will enhance reliability of Block's platform, improve incident response using AI tools, and coordinate incident management. Responsibilities include building reliable systems, standardizing tools, and leading high-severity incidents during on-call rotations.
Top Skills:
Amazon Web ServicesDatadogDynamoDBGrpcHTTPIstioJavaJSONKotlinKubernetesLaunchdarklyMySQLProtocol BuffersTerraformVitess
Software
Lead and resolve high-severity post-sales technical escalations through deep-dive investigation, reproduction, and engineering handoff. Own customer calls, collect logs/evidence, file bug reports, monitor patterns, and improve tooling and runbooks in partnership with Support and Engineering.
Top Skills:
Ci/CdCloud IaasContainerized ApplicationsDnsFirewallsGoJSONKubernetesLinuxLoad BalancersmacOSNatPacket-Level AnalysisPprofPythonRoutingStunTcp/IpUdp Hole PunchingVpnWindowsWireguard
Artificial Intelligence • Blockchain • Fintech • Financial Services • Cryptocurrency • NFT • Web3
Lead design and delivery of reliability projects for Coinbase's platform: secure service configuration and secrets management, improve canary-based deployments, partner with core services to increase scalability and reduce incidents, drive reliability best practices, and participate in on-call rotations.
Top Skills:
AWSAzureDatadogGCPGoKibanaRubyTerraform
Information Technology • Software • Database • Automation
Owner of on-prem reliability and escalations: reproduce and resolve L2/L3 issues across heterogeneous Kubernetes environments, build diagnostics and automation, improve CI and e2e test stability, establish performance baselines, harden install/upgrade flows, and write tooling in Python/Go/Rust to reduce repeat incidents.
Top Skills:
BenchmarkingCiCi/CdContainersE2E TestingGoHealth ChecksHelmInstallersIntegration TestingKubernetesLoad GenerationLogsMetricsNetworkingObservabilityPackagingProfilingPythonRbacRustStorageSupport BundlesTraces
Insurance
As a Reliability Engineer, you'll design, implement, and maintain AWS cloud environments, ensuring systems' reliability and performance, while enhancing monitoring and incident response capabilities.
Top Skills:
AWSNginxPythonUnixWindows
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure, Terraform IaC, Kubernetes/EKS clusters, CI/CD, and observability for a platform serving millions. Lead incident response, improve developer tooling, and set reliability standards across engineering teams.
Top Skills:
AWSCdnCircleCICloud NetworkingCronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Information Technology • Software
The Reliability Engineer will develop and implement reliability test plans for IVD medical devices, conduct analyses, and lead cross-functional projects while ensuring compliance with regulatory standards.
Top Skills:
JmpMatlabMinitabPythonR
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills:
AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Database • Analytics
The Senior Site Reliability Engineer will ensure reliability and scalability of cloud infrastructure, enhance incident management, and optimize operational efficiencies through collaboration with various teams.
Top Skills:
AnsibleAWSAzureDocker SwarmGoGoogle Cloud PlatformKubernetesPuppetPythonTerraform
Payments • Software
Own and scale payments partner integrations and reporting; automate manual processes; build tooling and anomaly detection; collaborate with engineers, accounting, and partners to trace and reconcile large-scale money flows; troubleshoot and implement code changes.
Top Skills:
GitJavaRubySQL
Software
Partner with product and platform teams to improve reliability, observability, and developer autonomy. Design for reliability, establish standards, build shared tooling, mentor teams on SRE practices, and lead incident response and service health initiatives to scale platform resilience.
Top Skills:
Cloud-Native PlatformsInfrastructure-As-CodeNode.jsTypescript
Blockchain • Financial Services • Cryptocurrency • Web3
Operate and improve a shared telemetry platform for metrics, logs, traces, alerting, dashboards, and profiling. Maintain collection, storage, querying, routing, and pipelines; troubleshoot production telemetry issues; deploy and manage services with IaC and container orchestration; build automation for dashboards and alerts; participate in incident response, on-call, and runbook documentation.
Top Skills:
AlertmanagerAWSCi/CdClaude (Ai Tools/Agents)ConsulContainer OrchestrationGrafanaGrafana AlloyKubernetesLogqlLokiNomadOpentelemetryPrometheusPromqlPyroscopeSplunkTempoTerraformTerragruntVaultVectorVictoriametrics
Artificial Intelligence • Information Technology • Software
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
Top Skills:
Anthropic ApiAWSDistributed TracingGCPGoGoogle Llm ApisKubernetesOpenai ApiPostgresRedisRustTerraform
Aerospace • Hardware • Software • Defense • Manufacturing
As a Site Reliability Engineer, you'll ensure robotics system reliability, build telemetry integration, and develop tools for diagnostics and automation, collaborating with engineering teams for enhanced production reliability.
Top Skills:
C++DatadogGoKubernetesOpentelemetryPrometheusPythonRos2TelegrafTypescript
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Transportation
Design and develop Waabi's observability stack, optimize performance, build automation tooling, and support application requirements while leading projects and mentoring teams.
Top Skills:
AWSC/C++DockerGoGrafanaJavaKubernetesOpentelemetryPythonRust
Information Technology • Cybersecurity
Lead global SRE team to design, operate, and optimize scalable cloud infrastructure (AWS/Azure/GCP). Own reliability, observability, incident response, IaC, CI/CD, cost optimization (FinOps), security hygiene, and mentor engineers while contributing hands-on and researching AI-driven SRE tooling.
Top Skills:
AlloyAWSAzureCi/CdCloudFormationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkTerraform
Information Technology • Security • Cybersecurity
Lead design, build, and scale of secure, multi-tenant Kubernetes infrastructure and CI/CD systems. Own AI tooling infrastructure and MCP servers, implement GitOps and IaC, run streaming analytics (Kafka, Flink, ClickHouse), improve observability and SLOs, lead incident response and postmortems, and mentor engineering teams on reliability, automation, and progressive delivery strategies.
Top Skills:
Ai/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesMcp ServersOpentelemetryPrometheusPulumiPythonTerraform
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills:
Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Calgary Companies Hiring Reliability Engineers
See AllPopular Job Searches
Tech Jobs & Startup Jobs in Calgary
Remote Jobs in Calgary
Hybrid Jobs in Calgary
Account Executive Jobs in Calgary
Account Manager Jobs in Calgary
Accounting Jobs in Calgary
AI Jobs in Calgary
Analyst Jobs in Calgary
Analytics Jobs in Calgary
Automation Engineer Jobs in Calgary
AWS Jobs in Calgary
Azure Jobs in Calgary
Business Analyst Jobs in Calgary
Business Development Jobs in Calgary
Cloud Jobs in Calgary
Communications Jobs in Calgary
Content Writer Jobs in Calgary
Controller Jobs in Calgary
Copywriting Jobs in Calgary
Customer Service Jobs in Calgary
Customer Service Manager Jobs in Calgary
Cyber Security Jobs in Calgary
Data Analyst Jobs in Calgary
Data Engineer Jobs in Calgary
Data Jobs in Calgary
Data Science Jobs in Calgary
Database Administrator Jobs in Calgary
Design Jobs in Calgary
DevOps Jobs in Calgary
Engineering Jobs in Calgary
Engineering Manager Jobs in Calgary
Executive Assistant Jobs in Calgary
Finance Jobs in Calgary
Finance Manager Jobs in Calgary
Financial Analyst Jobs in Calgary
Front End Developer Jobs in Calgary
Full Stack Developer Jobs in Calgary
Graphic Design Jobs in Calgary
HR Jobs in Calgary
HR Manager Jobs in Calgary
IT Jobs in Calgary
IT Support Jobs in Calgary
Java Developer Jobs in Calgary
Legal Counsel Jobs in Calgary
Legal Jobs in Calgary
Linux Jobs in Calgary
Machine Learning Jobs in Calgary
Marketing Jobs in Calgary
Marketing Manager Jobs in Calgary
NET Jobs in Calgary
Network Engineer Jobs in Calgary
Operations Jobs in Calgary
Operations Manager Jobs in Calgary
Outside Sales Jobs in Calgary
Payroll Jobs in Calgary
Product Manager Jobs in Calgary
Product Owner Jobs in Calgary
Program Manager Jobs in Calgary
Project Engineer Jobs in Calgary
Project Manager Jobs in Calgary
Python Developer Jobs in Calgary
Quality Assurance Jobs in Calgary
Quality Engineer Jobs in Calgary
Recruiter Jobs in Calgary
Reliability Engineer Jobs in Calgary
Research Jobs in Calgary
Sales Jobs in Calgary
Sales Manager Jobs in Calgary
Sales Rep Jobs in Calgary
SEO Jobs in Calgary
Software Engineer Jobs in Calgary
Software Testing Jobs in Calgary
Staff Accountant Jobs in Calgary
Talent Acquisition Jobs in Calgary
Tax Jobs in Calgary
Technical Support Jobs in Calgary
UX Designer Jobs in Calgary
Web Developer Jobs in Calgary
Writing Jobs in Calgary
All Filters
Total selected ()
No Results
No Results








.png)



























