Arena (arena.ai) Logo

Arena (arena.ai)

Site Reliability Engineer

Reposted 22 Days Ago
Remote or Hybrid
Hiring Remotely in CA
Senior level
Remote or Hybrid
Hiring Remotely in CA
Senior level
Build and operate the core infrastructure for Arena's online evaluation systems: design low-latency, high-reliability APIs and gateways, implement enterprise-grade features (rate limiting, auth, metering, audit logging), instrument observability (tracing, latency, usage tracking), integrate with LLM providers and the evaluation platform, and collaborate with research and product teams to scale and harden systems for bursty, unpredictable traffic.
The summary above was generated by AI
About Arena Intelligence

Arena is the platform for evaluating how AI models perform in the real world. Founded by researchers from UC Berkeley's SkyLab, we're on a mission to measure and advance the frontier of AI for real-world use, and to build the foundation for everyone to understand, shape, and benefit from it.


Tens of millions of people use Arena each month to evaluate how frontier systems handle the work they actually do. The preferences they share power the most transparent, rigorous, and human-centered evaluations in AI. Leading AI labs, enterprises, and independent researchers rely on our work and open datasets to understand how models behave in real workflows: agentic coding, creative generation, professional productivity, and beyond. We go beyond leaderboards and decompose what human experience reveals about AI, so models advance toward the work people actually do.


We're a team of researchers, academics, builders, and creatives from UC Berkeley, Google, Stanford, and DeepMind. We seek truth, move fast, and value craftsmanship, curiosity, and impact over hierarchy. We're building a company where thoughtful, curious people from all backgrounds can do their best work together, in an office culture that radiates excellence, energy, and focus.

About the Role

Arena Intelligence is looking for an engineer to build the core infrastructure that sits beneath our online evaluation systems — the AI gateways, automated arena runtimes, and serving layers that make real-world model evaluation possible at scale.

This is a critical part of the Arena Service. Arenas are live, online systems: they route traffic across frontier models from many providers, handle bursty and unpredictable load, need to fail gracefully when upstream models do, and have to remain fair and consistent under all of it. We exist to build foundational infrastructure for our users that scales, is reliable, and makes the complexities of operating this infrastructure at scale disappear. We need a practitioner who's shipped this kind of infrastructure before and knows where the sharp edges are.

You'll be an early member of our infrastructure team, working closely with researchers, engineers, and product leadership. The work is zero-to-one in places and scale-it-up in others. We move fast and stay rigorous.

What You'll Do
  • Build and orchestrate API-based products from the ground up. Design and implement low-latency, high-reliability APIs for leaderboards, models, and arenas.

  • Ship enterprise-grade infrastructure. Build the systems enterprise customers expect: rate limiting, authentication, usage metering, cost attribution, audit logging, and SOC 2 compliance.

  • Build deep observability. Instrument infrastructure with distributed tracing, latency breakdowns, token-level usage tracking, and real-time dashboards so customers (and we) can see exactly what's happening.

  • Build AI-centered products. Integrate with our core evaluation platform, Arena data, and customer-specific benchmarks. Collaborate with the research team to turn novel ideas into full-featured products.

What We're Looking For
  • 6+ years of backend engineering experience, with meaningful time spent on distributed systems, infrastructure, or developer-facing platforms.

  • Strong proficiency in Go and/or Rust, with hands-on experience building high-throughput APIs or proxy/gateway systems.

  • Experience with LLM provider APIs (OpenAI, Anthropic, Google, etc.) and a working understanding of the challenges: streaming, token management, rate limits, model-specific quirks.

  • Solid cloud infrastructure skills — you're comfortable with AWS or GCP, Kubernetes, Terraform, and database systems like Postgres and Redis.

  • A product-oriented mindset. You think about the developer experience of your APIs, not just the implementation. You ask "why" before "how."

  • Comfort with ambiguity. We're a startup. Scope is fluid, context shifts, and you'll wear many hats. That should sound exciting, not stressful.

Nice to Have
  • Experience building API gateways, proxies, or developer tools (Bifrost, Kong, Envoy, Tyk, or custom).

  • Background in ML infrastructure, model serving, or evaluation frameworks.

  • Experience building enterprise-ready features: SSO, RBAC, audit logs, multi-tenancy.

  • Experience building billing infrastructure around systems like Stripe, Metronome and Orb

  • Familiarity with the modern AI infra stack (vLLM, LiteLLM, LangChain, etc.).

 
 

What we offer
  • We offer competitive compensation and equity aligned to the markets where our team members are based. The base salary range will depend on the candidate’s permanent work location.

  • Comprehensive health and wellness benefits, including medical, dental, vision, and additional support programs.

  • The opportunity to work on cutting-edge AI with a small, mission-driven team

  • A culture that values transparency, trust, and community impact

Come help build the space where anyone can explore and help shape the future of AI.

Arena Intelligence provides equal employment opportunities (EEO) to all employees and applicants for employment without regard to race, color, religion, sex, national origin, age, disability, genetics, sexual orientation, gender identity, or gender expression. We are committed to a diverse and inclusive workforce and welcome people from all backgrounds, experiences, perspectives, and abilities.

Similar Jobs

Yesterday
Remote
CA
Entry level
Entry level
Cloud • Information Technology • Business Intelligence • Consulting
Designs, builds, and operates cloud infrastructure and reliability capabilities for an enterprise AI platform. Responsibilities include infrastructure as code, landing zones, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, on-premises, and customer-controlled environments. This remote, client-facing consulting role requires strong communication, production ownership, and the ability to balance reliability, delivery speed, security, and operational cost.
Top Skills: AlertingAzureAzure ArcAzure DevopsBicepCi/CdCloud InfrastructureDashboardsGithub ActionsGpu WorkloadsIdentity And Access ManagementInfrastructure As CodeKubernetesLogsMetricsNetworkingObservabilityTerraformTraces
2 Days Ago
Remote or Hybrid
Junior
Junior
Healthtech • Software
Supports critical production services, monitoring, incident response, deployment validation, observability, and reliability improvements. Develops scripts, automation, AI-powered operational tools, runbooks, and workflows to reduce operational toil. Partners with development teams on reliability best practices, operational readiness, and production system understanding while learning SRE, distributed systems, and cloud operations under mentorship.
Top Skills: AppdynamicsAzureDatadogElkGrafanaKubernetesPrometheus
2 Days Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Information Technology • Consulting
Merge medical imaging solutions, offered by Merative, combine intelligent, scalable imaging workflow tools with deep and broad expertise to help healthcare organizations improve their confidence in patient outcomes and optimize care delivery. This principal-level individual contributor sets reliability, observability, and automation direction for business-critical Azure cloud platforms. The role establishes SLOs, leads incident management, designs resilience and disaster recovery, builds infrastructure automation, reduces operational toil, mentors engineers, and influences cross-functional teams without direct authority.
Top Skills: AksAnsibleArmAzure MonitorBashBicepCi/CdDicomDistributed TracingGoGrafanaHipaaHitrustHl7Iec 62304Iso 13485IstioJavaKafkaKqlKubernetesLinuxAzureMS OfficePrometheusPythonService MeshTerraform

What you need to know about the Calgary Tech Scene

Employees can spend up to one-third of their life at work, so choosing the right company is crucial, not just for the job itself but for the company culture as well. While startups often offer dynamic culture and growth opportunities, large corporations provide benefits like career development and networking, especially appealing to recent graduates. Fortunately, Calgary stands out as a hub for both, recognized as one of Startup Genome's Top 100 Emerging Ecosystems, while also playing host to a number of multinational enterprises. In Calgary, job seekers can find a wide range of opportunities.

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account