Back to search

On Running · Consumer Tech

Principal Engineer (AI & Cloud Infrastructure)

Verified·1 day ago

In short

In Technology at On, everything we build fuels our mission to ignite the human spirit through movement. We craft technology that improves with every interaction, enhancing the experience for everyone who moves. Delivering Wow through our platforms, apps, and services is a core value.

Your focus will not be on individual use cases or one-off prototypes. Instead, your focus will be AI Platform Engineering: building the paved roads, shared services, and core capabilities that make every AI initiative at On faster to build, cheaper to run, safer to operate, and easier to scale. You will turn scattered experimentation into durable, production-grade capability. You will be a hands-on Principal Engineer who can move from architecture to implementation, from proof of concept to production platform, and from a single workload to a multi-tenant capability serving the entire organisation.

You will work closely with our AI strategy, data science, machine learning, applied AI, and infrastructure leaders, ensuring that On's AI foundations are technically excellent, deeply integrated with our data and systems landscape, and built to serve athletes and customers for years to come.

What Success Looks Like:

You will be successful in this role if you turn On's AI ambition into durable engineering capability. You will build a platform that teams across On trust and choose to build on. AI initiatives that once took months of bespoke infrastructure work will ship in days on paved roads. Cost, quality, security, and governance will be managed by design rather than by exception. And as the AI landscape evolves, On's foundations will evolve with it: stable where it matters, adaptable where it counts. You will be the engineering backbone that turns AI possibility into production reality.

Your mission

Platform & Infrastructure Delivery

Build On's Core AI Platform

  • Design and deliver the shared services at the heart of On's AI capability: model gateway and orchestration layers, inference serving, retrieval infrastructure, vector and feature stores, evaluation pipelines, and the developer tooling that ties them together.

Provide Paved Roads for AI Development

  • Create golden paths, SDKs, templates, and reference architectures that allow product and applied AI teams to ship AI features quickly without reinventing infrastructure, security, or governance for each initiative.

Engineer for Production, Not Just Demos

  • Ensure that AI workloads at On meet the standards expected of any critical system: reliability, observability, latency, cost efficiency, failover, and graceful degradation. Take capabilities from promising prototype to hardened production service.

Own Model Lifecycle Infrastructure

  • Build and operate the machinery for the full model lifecycle: versioning, deployment, A/B rollout, monitoring, drift detection, evaluation, and retirement, across both third-party foundation models and internally developed models.

Optimise Cost, Performance, and Scale

  • Establish the practices and tooling to manage inference cost, token consumption, caching, routing between models, and capacity planning, so that On's AI usage scales sustainably with the business.

Engineering & Technical Leadership

Act as a Hands-on Principal Engineer

  • Lead by example through deep technical contribution. Be comfortable writing production code, designing distributed systems, reviewing critical architecture, and unblocking teams on the hardest infrastructure problems.

Set the Technical Direction for AI Infrastructure

  • Define the architecture, standards, and long-term technical roadmap for On's AI platform. Make deliberate build-vs-buy decisions across the rapidly evolving AI tooling landscape, and keep the platform coherent as it grows.

Integrate AI Into On's Systems Landscape

  • Connect the AI platform deeply with On's data platform, identity and access management, event streams, and application ecosystem, so intelligent capabilities can draw on trusted data and act safely within existing systems.

Set a High Technical Bar

  • Ensure platform components are designed with sound engineering judgement. Balance velocity with quality, simplicity, security, maintainability, and long-term scalability, and raise the bar for AI engineering practice across the organisation.

Navigate Ambiguity With Pragmatism

  • Work in a space where the ecosystem changes monthly. Use technical judgement, benchmarking, and first-principles thinking to make durable architectural decisions in a fast-moving landscape.

Enablement, Governance & Adoption

Enable Applied AI Teams

  • Partner closely with applied AI, product, and functional teams as your customers. Understand their needs, remove friction, and shape the platform roadmap around what accelerates them most.

Build Governance Into the Platform

  • Embed security, privacy, data protection, access control, auditability, and responsible AI guardrails directly into platform primitives, so doing the right thing is the default, not an afterthought.

Establish Evaluation & Quality Infrastructure

  • Provide the shared evaluation frameworks, benchmarks, regression suites, and observability that let teams measure model and agent quality objectively, and give leadership confidence in what is deployed.

Partner With the AI Circle and AI Kitchen

  • Contribute actively to On's AI operating rhythm, including the AI Kitchen and AI Circle. Represent the platform perspective in prioritisation discussions and help turn strategic ambition into robust, scalable execution.

Influence Without Authority

  • Operate as a senior individual contributor who earns trust through expertise, clarity, delivery, and collaboration. Guide teams through professional influence rather than formal management authority.

Innovation & Continuous Improvement

Scout Emerging AI Infrastructure

  • Continuously evaluate new serving frameworks, orchestration tools, model providers, agent runtimes, and infrastructure patterns, assessing where they can strengthen On's platform.

Evolve the Platform From First Principles

  • Avoid accumulating tooling for its own sake. Continuously simplify, consolidate, and redesign the platform as the ecosystem matures, retiring what no longer earns its place.

Industrialise What Works

  • Identify when a capability proven by applied teams should be absorbed into the platform as a shared service, and lead that transition from bespoke solution to reusable foundation.

Promote Responsible AI Use

  • Ensure platform capabilities are developed and operated with appropriate consideration for security, privacy, data protection, governance, cost, and brand trust.

Mentor and Inspire

  • Support engineers across On in building strong AI infrastructure skills. Share practical knowledge on distributed systems, MLOps, and LLMOps, and help others develop the judgement to build AI systems well.

Your story

Technical Background

  • You hold a Bachelor's or Master's degree in Computer Science, Machine Learning, Software Engineering, Data Science, or a related technical field, or equivalent practical experience.

Principal-Level Engineering Experience

  • You have 10+ years of experience designing and building high-quality, large-scale software systems, with a proven track record as a senior or principal-level engineer operating across complex organisations.

AI/ML Infrastructure Practitioner

  • You have hands-on experience building the infrastructure behind AI systems: model serving and inference optimisation, retrieval and vector infrastructure, orchestration and agent runtimes, evaluation pipelines, MLOps/LLMOps tooling, and the integration of foundation model APIs into production systems.

Distributed Systems Engineer

  • You are deeply comfortable with cloud-native architecture, Kubernetes and containerised workloads, event-driven systems, API design, CI/CD, infrastructure as code, and the operational disciplines: observability, SLOs, incident response, that keep critical platforms healthy.

Platform Builder

  • You have built platforms or internal developer tooling that other engineering teams depend on. You think in terms of paved roads, self-service, multi-tenancy, and developer experience, and you measure your success by the teams you accelerate.

Systems Thinker

  • You understand how data, platforms, workflows, and business processes connect. You can design shared capabilities that create leverage across an entire ecosystem rather than solving one problem at a time.

Pragmatic Problem Solver

  • You are excited by complexity, but you do not add unnecessary complexity yourself. You favour simple, elegant, high-leverage architectures that can be understood, adopted, and evolved.

Technology Translator

  • You can speak equally well with engineers, executives, product managers, and business stakeholders. You make infrastructure trade-offs, costs, and risks understandable, relevant, and actionable.

Impact-Oriented Mindset

  • You care deeply about outcomes. You are not satisfied with elegant infrastructure that nobody uses. You want to see the platform measurably accelerate how On builds, ships, and operates AI.

Exceptional Communicator

  • Fluent in English, you can articulate complex technical ideas clearly and persuasively, whether in an architecture review, executive conversation, design document, or cross-functional workshop.