Staff Site Reliability Engineer, Ads

Reddit·San Francisco, CA·onsite
crypto:applicationengineeringIC6BE Platform
Compensation
Not disclosed
Reddit is a community of communities. It’s built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. Every day, Reddit users submit, vote, and comment on the topics they care most about. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information. For more information, visit www.redditinc.com . This role is remote friendly. Reddit has a flexible first workforce The Ads organization powers Reddit's advertising platform, enabling advertisers to reach highly engaged communities while helping Reddit grow its business. The reliability of our Ads systems directly impacts advertiser success, revenue generation, and user experience. The Ads Reliability team partners closely with Ads Engineering teams to improve reliability, scalability, operational excellence, and developer productivity across Reddit's advertising ecosystem. We're looking for a Staff Site Reliability Engineer who will define and provide technical leadership for reliability initiatives across the Ads organization and help shape the future of Ads infrastructure at Reddit. What you’ll do: Lead reliability initiatives across multiple Ads domains including ad serving, auctions, targeting, reporting, measurement, and billing. Partner with engineering leadership to develop a roadmap to improve reliability, scalability, operational excellence, and engineering efficiency across the Ads organization. Design and build platforms, tooling, and automation that improve reliability and developer productivity at scale. Drive architecture reviews and influence technical decisions impacting critical revenue-generating systems. Participate in on-call rotations, lead complex incident investigations and coordinate cross-functional response efforts during major production events. Identify systemic reliability risks and drive long-term solutions that improve plat