Senior Site Reliability Engineer
crypto:applicationengineeringIC510402 Engineering - Product Platform Engineering
Compensation
Not disclosed
Block is one company built from many blocks, all united by the same purpose of economic empowerment. The blocks that form our foundational teams — People, Finance, Counsel, Hardware, Information Security, Platform Infrastructure Engineering, and more — provide support and guidance at the corporate level. They work across business groups and around the globe, spanning time zones and disciplines to develop inclusive People policies, forecast finances, give legal counsel, safeguard systems, nurture new initiatives, and more. Every challenge creates possibilities, and we need different perspectives to see them all. Bring yours to Block.
The Role
As a member of the SRE team, you will proactively and reactively improve the reliability of Block's platform and critical infrastructure. You are metrics-driven, systems-oriented, and focused on building distributed platforms that enable safe, scalable product development.
You will leverage and continuously improve AI-driven tooling and automation to enhance observability, accelerate incident detection and response, and reduce operational toil. This includes applying AI to incident analysis, alert tuning, and operational workflows.
You will participate in primary platform oncall (12 hours per day, one week every few weeks, depending on team size), supporting Block's most critical (Tier 0) services. In this role, you will lead incident command, coordinate mitigation, and drive effective escalation during high-severity events.
You Will
Build and extend platforms to improve system reliability
Work on team goals that encompass reliability for the entire company
Standardize reliability tools across multiple platforms and organizations
Triage, coordinate, and lead stabilization of sev 0–1 incidents
Serve as primary oncall, maintaining structured escalation paths and exercising leadership escalation
Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns
Use AI-driven systems to