Lead Site Reliability Engineer
Apply on JP Morgan Chase's site
Description -----------
Assume a critical role in defining the future of a globally recognized firm and have a direct and significant effect in a realm tailored for top achievers in site reliability.
As a Lead Site Reliability Engineer at JPMorgan Chase within the Corporate & Investment Bank (CIB) Management and Support Functions Digital & Platform Services team, you hold a leadership role in your team, demonstrate strong knowledge across multiple technical domains, and advise others on the technical and business issues facing them.
Job responsibilities
- Define the Production Management SRE vision, north-star outcomes, and multi‑year roadmap aligned to CIB and JPMorganChase Global Technology priorities, and establish the global SRE operating model (ways of working, intake, prioritization, and engagement with engineering/production support).
- Build and scale a high-impact SRE capability by developing a small core team and/or a global community of practice, and by partnering with business-aligned Production Support leads to embed reliability practices and “engineer out” operational load.
- Set and implement firmwide reliability standards and patterns across service cataloging, SLO/SLI and error budgets, incident response maturity, blameless post-incident reviews, resiliency patterns, and capacity/performance/scalability engineering.
- Drive evidence-based service health and reliability governance through regular service reviews and reporting (availability, latency, incident trends, MTTR/MTTD, change failure rate, and customer impact), and continuously improve observability and alerting quality (logs/metrics/traces, golden signals, end-user journey monitoring, actionable routing, and reduced false positives).
- Champion enterprise-authorized AI adoption to reduce operational toil and improve incident response (triage, troubleshooting, post-incident analysis), and lead reuse-first AI-assisted reliability workflows across the SDLC/toolchain (CI/CD quality checks, automation, operational readiness) with traceability/auditability and required resiliency/security controls.
- Uses enterprise-authorized AI capabilities within the work environment to accelerate major-incident triage, troubleshooting, and post-incident analysis, validating outputs and handling operational data according to sensitivity and security requirements.
- Leads reuse-first adoption of AI-assisted reliability workflows across SDLC/toolchain practices (e.g., CI/CD quality checks, test/validation automation, and operational readiness), ensuring traceability/auditability, resiliency, and security controls.
Required qualifications, capabilities, and skills
- Formal training or certification on site reliability engineering concepts and 5+ years applied experience ( NAMR/APAC – India/ LATAM/ Hong Kong)
- Demonstrated experience leading SRE/reliability engineering or production engineering transformations in a complex enterprise environment.
- Strong engineering background: ability to design, build, and deliver automation and reliability solutions.
- Fluency & expertise in Python
- Deep practical knowledge of: SLOs/SLIs, error budgets, incident management, postmortems, observability design across metrics/logs/traces and distributed systems troubleshooting, resilience engineering, performance/capacity management, and change risk reduction.
- Proficiency and experience with telemetry (logs/metrics/traces) collection using tools and standards such as Prometheus, OpenTelemetry, Datadog, Dynatrace, Splunk.
- Experience delivering automation at scale (scripting, workflow automation, runbook automation, CI/CD-integrated guardrails).
- Proven leadership skills: influencing without authority, coaching leaders, and building communities of practice.
- Strong judgment around risk, security, and controls—especially when applying AI to production workflows.
- Demonstrated experience using enterprise-authorized AI capabilities within the work environment to improve SRE workflows (e.g., incident investigation support and knowledge capture) with strong validation habits and awareness of data sensitivity.
- Ability to evaluate AI-assisted operational recommendations for correctness and risk, define appropriate guardrails for team usage, and ensure outcomes align to resiliency and security expectations.
Preferred qualifications, capabilities, and skills
- Proficient with container and container orchestration
- Experience with troubleshooting common networking technologies and issues
- Advanced knowledge of software applications and technical processes with emerging depth in one or more technical disciplines, and actively self-educates to evaluate and recommend suitable new technologies
- Familiarity with Athena / prior experience in Athena
About JP Morgan Chase
JP Morgan Chase is a global financial services provider that offers investment banking, asset management, treasury, and other services.
More jobs at JP Morgan Chase
- Product Delivery Manager- Asset Based Lending Wholesale Lending Services
- Lead Software Engineer — Workflow Orchestration (React | Camunda | Java/Python) Full Stack
- Lead Software Engineer (SWES04)- SDET with Playwright
- Vice President, Strategic Program Manager - Affluent Program Office
- Lead Software Engineer- Data Engineer/Pyspark/Databricks
- Full-Stack Java/Python React Software Engineer III - Trading Platform
- Sr. Lead Software Engineer: Data Engineering
- Software Engineer III - Automation Engineer
- Site Reliability Engineer III
- Sr Director of Security Engineering