Description
At JPMorganChase we understand that customers seek exceptional value and a seamless experience from a trusted financial institution. Thats why we launched Chase UK to transform digital banking with intuitive and enjoyable customer journeys. With a strong foundation of trust established by millions of customers in the US we have been rapidly expanding our presence in the UK and soon across Europe. We have been building the bank of the future from the ground up offering you the chance to join us and make a significant impact.
Job responsibilities
- Own and drive continuous improvement of reliability monitoring and alerting across the service portfolio.
- Define and operationalise service-level indicators service-level objectives error budgets user journeys and reliability scorecards.
- Reduce operational toil by building automation self-healing mechanisms and scalable reliability tooling.
- Lead performance engineering and capacity planning including load testing strategy bottleneck analysis and scaling plans.
- Establish resiliency patterns such as graceful degradation timeouts retries circuit breakers rate limiting and failover strategies.
- Set technical direction and standards for observability alerting automation and site reliability engineering practices.
- Provide hands-on implementation for complex or high-impact engineering work while delegating effectively to grow team capability.
- Partner across product engineering and platform teams to align reliability priorities with delivery roadmaps.
- Participate in feature planning and design reviews to ensure reliability values are built in from the start.
- Coach and mentor site reliability engineers by providing technical guidance feedback and support for development goals.
- Build practical AI capability within the team by encouraging effective practices guardrails validation and safe usage patterns.
- Drives adoption and governance of approved AI-assisted engineering practices across teams to improve code quality delivery speed and operational outcomes (e.g. AI-assisted code review/refactoring test acceleration release readiness incident/root-cause analysis) while establishing measurable validation standards (secure coding peer review automated testing) and promoting reuse of proven patterns and automation within the SDLC/TLM toolchain.
- Applies knowledge of tools within the Software Development Life Cycle toolchain including approved AI-assisted development and automation capabilities to improve the value realized by automation at scale.
Required qualifications capabilities and skills
- Formal training or certification on software engineering concepts and advanced applied experience
- Proven experience as a software engineer including proficiency in at least one programming language such as Python Go or Java.
- Demonstrated experience operating and improving production systems in a site reliability engineering or site reliability engineering capacity.
- Strong distributed-systems debugging and troubleshooting skills across services infrastructure and delivery pipelines.
- Experience with Kubernetes.
- Experience with cloud computing services.
- Familiarity with observability and reliability toolchains such as Grafana Prometheus Elasticsearch Kibana or Jaeger.
- Proven ability to lead technically by setting direction making pragmatic trade-offs guiding designs and improving engineering standards.
- Strong communication skills with the ability to influence stakeholders and translate operational risk into clear engineering priorities.
- Ability to use AI-assisted engineering tools responsibly including validating outputs understanding failure modes and adhering to secure handling practices.
Preferred qualifications capabilities and skills
- Experience with AWS.
- Prior experience leading a site reliability engineering or site reliability engineering team or acting as a technical lead for platform or production engineering.
- Experience building internal platforms or tooling including operators controllers automation frameworks or reliability guardrails.
- Experience applying AI to operational workflows such as alert enrichment incident summarisation or anomaly triage using approved tools and patterns.
- Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools within the work environment (e.g. for coding code review test acceleration troubleshooting) with the ability to set team expectations for validating AI outputs for correctness performance and security
- Strong understanding of responsible AI use in engineering workflows including data sensitivity considerations secure handling of inputs/outputs and adherence to resiliency and security expectations; experience coaching senior engineers/leads on compliant usage patterns and controls.
#ICBCareers #ICBEngineering
Employment Type : Full-Time
Experience: years
Vacancy: 1