The role
From JPMorgan Chase's own posting.
Job Description
If you take ownership of outcomes in production — not just implementation — and thrive on turning ambiguous requirements into stable, well-modeled service designs, this role was built for you.
As a Senior Lead Software Engineer at JPMorganChase within the Corporate Artificial Intelligence and Machine Learning Data Platforms – Machine Learning Center of Excellence, you will design, build, and optimize high-performance, low-latency distributed systems that serve as the backbone of our machine learning and data infrastructure. You will collaborate across engineering, data science, and platform teams to deliver resilient, cloud-native solutions that enable the firm to operate at the forefront of AI-driven innovation. Your work will directly shape the reliability, scalability, and performance of systems that process critical data across the enterprise, and your voice will carry weight in the architectural and engineering decisions that define how the platform evolves. You will have meaningful latitude to influence architecture, engineering standards, and reliability posture across services, with expectations and recognition aligned to senior-level impact.
Job responsibilities
Architect and implement low-latency, high-throughput Java Spring Boot–based distributed services using object-oriented principles, delivering production-grade performance with strong, well-defined APIs
Design and build resilient, cloud-native service architectures with high-availability requirements from 99.9% to 99.999%, leveraging AWS compute, messaging, streaming, database, and storage services including Managed Streaming for Apache Kafka (MSK), Simple Queue Service (SQS), S3, Elastic Container Service (ECS), Elastic Kubernetes Service (EKS), Lambda, Kinesis Video/Data Streams, Relational Database Service (RDS), DynamoDB, and Redshift
Develop and maintain infrastructure-as-code solutions using Terraform and/or CloudFormation to support scalable, repeatable, and auditable cloud deployments
Implement and continuously improve observability solutions — including alerting, monitoring, and reporting — using Datadog, Dynatrace, and Splunk to deliver actionable production intelligence across microservices platforms
Translate ambiguous or evolving requirements into stable, well-modeled service designs, clearly articulating engineering tradeoffs to both technical and non-technical stakeholders
Lead technical design reviews, establish engineering best practices, and drive adoption of standards that improve platform operability, reliability, and maintainability
Own production outcomes end-to-end — identifying and resolving performance bottlenecks, reliability gaps, and scalability constraints through automation and runbook-driven operations
Partner with machine learning engineers and data scientists to understand platform requirements and deliver robust, production-ready engineering solutions
Mentor and provide technical guidance to engineers across the team, fostering a culture of ownership, continuous learning, and engineering excellence
Drive adoption and governance of approved AI-assisted engineering practices to improve code quality, delivery speed, and operational outcomes — including AI-assisted code review, test acceleration, release readiness, and incident analysis — while establishing measurable validation standards and promoting reuse of proven patterns within the software development lifecycle toolchain
Required qualifications, capabilities, and skills
Formal training or certification on software engineering concepts and 5+ years applied experience, with very strong Java development skills using object-oriented principles and significant experience with Spring Boot
Demonstrated experience designing and tuning for low-latency processing in production distributed systems
Hands-on experience leveraging AWS services including MSK (Kafka), SQS, S3, ECS, EKS, Lambda, Kinesis Video/Data Streams, RDS, DynamoDB, and Redshift in large-scale, resilient service architectures
Practical experience implementing alerting, monitoring, and reporting solutions using Datadog, Dynatrace, and/or Splunk in production-grade environments
Strong engineering fundamentals including API design, testing discipline, and debugging in production contexts
Proficiency in one or more modern programming languages — with heavy emphasis on Java — writing clean, maintainable, object-oriented, and testable code
Strong experience with containerization and orchestration technologies, including Docker and Kubernetes
Demonstrated ability to communicate engineering tradeoffs clearly to both technical and non-technical stakeholders
Demonstrated experience leading effective use of enterprise-authorized AI-assisted software development tools for coding, code review, test acceleration, and troubleshooting, with the ability to set team expectations for validating AI outputs for correctness, performance, and security
Strong understanding of responsible AI use in engineering workflows, including data sensitivity considerations, secure handling of inputs and outputs, and adherence to resiliency and security expectations, with experience coaching engineers on compliant usage patterns and controls
Preferred qualifications, capabilities, and skills
Deep familiarity with low-latency, highly transactional architectures and advanced usage of AWS managed services — particularly Kinesis Video/Data Streams — for real-time processing, distributed event handling, and efficient data storage and retrieval
Expertise designing and automating observability and reporting workflows using Datadog, Dynatrace, and Splunk to deliver actionable monitoring and production intelligence across microservices platforms
Experience with modern delivery practices including continuous integration and delivery, infrastructure-as-code, and containerized deployments that support reliable service delivery at scale
Experience with Terraform and/or CloudFormation for building and maintaining cloud infrastructure in an enterprise environment