The role
From Gemini's own posting.
About the Company
Gemini is a global crypto and Web3 platform founded by Cameron and Tyler Winklevoss in 2014, offering a wide range of simple, reliable, and secure crypto products and services to individuals and institutions in over 70 countries. Our mission is to unlock the next era of financial, creative, and personal freedom by providing trusted access to the decentralized future. We envision a world where crypto reshapes the global financial system, internet, and money to create greater choice, independence, and opportunity for all — bridging traditional finance with the emerging cryptoeconomy in a way that is more open, fair, and secure. As a publicly traded company, Gemini is poised to accelerate this vision with greater scale, reach, and impact.
The Department: Platform
Our Platform organization’s purpose is to enable Gemini to scale effectively and empower our engineering teams to focus on building innovative financial products and experiences for individuals around the world. Platform builds a scalable and secure foundation that enables Engineering to develop, deploy, validate, and operate services in production; improve reliability; reduce repetitive operational work; and improve system efficiency through architectural improvements.
The Platform team builds products for engineers. We work directly with product, application engineering, data, and security teams to understand their needs, establish safe defaults, and provide standard workflows for building and operating services. We treat internal platforms as production software with users, service-level objectives, roadmaps, documentation, and feedback loops.
The Role: Staff Platform Engineer
You will provide technical leadership for the software and infrastructure that form Gemini’s internal engineering platform. This is a software engineering role grounded in deep production and infrastructure experience. You will build self-service systems, developer tools, and automation that improve the way teams create, deploy, scale, observe, and operate services.
You will shape platform strategy across cloud, networking, compute, storage, infrastructure automation, CI/CD, and the systems required to run highly available services in production. You will partner with engineering, security, data, and product leaders to turn broad operational needs into reliable platform products, establish clear technical direction, and drive adoption across teams.
You will make infrastructure programmable, secure, observable, and resilient; use SLOs and error budgets to guide tradeoffs; and continuously replace manual operational work with well-designed software. The role spans software, systems, and infrastructure engineering, with an emphasis on solving cross-functional problems and delivering measurable improvements to developer productivity and production reliability.
What success looks like in this role: Engineers can safely self-serve common workflows through documented, well-supported platform products. Platform capabilities are adopted because they reduce complexity and improve outcomes. Reliability, deployment safety, developer experience, and infrastructure efficiency improve measurably over time. Operational load grows more slowly than the number and scale of services supported by the team. Platform engineers remain close to production while spending the majority of their time on durable software and systems improvements.
Responsibilities:
Set technical direction for platform capabilities and contribute to a long-term plan for cloud, networking, infrastructure, delivery, and reliability engineering.
Lead architecture recommendations and technical engagements across the software development lifecycle, from early design through production operation.
Design, build, and maintain platform services, APIs, CLIs, controllers, workflows, and automation in languages such as Go, TypeScript, Python, or similar; own them through design, implementation, testing, deployment, and production operation.
Build self-service capabilities that let engineers provision environments, deploy services, manage infrastructure, and respond to operational conditions safely without relying on manual platform intervention.
Experience with build tooling such as Bazel or Buck2.
Create and evolve standard workflows for service development and delivery, including templates, libraries, deployment workflows, policy-as-code, and reusable platform building blocks.
Design and improve the cloud, compute, storage, identity, and infrastructure building blocks used by engineering teams.
Design and operate core network infrastructure across public clouds and regions, including routing, connectivity, traffic management, and security controls.
Operate and improve the compute, container, networking, and cloud infrastructure that supports critical services and new workloads, including distributed and AI-enabled systems where appropriate.
Improve the systems and workflows used to provision infrastructure, connect services, deliver changes, and recover from failures across environments.
Define and use service-level indicators, SLOs, error budgets, and production-readiness standards with engineering teams to make reliability and launch decisions explicit and measurable.
Improve observability, alerting, capacity planning, performance, and failure handling; design systems that degrade gracefully, roll out progressively, and recover automatically where safe.
Lead or coordinate response to significant incidents, ensure effective post-incident follow-through, and turn recurring operational work into permanent improvements through postmortems.
Reduce repetitive operational work. Measure it, automate or eliminate it, and keep the team focused primarily on engineering work that improves reliability, scale, security, cost, or developer experience.
Influence engineering standards and practices across teams through design reviews, technical writing, prototypes, internal education, and direct collaboration.
Build alignment across competing priorities and stakeholders; make pragmatic tradeoffs while preserving clear interfaces and room for future evolution.
Mentor engineers, develop technical leaders, and raise the quality of platform engineering practices across Gemini.
Qualifications:
8+ years of professional software engineering, infrastructure engineering, SRE, or an equivalent combination of experience, including substantial ownership of production systems.
Strong software engineering ability in at least one language such as Go, Python, TypeScript, Rust, or similar, with experience building maintainable services, automation, CLIs, APIs, or developer tools.
Demonstrated experience operating infrastructure and distributed systems in production, including diagnosing failures, managing capacity, improving performance, and designing for high availability and recovery.
Experience designing, operating, or improving production systems across cloud infrastructure, networking, compute, storage, deployment platforms, or related infrastructure domains.
Hands-on experience with cloud platforms such as AWS, GCP, or Azure and infrastructure as code, such as Terraform or an equivalent system.
Experience with containers and orchestration, such as Kubernetes, EKS, GKE, AKS, Docker, or Nomad, including workload deployment, networking, security, and operational troubleshooting.
Experience owning and managing cloud network security controls, including network firewalls, DNS filtering, and network segmentation.
Experience designing or improving CI/CD and automated software delivery systems, including testing, deployment, progressive rollout, rollback, and change validation.
Practical experience with observability: metrics, logs, traces, alerting, dashboards, and production health models. Familiarity with OpenTelemetry or comparable tooling is valuable.
Working knowledge of SLOs, error budgets, incident response, blameless postmortems, and techniques for reducing operational toil.
Experi