Docker

Software Engineer, Infrastructure Platform

Dockerโ€ข๐Ÿ’ฐ USD 136,912 - USD 220,687
Full Time๐Ÿ“ Worldwide - RemoteCloud & DevOps
Go (30%)Kubernetes (25%)Terraform (20%)Platform Engineering (15%)Distributed Systems (10%)
โฑ๏ธ Posted 4mo agoโœ… Verified 6 hours ago
โœจ AI Summary
๐Ÿ“Œ Before You Apply

Know the key requirements and restrictions.

Eligibility Requirements

  • โœ“
    On-call rotationThis role may require participation in an on-call rotation to provide support outside of standard business hours
  • โœ“
    Experience4+ years of backend software engineering experience building large-scale cloud or distributed systems
  • โœ“
    CommunicationClear written and verbal communication in a remote environment, including RFCs, incident writeups, and async collaboration

Company Constraints & Limitations

  • โœ“
    Visa Sponsorship AvailableDocker considers visa sponsorship on a case-by-case basis based on business needs
๐Ÿ˜Ž Benefits & Perks

See the benefits and perks offered for this role.

๐Ÿ“ˆ
Equity / Stock Options
Equity; we are a growing start-up and want all employees to have a share in the success of the company
๐Ÿ’ป
Remote Stipend / Home Office Budget
Home office setup; we want you comfortable while you work
๐ŸŒด
Paid Time Off
PTO plan that encourages you to take time to do the things you enjoy
๐ŸŒด
Paid Parental Leave
16 weeks of paid Parental leave (after 6 months of employment)
๐Ÿ“š
Learning & Development Budget
Training stipend for conferences, courses and classes
๐ŸŽ
Technology Package
Technology stipend equivalent to $100 USD net/month
โค๏ธ
Medical
Medical benefits, retirement and holidays vary by country
๐ŸŽ
Whaleness Days
Designated quarterly Whaleness Days plus end of year Whaleness break

Docker has been one of the most loved brands in developer tooling, trusted by more than 20 million monthly users and over 20 billion container image pulls. From solo founders to the world's largest companies, developers rely on Docker to build, share, and run their applications across our suite of products including Docker Desktop, Docker Hub, and Docker Scout.

We are a globally distributed, remote-first team building the tools that define how software gets built and delivered. As AI agents redefine software development, Docker is at the center of that shift, providing the sandboxed environments, verified images, and secure infrastructure that make autonomous workflows trustworthy by default.

Our Infrastructure Engineering team builds and operates the cloud-native platform that powers Dockerโ€™s suite of products. We design resilient services, automate where it helps most, and measure what matters so hundreds of engineers can ship safely to millions of users every day.

A core focus is self-service. We build paved-road platform capabilities that let internal teams provision, deploy, observe, and operate services with minimal friction and strong guardrails. We treat the platform as a product with clear contracts, well-defined defaults, and great documentation. Success is measured by adoption and fewer support requests.

How We Work

Write it down, ship it, iterate: RFCs and design docs, code review, and small safe releases.

Sustainable reliability: we prioritize root-cause fixes, good alerts, and automation over heroics.

Cross-functional by default: we partner closely with product and security teams.

AI-accelerated execution: we build agentic workflows to reduce toil and improve incident response, with guardrails, auditability, and human review.

What Youโ€™ll Work On

Reducing toil through automation, including AI-assisted and agentic operational workflows.

Building self-service onboarding and deployment workflows that reduce tickets and speed delivery.

Scaling Kubernetes foundations and evolving our traffic and ingress stack.

Responsibilities

1) Self-Service Platform Services

Build and operate internal platform services and APIs in Go, including provisioning, quotas and policies, cost insights, and platform workflows.

Deliver golden paths for self-serve onboarding and day-2 operations, including access, deployment setup, observability defaults, and governance guardrails.

Partner with teams to drive adoption through clear docs, examples, and measurable outcomes.

2) Infrastructure as Code and Reliability

Codify infrastructure with Terraform and GitOps practices, and contribute to platform tooling in Go.

Define and improve SLOs, alerting, and operational readiness. Participate in incident response and preventive follow-ups.

Help standardize safe delivery patterns, including testing gates, canaries, and rollback triggers, so deployments are routine and low-risk.

3) Kubernetes and Networking Foundations

Operate and scale multi-tenant EKS clusters and traffic and ingress systems to deliver secure, reliable routing.

Evaluate and adopt improvements with a bias toward incremental rollout and measurable impact.

4) AI and Agentic Workflows for Reliability

Build and iterate on agentic workflows that reduce operational toil, including triage support, context gathering, safe runbook execution, and remediation suggestions.

Integrate automation into delivery and operations in a way that is safe, observable, and auditable.

5) On-Call and Incident Response

Operational ownership is part of this role.

This role may require participation in an on-call rotation to provide support outside of standard business hours, including evenings, weekends, and holidays, as needed.

Youโ€™ll join an on-call rotation after onboarding and shadowing, and participate in incident response during your shifts.

We aim for sustainable on-call through good alerting, automation, and blameless postmortems focused on prevention.

Qualifications

Core Engineering Skills (must-have)

4+ years of backend software engineering experience building large-scale cloud or distributed systems

Strong software development skills in Go or a similar language, including design, testing, debugging, and code review.

Experience shipping and operating cloud services in production, often 3+ years. We hire for skill and impact, not years alone.

Solid foundation in Linux, networking fundamentals, and cloud security.

Experience building operational automation, including AI-assisted or agentic workflows, with an emphasis on safety, guardrails, and auditability.

Clear written and verbal communication in a remote environment, including RFCs, incident writeups, and async collaboration.

Nice-to-have

Kubernetes and EKS experience, plus ingress, CNI, service mesh, and familiarity with L4 and L7 load balancing.

Observability tooling such as OpenTelemetry, Prometheus, and Grafana, plus alerting and SLO practice.

CI/CD and progressive delivery, including GitHub Actions or Argo CD, canaries, and automated rollback.

Cost optimization at scale, including FinOps and capacity modeling.

Distributed systems, containers, and Go-based platform tooling.

We value depth in one area and curiosity across others, and we will help you grow in the rest.

What to Expect

First 30 Days

Ship your first change to a Terraform module or internal service and learn how we operate.

Shadow on-call and build context on our platform and reliability priorities.

First 90 Days

Own a component and deliver an improvement from design to production with measurable impact.

Join the on-call rotation and contribute effectively during your shifts.

First Year

Lead or co-lead a meaningful platform initiative, with scope that scales by level, and help reduce toil through automation.

Become a trusted contributor in one or more areas such as platform services, Kubernetes and networking foundations, or reliability automation.

Docker considers visa sponsorship on a case-by-case basis based on business needs.

Perks

Freedom & flexibility; fit your work around your life

Designated quarterly Whaleness Days plus end of year Whaleness break

Home office setup; we want you comfortable while you work

16 weeks of paid Parental leave (after 6 months of employment)

Technology stipend equivalent to $100 USD net/month

PTO plan that encourages you to take time to do the things you enjoy

Training stipend for conferences, courses and classes

Equity; we are a growing start-up and want all employees to have a share in the success of the company

Docker Swag

Medical benefits, retirement and holidays vary by country

Remote-first culture, with offices in Seattle and Paris

Docker embraces diversity and equal opportunity. We are committed to building a team that represents a variety of backgrounds, perspectives, and skills. The more inclusive we are, the better our company will be.

#LI-REMOTE

๐Ÿ’ผ Similar Jobs You Might Like

ClickHouse
Full Time
Cloud Software Engineer - Observability PlatformClickHouse
USD 141,000 - USD 230,000๐Ÿ“ United States๐Ÿ“ Worldwide
Cloud & DevOps
โฑ๏ธ Posted 1w ago
โœ… Verified 6 hours ago
ClickHouse
Full Time
Senior Site Reliability Engineer- Remote ClickHouse
USD 141,000 - USD 230,000๐Ÿ“ United States๐Ÿ“ Worldwide
Cloud & DevOps
โฑ๏ธ Posted 4mo ago
โœ… Verified 6 hours ago
ClickHouse
Full Time
Senior Software Engineer - Cloud InfrastructureClickHouse
USD 141,000 - USD 230,000๐Ÿ“ United States๐Ÿ“ Worldwide
Cloud & DevOps
โฑ๏ธ Posted 1w ago
โœ… Verified 6 hours ago
ClickHouse
Full Time
Senior Cloud Engineer ClickHouse
USD 141,000 - USD 230,000๐Ÿ“ United States
Cloud & DevOps
โฑ๏ธ Posted 2mo ago
โœ… Verified 6 hours ago
Docker
Full Time
Staff Software Engineer, InfrastructureDocker
CAD 238,250 - CAD 382,250๐Ÿ“ United States๐Ÿ“ Canada
Cloud & DevOps
โฑ๏ธ Posted 2mo ago
โœ… Verified 6 hours ago
Render
Full Time
Software Engineer, Dev Velocity (all levels)Render
USD 170,000 - USD 290,000๐Ÿ“ United States๐Ÿ“ Canada
Cloud & DevOps
โฑ๏ธ Posted 1mo ago
โœ… Verified 6 hours ago