Skydio company logo

Skydio

Website

Site Reliability Engineer

Oct 8, 2026
Full Time
DevOps/SysAdmin
Remote & Hybrid Options
180k-220k / Yr
All U.S.

Skydio is the leading US drone company and the world leader in autonomous flight, the key technology for the future of drones and aerial mobility. The Skydio team combines deep expertise in artificial intelligence, best-in-class hardware and software product development, operational excellence, and customer obsession to empower a broader, more diverse audience of drone users, from utility inspectors to first responders, soldiers in battlefield scenarios, and beyond.

About the role:

We are looking for a hands-on Site Reliability Engineer to build, operate, and scale the cloud infrastructure that powers our products. This role is focused on owning production infrastructure, including Kubernetes, AWS, infrastructure as code, CI/CD, observability, networking, and reliability.

You don't need to be an expert in every area, but you should have strong Kubernetes and cloud fundamentals with meaningful depth in at least one infrastructure domain. Our technology helps save lives. You’ll play a critical role in keeping the infrastructure behind it reliable, scalable, and available when it matters most.

How you'll make an impact:

  • Build, operate, and troubleshoot production Kubernetes/EKS clusters.

  • Perform Kubernetes upgrades, node rollouts, and cluster maintenance.

  • Build and manage AWS infrastructure including VPCs, networking, subnets, load balancers, IAM, EKS, databases, and storage.

  • Define and maintain infrastructure using Terraform.

  • Build and operate CI/CD and deployment infrastructure.

  • Troubleshoot production issues across Kubernetes, AWS, Linux, networking, and databases.

  • Build monitoring, alerting, and observability for critical infrastructure.

  • Participate in on-call rotations and respond to production incidents.

  • Identify and solve infrastructure scaling and reliability problems.

  • Automate operational work using Python, Go, or similar languages.

  • Help expand infrastructure across new regions and deployment environments.

What makes you a good fit:

  • 3+ years of experience as a Production Engineer, SRE, DevOps, or equivalent infrastructure role.

  • Strong hands-on experience operating Kubernetes, not simply deploying applications to existing clusters.

  • Experience managing Kubernetes/EKS upgrades and production clusters.

  • Strong AWS fundamentals, including VPCs, public/private subnets, networking, load balancers, EKS, IAM, and databases.

  • Production experience with Terraform or similar infrastructure-as-code tooling.

  • Experience owning or maintaining CI/CD and deployment systems such as Argo CD, Spinnaker, GitHub Actions, GitLab CI/CD, or Jenkins.

  • Experience diagnosing production infrastructure and networking problems.

  • Experience solving meaningful scaling or reliability challenges.

  • Obtaining FAA Part 107 certification within the first 60 days of employment is strongly encouraged for all Skydio employees and required for certain positions.

  • This position requires access to export-controlled technical data, restricted government information, and/or information systems subject to U.S. government security and access-control requirements. Employment in this role is contingent upon verification of U.S. person status and the ability to access controlled or restricted information as required for the position.

Bonus points:

  • Helm and GitOps experience.

  • Datadog or similar observability tooling.

  • PostgreSQL/database operations experience.

  • Multi-region infrastructure experience.

  • On-premises or disconnected deployment experience.

  • Streaming or high-throughput distributed systems experience.

Compensation:

At Skydio, our compensation packages for regular, full-time employees include competitive base salaries, equity in the form of stock options, and comprehensive benefits packages. Compensation will vary based on factors, including skill level, proficiencies, transferable knowledge, and experience. Relocation assistance may also be provided for eligible roles. The annual base salary range for this position is $180,000 - $220,000*. Fundamentally, we believe that equity is the key to long-term financial growth, and we ensure all regular, full-time employees have the opportunity to significantly benefit from the company's success. Regular, full-time employees are eligible to enroll in the Company’s group health insurance plans. Regular, full-time employees are eligible to receive the following benefits: Paid vacation time, sick leave, holiday pay and 401K savings plan. This position and all associated benefits are subject to applicable federal, state, and local laws, as well as the Company’s policies and eligibility criteria.

*Compensation for certain positions may vary based on the position's location.

#Li-wa1

At Skydio we believe that diversity drives innovation. We have created a multidisciplinary environment that embraces the power of diverse perspectives to create elegant solutions for complex problems. We are committed to growing our network of people, programs, and resources to nurture an inclusive culture.

Qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, disability, protected veteran status, or other characteristics protected by federal, state or local anti-discrimination laws.

For positions located in the United States of America, Skydio, Inc. uses E-Verify to confirm employment eligibility. To learn more about E-Verify, including your rights and responsibilities, please visit https://www.e-verify.gov/

MoreDevOps/SysAdminJobs

RunPod company logo
Site Reliability Engineer
Oct 8
RunPod
Oct 8
All U.S.
Fully Remote
150k-200k / Yr
<p>Runpod is the AI Developer Cloud. More than one million developers, from indie researchers to teams running frontier models in production, use Runpod to experiment, train, fine-tune, deploy, and scale AI on one platform. The platform has processed more than 20 billion inference requests. We closed a $100M Series A in June 2026. We're at an inflection point for AI infrastructure, and we're building the platform the next generation of developers will depend on.</p><p>We're a small, remote-first team. We take ownership seriously, move fast, and ship work that more than a million developers rely on every day. We're looking for people who care deeply, build with urgency, and want to matter at scale.</p><p>Learn more in our CEO's funding announcement: <a href='https://www.runpod.io/blog/one-million-developers'>https://www.runpod.io/blog/one-million-developers</a>.</p><p>The Reliability team owns the availability, performance, and operational excellence of Runpod’s global platform. While infrastructure teams build the systems, the Reliability team ensures those systems remain resilient, observable, and scalable under real-world production conditions.</p><p>This team is responsible for:</p><ul><li><p>Defining and enforcing reliability standards across engineering</p></li><li><p>Designing incident response processes and improving recovery times</p></li><li><p>Building observability systems and reliability tooling</p></li><li><p>Driving SLO adoption and production readiness reviews</p></li><li><p>Reducing operational toil through automation</p></li></ul><p>The Reliability team works cross-functionally with Infrastructure, Product Engineering, and Support to ensure our systems remain stable and performant as we scale rapidly. We value proactive problem solving, automation-first thinking, and strong ownership of production systems.</p><p>As a Site Reliability Engineer on the Reliability team, you will focus on ensuring the stability and resilience of Runpod’s distributed platform. You will partner with engineering teams to improve system design, strengthen observability, and prevent incidents before they happen.</p><p>This role blends software engineering with production operations. You’ll work on reliability frameworks, SLO design, automation, and production hardening, reducing errors and improving performance across different services and infrastructure.</p><p>This is a high-impact role central to maintaining trust with developers running critical AI workloads on Runpod.</p><h2>Your Impact</h2><ul><li><p>Increase platform uptime and reduce incident frequency and duration</p></li><li><p>Establish and operationalize SLIs/SLOs across services</p></li><li><p>Improve MTTR through better tooling, automation, and runbooks</p></li><li><p>Strengthen production readiness standards</p></li><li><p>Drive long-term systemic reliability improvements</p></li></ul><p>You will influence how reliability is defined and measured across Runpod and help build the operational backbone of the company.</p><h2>Responsibilities:</h2><h2>Reliability Engineering</h2><ul><li><p>Define and implement SLIs/SLOs for critical services</p></li><li><p>Lead incident response and coordinate cross-team mitigation efforts</p></li><li><p>Conduct blameless postmortems and ensure corrective actions are completed</p></li><li><p>Perform production readiness reviews for new services and features</p></li><li><p>Identify systemic risks and drive preventative improvements</p></li></ul><h2>Observability & Monitoring</h2><ul><li><p>Design and improve monitoring, alerting, and dashboards (Prometheus, Grafana, etc.)</p></li><li><p>Improve signal-to-noise ratio in alerts and reduce alert fatigue</p></li><li><p>Build internal tooling for reliability tracking and reporting</p></li><li><p>Improve visibility into GPU performance and distributed systems health</p></li></ul><h2>Automation & Toil Reduction</h2><ul><li><p>Automate recurring operational workflows</p></li><li><p>Build tools and scripts (Python, Go, Bash) to eliminate manual processes</p></li><li><p>Improve deployment safety through automation and guardrails</p></li><li><p>Strengthen CI/CD reliability and release processes</p></li></ul><h2>Cross-Functional Reliability Advocacy</h2><ul><li><p>Partner with engineering teams to improve system resilience</p></li><li><p>Provide guidance on fault tolerance, scalability, and failure handling</p></li><li><p>Contribute to architectural discussions with a reliability-first mindset</p></li></ul><h2>Requirements:</h2><ul><li><p>5+ years of experience in SRE, Reliability Engineering, or Production Engineering</p></li><li><p>Strong Linux systems and Networking expertise</p></li><li><p>Experience managing containerized production systems</p></li><li><p>Strong understanding of distributed systems and failure modes</p></li><li><p>Experience defining and managing SLIs/SLOs</p></li><li><p>Proven incident response and postmortem leadership experience</p></li><li><p>Strong scripting or programming skills</p></li><li><p>Experience with monitoring and alerting systems</p></li><li><p>Excellent written communication skills</p></li><li><p>Successful completion of a background check</p></li></ul><h2>Preferred:</h2><ul><li><p>Experience with GPU infrastructure or AI/ML platforms</p></li><li><p>Experience improving reliability in high-growth or large scale environments</p></li><li><p>Familiarity with GPU observability tooling</p></li><li><p>Experience with Infrastructure as Code</p></li><li><p>Experience working in startup environments</p></li><li><p>Experience building internal reliability platforms or frameworks</p></li></ul><h2>What You’ll Receive:</h2><ul><li><p>The competitive base pay for this position ranges from $150,000- $200,000 usd. This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location</p></li><li><p>Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside.</p></li><li><p>Generous medical, dental & vision plans</p></li><li><p>Flexible PTO- take the time you need to recharge</p></li><li><p>Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication </p></li><li><p>Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale.</p></li></ul><p>Runpod is committed to maintaining a workplace free from discrimination and upholding the principles of equality and respect for all individuals. We believe that diversity in all its forms enhances our team. As an equal opportunity employer, Runpod is committed to creating an inclusive workforce at every level. We evaluate qualified applicants without regard to race, color, religion, sex, sexual orientation, gender identity, national origin, age, marital status, protected veteran status, disability status, or any other characteristic protected by law. We welcome every qualified candidate eligible to work in the United States; however, we are currently unable to sponsor employment visas.</p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Bitwarden company logo
Senior Site Reliability Engineer - FedRAMP
Oct 4
Bitwarden
Oct 4
All U.S.
Fully Remote
140k-185k / Yr
<p>Bitwarden empowers enterprises, developers, and individuals to securely store and share sensitive data. With a transparent, open-source approach to password management, secrets management, and passwordless and passkey innovations, Bitwarden makes it easy for users to extend robust security practices across all online activities. Founded in 2016 with headquarters in Santa Barbara, California, Bitwarden is supported by a passionate global community of security experts and enthusiasts.</p><p>As a Senior Site Reliability Engineer - FedRAMP at Bitwarden, you will be responsible for operating and building Bitwarden Gov. You will be responsible for every element of the infrastructure and operations within the new region, including availability, performance, change management, monitoring, incident response, and capacity planning, all while maintaining a strong emphasis on security/compliance/governance.</p><p>We’re looking for team members who can own the current and future state of our Gov cloud infrastructure distributed across multiple cloud providers, design and build out new infrastructure, as well as operationalize tools and technologies that enable our team and application to scale.</p><p><strong>In this role, you will be managing services in a FedRAMP compliant environment, which </strong><strong>requires you to be a U.S. citizen. Additionally, this position may include work specified by the U.S. government that is restricted to U.S. citizens and must be conducted on U.S. territory.</strong></p><h2>Responsibilities</h2> <ul><li>Take ownership of the Bitwarden Gov cloud infrastructure, with emphasis on security, reliability and quality</li><li>Evaluate current infrastructure and, on a regular basis, make recommendations for reliability, security, availability, scalability and cost management</li><li>Implement site reliability tools, monitoring, early warning and alert systems, and observability across Bitwarden Gov cloud environments</li><li>Respond to infrastructure based outages; participate and contribute to ongoing strategy for 24x7 support (This role has an on-call rotation with weekend shifts)</li><li>Architectural designs and engineering operations at scale</li><li>Active participation in code reviews, learning and spreading technical knowledge</li><li>Contribute and mature incident management/escalation processes</li><li>Collaborate with cross functional teams to refine priorities and deliverables</li><li>Evaluate and identify opportunities for new initiatives to support organizational needs</li><li>Create/maintain standard operational procedures</li><li>Ongoing vulnerability management and security remediations</li></ul> <h2>What You Bring To Bitwarden</h2> <ul><li>Minimum 5 years of experience with FedRAMP (Moderate/High levels) - continuous monitoring, vulnerability management</li><li>Sense of curiosity, resourcefulness, pragmatism</li><li>Production experience with incident response, on-call rotations</li><li>Strong working knowledge and understanding of cloud networking, traffic management (CNI, Istio)</li><li>Expertise with multi-region deployments in public cloud environments (AWS/Azure/GCP)</li><li>Demonstrable production Kubernetes experience (Managed Kubernetes, Helm, kubectl, kOps, etc.)</li><li>AI fluency with AI assisted workstreams and AI tooling</li><li>Competency with least one programming language, such as C#, Python, Go, etc</li><li>Experience with cloud deployment and automation tools/methodologies (i.e. GitOps, Terraform, Pulumi)</li><li>Ability to maintain discretion, handle sensitive information, and improve security best-practices</li><li>Technocrat at heart, staying current with trends and new technologies</li><li>Collaborative and adaptable mindset</li><li>Openness and authenticity combined with excellent communication skills</li><li>Excitement and enthusiasm for open source and for better internet security</li><li>Excellent problem-solving skills – you might not know all the answers, but you know how to find and communicate the possible solutions</li></ul> <h2>Nice-to-haves</h2> <ul><li>Experience with distributed monitoring</li><li>Experience with data engineering</li><li>Startup experience</li><li>Open source experience</li><li>User of Bitwarden</li><li>Prior SaaS experience</li></ul> <h2>What To Expect In The Interview Process</h2> <ul><li>Meeting with our Recruiting team</li><li>Interview with SRE Manager</li><li>Interview(s) with SRE team panel</li><li>Interview with Head of Security</li></ul> <h2>A Few Reasons To Work With Us</h2> <ul><li>Our user community loves us and we love them. Come to work each day with a sense of purpose as we bring a more secure internet experience to everyone from our friends and family to the world’s largest organizations.</li><li>Become an expert. You’ll get immersed in the prominent technology markets of security and open source software. We are dedicated to building an incredible team.</li><li>Work remotely with motivated andinnovative team members and take part in productive and fun meetups.</li><li>Learn and grow. Take on new challenges with the support of your team.</li></ul> <p>In the United States, the starting base compensation range for this role is $140,000 - $185,000. Actual compensation may vary based on level, relevant experience, and skill set as assessed in the interview process, as well as market data by location. See our careers page for a list of benefits. Please note that compensation outside the U.S. will differ based on the market.</p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Lambda company logo
Senior Site Reliability Engineer - SDN
Oct 1
Lambda
Oct 1
Multiple States
Hybrid
240k-312k / Yr
<p>Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.</p><p>If you'd like to build the world's best AI cloud, join us.</p><p>*Note: This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.</p><p>Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.</p><h2>What You'll Do</h2><ul><li><p>Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure</p></li><li><p>Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs</p></li><li><p>Develop tooling and automation to reduce operational toil and improve reliability</p></li><li><p>Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows</p></li><li><p>Deploy and maintain network monitoring, observability, and management tools</p></li><li><p>Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts</p></li><li><p>Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation</p></li></ul><h2>You</h2><ul><li><p>Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role</p></li><li><p>Have experience operating and supporting large-scale distributed systems in production</p></li><li><p>Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations</p></li><li><p>Have experience participating in on-call rotations and incident response</p></li><li><p>Have strong troubleshooting skills across Linux systems, Kubernetes, distributed systems, and networking</p></li><li><p>Have experience with observability platforms, monitoring, alerting, and metrics</p></li><li><p>Are comfortable working on the Linux command line and have a solid understanding of the Linux networking stack</p></li><li><p>Have experience with multi-datacenter and hybrid cloud environments</p></li><li><p>Have experience automating infrastructure and operational workflows using Python, Ansible, or similar tools</p></li><li><p>Have experience designing and operating CI/CD and GitOps deployment workflows</p></li></ul><h2>Nice To Have</h2><ul><li><p>Experience building and operating Software Defined Networks (SDN), including OpenStack Neutron, OVN, and OVS</p></li><li><p>Experience operating production-scale SDNs in a cloud environment (e.g., infrastructure powering AWS VPC-like networking services)</p></li><li><p>Software development experience in Go and/or Python (C is a plus)</p></li><li><p>Experience automating infrastructure and network configuration using Kubernetes, Helm, Terraform, and Ansible</p></li><li><p>Deep understanding of the Linux networking stack and its interaction with network virtualization technologies, SR-IOV, and DPDK</p></li><li><p>Understanding of the SDN ecosystem and modern cloud networking architectures</p></li><li><p>Experience diagnosing complex production issues across infrastructure, networking, and application layers</p></li></ul><h2>Salary Range Information</h2><p>The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.</p><h2>About Lambda</h2><ul><li><p>Founded in 2012, with 500+ employees, and growing fast</p></li><li><p>Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove</p></li><li><p>We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG</p></li><li><p>Our values are publicly available: <a href='https://lambda.ai/careers'>https://lambda.ai/careers</a></p></li><li><p>We offer generous cash & equity compensation</p></li><li><p>Health, dental, and vision coverage for you and your dependents</p></li><li><p>Wellness and commuter stipends for select roles</p></li><li><p>401k Plan with 2% company match (USA employees)</p></li><li><p>Flexible paid time off plan that we all actually use</p></li></ul><h2>Equal Opportunity Employer</h2><p>Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.</p>
CA, WA
DevOps/SysAdmin
Planet company logo
Senior Site Reliability Engineer
Sep 9
Planet
Sep 9
All U.S.
Remote & Hybrid Options
142.8k-178.5k / Yr
<p><strong>Welcome to Planet. We believe in using space to help life on Earth.</strong></p><p>Planet designs, builds, and operates the largest constellation of imaging satellites in history. This constellation delivers an unprecedented dataset of empirical information via a revolutionary cloud-based platform to authoritative figures in commercial, environmental, and humanitarian sectors. We are both a space company and data company all rolled into one.</p><p>Customers and users across the globe use Planet's data to develop new technologies, drive revenue, power research, and solve our world’s toughest obstacles.</p><p>As we control every component of hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety of domains.</p><p>We have a people-centric approach toward culture and community and we strive to iterate in a way that puts our team members first and prepares our company for growth. Join Planet and be a part of our mission to change the way people see the world.</p><p>Planet is a global company with employees working remotely world wide and joining us from offices in San Francisco, Washington DC, Germany, Austria, Slovenia, and The Netherlands.</p><p><strong>About the Role:</strong></p><p>Planet designs, builds, and operates the largest constellation of imaging satellites in history. This constellation delivers an unprecedented dataset of empirical information via a revolutionary cloud-based platform to authoritative figures in commercial, environmental, and humanitarian sectors. We are both a space company and data company rolled into one.</p><p>In this role, you will join Planet's Direct Access Service Infrastructure team, directly contributing to our next-generation Constellation as a Service platform. This platform represents a major new offering to our customers that goes beyond traditional cloud-based platforms and supports on-premises deployments. </p><p>You will be responsible for building, deploying, and operating critical compute software that supports end-to-end imaging operations within customer on-premises and/or cloud environments. You will use your understanding of internal compute requirements as well as customers' environmental-specific constraints to help design, implement, and support a robust system for reproducible deployments across operating environments, to guarantee the reliability, scalability, and availability of our services. To do this, you will partner closely with cross-functional engineering teams to enable and empower the integration of software solutions and the troubleshooting of distributed systems.</p><p>This is a full-time, remote position based in the United States or Canada. If located near an office, you are expected to work from that office 3 days per week.</p><p><strong>Impact You'll Own:</strong></p><ul><li>Build and deploy computing services and infrastructure in customer environments for a next-generation satellite operations and image processing end-to-end platform</li><li>Operate in a high-impact, tight knit team to architect novel systems for air-gapped deployments at scale</li><li>Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and external customers</li><li>Responsible for operations such as deployments, service orchestration, and documentation for cross platform stakeholders</li><li>Scale architecture while ensuring availability of services</li><li>Improve reliability and scalability by resolving edge cases, studying failure modes, and writing tests</li><li>Participate in on-call rotations to ensure operational excellence </li></ul> <p><strong>What You Bring:</strong></p><ul><li>6+ years of experience building services that leverage cloud-native infrastructure and tooling</li><li>Bachelor’s degree in Computer Science or similar</li><li>Experience deploying and maintaining bare-metal and cloud kubernetes through tools such as Talos, RKE2, Proxmox, or k3s</li><li>Proficiency with Terraform, Ansible, Helm, Kustomize, and/or similar IaC / GitOps tooling</li><li>Experience with CI/CD tooling, such as Jenkins, GitLab CI/CD, Argo CD, or CircleCI</li><li>Experience successfully building, releasing, and supporting highly available, consistently performant services</li><li>Knowledge of hardware and network level implications of on-prem compute</li><li>Experience with platform optimization, particularly resource optimization, management, and cluster tuning in a constrained environment</li><li>Ability to observe and troubleshoot distributed systems with tools such as Alloy, Prometheus, Grafana, and OpenTelemetry</li><li>Advanced skills in Python, Bash, and other tooling as appropriate to build services and meet product goals</li><li>Excellent communication skills and the ability to work through collaboration with cross-functional engineering teams</li><li>Experience working with Jira for task management and progress tracking</li></ul> <p><strong>What Makes You Stand Out:</strong></p><ul><li>Experience with CUDA-based GPU programs</li><li>Security expertise in sensitive environments, including implementing zero-trust architectures, hardening Kubernetes clusters, conducting security audits, and deploying workloads in air-gapped environments</li></ul> <p><strong>Application Deadline:</strong></p><p>November 11, 2026 by 11:59p / 23:59 CET (Central European Time)</p><p><strong>EAR/ITAR Requirements:</strong></p><p><em>This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.</em></p><p><strong>Benefits While Working at Planet:</strong></p><p><em>These offerings are dependent on employment type and geographical location, based upon applicable law or company policy.</em></p><ul><li>Comprehensive Medical, Dental, and Vision plans</li><li>Health Savings Account (HSA) with a company contribution</li><li>Generous Paid Time Off in addition to holidays and company-wide days off </li><li>16 Weeks of Paid Parental Leave</li><li>Wellness Program and Employee Assistance Program (EAP)</li><li>Home Office Reimbursement</li><li>Monthly Phone and Internet Reimbursement</li><li>Tuition Reimbursement and access to LinkedIn Learning</li><li>Equity</li><li>Commuter Benefits (if local to an office)</li><li>Volunteering Paid Time Off</li></ul> <p><strong>Compensation:</strong></p><p>The US base salary range for this full-time position at the commencement of employment is listed below. Additionally, this role might be eligible for discretionary short-term and long-term incentives (bonus and equity). The final salary range is determined by job related experience, skills and location. The range displays our typical hiring range for new hire salaries in US locations only. Your recruiter can share more about the specific salary range for your preferred location during the hiring process.</p><h2>#Li-remote</h2><p>New York City + California Salary Range</p><p>$153,000&mdash;$191,300 USD</p><p>San Francisco Salary Range</p><p>$162,600&mdash;$203,200 USD</p><p>US National Salary Range</p><p>$142,800&mdash;$178,500 USD</p><p><strong>San Francisco Fair Chance Ordinance<br></strong>Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.<strong><br></strong></p><p><strong>Why we care so much about Belonging. <br></strong>We’re dedicated to helping the whole Planet, and to do that we must strive to represent all of it within each of our offices and on all of our teams. That’s why Planet is guided by an ultimate north star of Belonging—dreaming big as we approach our ongoing work. If this job intrigues you, but you’re thinking you might not have all the qualifications, please... do apply! At Planet, we are looking for well-rounded people from around the world who can contribute to more ways than just what is listed in this job description. We don’t just fill positions, we aspire to fulfill people’s careers, most excited about folks who are motivated by our underlying humanitarian efforts. We are a few orbits around the sun before we get to where we want to be, so we hope you’re excited to come along for the ride. </p><p><strong>EEO statement: <br></strong>Planet is committed to building a community where everyone belongs and we invite people from all backgrounds to apply. Planet is an equal opportunity employer, and committed to providing employment opportunities regardless of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, pregnancy, childbirth and breastfeeding, age, sexual orientation, military or veteran status, or any other protected classification, in accordance with applicable federal, state, and local laws. <a href='https://www.eeoc.gov/sites/default/files/2023-06/22-088_EEOC_KnowYourRights6.12.pdf'>Know Your Rights.</a></p><p><strong>Accommodations: <br></strong>Planet is an inclusive community and we know that everyone has their own needs. If you have a disability or special need that requires accommodation during the hiring process, please reach out to [email protected] or contact your recruiter with your request. Your message will be confidential and we will be happy to assist you.</p><p><strong>Privacy Policy</strong>: By clicking 'Apply Now' at the top of this job posting, I acknowledge that I have read the <a href='https://www.planet.com/privacy/#california'>Planet Data Privacy Notice for California Staff Members and Applicants</a>, and hereby consent to the collection, processing, use, and storage of my personal information as described therein.</p><p><strong>Privacy Policy (European Applicants):</strong> By clicking 'Apply Now' at the top of this job posting, I acknowledge that I have read the <a href='https://www.planet.com/privacy/'>Candidate Privacy Notice GDPR Planet Labs Europe</a>, and hereby consent to the collection, processing, use, and storage of my personal information as described therein.</p><p><strong>AI in Our Interviewing Process</strong>: Planet is committed to providing an exceptional interview experience for all candidates. We currently use <a href='https://www.metaview.ai/'>Metaview</a> to better focus on candidates and less on trying to capture notes. As such, with the candidate's consent, select interviews may be recorded and include a 'Planet AI Notetaker' for transcription and summarization purposes. Should an interview involve use of AI interview technologies, the candidate will receive notification and have the ability to opt out both in advance and/or real-time. Opting out will not affect one's candidacy.</p><p><strong>Candidate AI Policy</strong>: Planet embraces Artificial Intelligence (AI) tools, and we encourage its responsible use. We understand that candidates may use various resources, including AI tools, to <em>prepare</em> for interviews and assessments. However, <em>during any live interview stage or when actively completing assessments for this position, the use of AI tools—e.g. Large Language Models (LLMs), deep fake technology, etc.—is strictly prohibited unless explicitly prompted by an interviewer or assessment instructions</em>. If you are unsure about acceptable use, please contact your recruiter for clarification. If an AI tool or similar technology is desired as an accommodation, please contact [email protected] with your request for assistance. Your message will be confidential, and we will be happy to assist you. Violation of this policy may result in disqualification of your application.</p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
AuthZed company logo
Sr. Site Reliability Engineer
Sep 9
AuthZed
Sep 9
All U.S.
Fully Remote
150k-195k / Yr
<h2>About AuthZed:</h2><p>We are the creators and maintainers of SpiceDB and the authorization infrastructure that companies around the world depend on to keep their engineering teams focused on what matters most - their own product.</p><p>We are a Series A company, fixing broken access control with products that eliminate complex permission management while delivering enterprise-scale performance and consistent access control.</p><p>AuthZed is a fully remote company with employees across the US, Canada, and Europe. We’re a hardworking and close-knit group with a software-driven culture (yep, even our GTM team understands and loves this technology)! We bring integrity to all our interactions, fostering confidence in decision making - trusting and respecting each voice on our team, every day.</p><h2>Company Values:</h2><ul><li><p><strong>Agency: </strong>Everyone should have the capability, freedom, and confidence to bring about changes to our business and product. Organizational processes exist to clearly define our goals, but not restrict how progress is made.</p></li><li><p><strong>Collaboration:</strong> Success is defined in various dimensions and no single person can be an expert in all of them. Without valuing the opinions of others, finding compromises, and sharing mutual trust and respect, you cannot arrive at the best possible solution.</p></li><li><p><strong>Open-mindedness: </strong>Without asking questions, testing assumptions, and questioning our pre-existing biases we risk operating within an echo-chamber. We celebrate the representation of diverse perspectives and backgrounds as a catalyst for creating an inclusive work environment that everyone can appreciate.</p></li></ul><h2>About the Role:</h2><p>As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our systems. You will be responsible for designing, implementing, and maintaining scalable infrastructure solutions to support our growing customer base. This is an exciting opportunity to work in a fast-paced environment and contribute to the success of a company bringing a Google-inspired authorization system to companies around the globe.</p><h2>What you’ll own:</h2><ul><li><p>Design, implement, and maintain highly available and scalable infrastructure solutions for our projects, products, and customers.</p></li><li><p>Monitor and analyze system performance, identifying and resolving bottlenecks and issues to ensure optimal performance and reliability.</p></li><li><p>Automate infrastructure deployment and configuration management processes.</p></li><li><p>Continuously improve system reliability, security, and efficiency through proactive monitoring, capacity planning, and performance tuning.</p></li><li><p>Troubleshoot and resolve complex infrastructure and application issues in production and test environments.</p></li><li><p>Collaborate with software engineering teams to design and implement systems that are resilient, scalable, and secure.</p></li><li><p>Participate in on-call rotation and respond to production incidents in a timely manner.</p></li><li><p>Document system configurations, troubleshooting procedures, and operational guidelines.</p></li></ul><h2>What you bring:</h2><ul><li><p>Proven experience as a Site Reliability Engineer or in a similar role.</p></li><li><p>Strong understanding of networking, operating systems, and cloud infrastructure.</p></li><li><p>Experience with Site Reliability Engineering, System Design, and Distributed Computing.</p></li><li><p>Experience in various programming languages — we currently have SDKs for NodeJS, Java, Python, Ruby, and Go.</p></li><li><p>Experience with containerization technologies such as Docker and Kubernetes.</p></li><li><p>Knowledge of infrastructure-as-code tools like Terraform and Pulumi.</p></li><li><p>Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack).</p></li><li><p>Experience with lower-level implementation details of relational databases (bonus if you have have experience with distributed SQL databased like Google Cloud Spanner or CockroachDB).</p></li><li><p>Experience working with Git and GitHub.</p></li><li><p>Experience with continuous integration and deployment systems.</p></li><li><p>Strong problem-solving and troubleshooting skills.</p></li><li><p>Excellent communication and collaboration abilities.</p></li></ul><h2>Extra shine:</h2><ul><li><p>Experience with Authorization systems.</p></li></ul><h2>Life at AuthZed:</h2><ul><li><p>Opportunity to work with cutting-edge technology in a rapidly growing sector.</p></li><li><p>A supported environment where your ideas lead to real impact.</p></li><li><p>Competitive salary based on experience.</p></li><li><p>Stock options at an early-stage startup.</p></li><li><p>Comprehensive benefits including healthcare (US-based) and other insurance.</p></li><li><p>A full remote and flexible schedule to accommodate different timezones</p></li><li><p>Twice-yearly travel for team offsites focused on team bonding, collaboration, and having fun!</p></li></ul>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Sanity company logo
Senior Site Reliability Engineer
Sep 5
Sanity
Sep 5
Multiple States
Fully Remote
<p>At <a href='http://Sanity.io'><strong><u>Sanity.io</u></strong></a>, we’re building the future of AI-powered Content Operations. Our AI Content Operating System gives teams the freedom to model, create, and automate content the way their business works, accelerating digital development and supercharging content operations efficiency. <br><br>It's always peak hour somewhere: with infrastructure and customers spanning every continent, a Sanity SRE makes sure the platform we build is scalable and fast, safe to deploy, and inspiring to use. The scale is real: Content Lake alone handles around 75,000 requests a second, about 4.5m a minute, and companies like <strong>Skims</strong>, <strong>Figma</strong>, <strong>Riot Games</strong>, <strong>Anthropic</strong>, <strong>Complex</strong>, <strong>Nordstrom</strong>, <strong>Arc’teryx,</strong> and <strong>Morningbrew</strong> run their content operations on it.</p><p>Our stack is built on a mix of the tried and tested and the bleeding edge. The core technologies we currently use include Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, and Google Cloud Platform.</p><p>The SRE role involves close partnership with our development teams to design and build infrastructure that supports our goal: to be the best platform for authoring, processing, and distributing content worldwide in real time. You will work close to the metal on the security, stability, and performance our customers have come to expect, and help raise the reliability bar as we grow.</p><h2>What you would do:</h2><ul><li><p>Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability.</p></li><li><p>Diagnose and troubleshoot complex distributed systems running at high request volume.</p></li><li><p>Ensure observability and analyze the behavior of our stack.</p></li><li><p>Contribute to in-flight work like modernizing our edge, caching, and gateway layers onto Fastly and tightening observability across the platform.</p></li><li><p>Raise the reliability bar through better dashboards, alert severity, paging standards, on-call readiness, and incident response.</p></li><li><p>Make deployment boring in the best way: build golden paths, production readiness checks, safe rollouts, and useful automation so engineers have fewer places to look before they ship.</p></li><li><p>Mentor engineers and raise the technical bar through code review, design review, and pairing.</p></li><li><p>Participate in our on-call rotation and help our developer on-call rollout land well.</p></li></ul><h2>About you:</h2><ul><li><p>Based in the United States, with reasonable overlap with European engineering hours.</p></li><li><p>Experience with SRE/DevOps tools, processes, and culture.</p></li><li><p>5+ years of experience as part of an SRE on-call rotation.</p></li><li><p>Analytical approach to designing, diagnosing, and optimizing infrastructure.</p></li><li><p>Experience with managing scalable, highly available, cloud-based applications, ideally with high request volume and customer-facing uptime expectations.</p></li><li><p>Experience with Kubernetes for orchestrating, scaling, and managing containerized applications in cloud-based environments.</p></li><li><p>Experience building CI/CD pipelines.</p></li><li><p>Experience with an observability stack (Prometheus, et al.).</p></li><li><p>Comfortable working across CDNs, edge, gateways, and caching layers, or eager to go deep there.</p></li><li><p>You improve on-call and reliability by building systems, standards, and feedback loops that make production healthier over time.</p></li><li><p>You are comfortable dealing with incidents and outages and have built a practical, thoughtful communication style for handling high-pressure situations.</p></li><li><p>An open but considered approach to new technologies.</p></li></ul><p>There are many roads leading up to being an SRE. Our team is already a mix of self-taught and formally educated people. Don't self-select out!</p><h2>What we can offer:</h2><ul><li><p>A highly-skilled, inspiring, and supportive team</p></li><li><p>Real infrastructure scale and meaningful, hands-on work changing how it runs</p></li><li><p>Positive, flexible, and trust-based work environment that encourages long-term professional and personal growth</p></li><li><p>A global, multi-culturally diverse group of colleagues and customers</p></li><li><p>Comprehensive health plans and perks</p></li><li><p>A healthy work-life balance that accommodates individual and family needs</p></li><li><p>Competitive stock options program and location-based salary</p></li></ul><h2>Who we are:</h2><p><a href='http://Sanity.io'>Sanity.io</a> is a modern content operating system that replaces rigid legacy content management systems. We treat content as data, so teams can keep one governed source of truth and adapt it across websites, apps, workflows, and AI agents with less duplicated content work.</p><p>Sanity recently raised an $85m Series C led by GP Bullhound and is backed by ICONIQ Growth, Threshold Ventures, Heavybit, Shopify, and founders from Vercel, WP Engine, Twitter, Mux, Netlify, and Heroku.</p><p>Sanity is a 200+ person company with committed, ambitious people. We are pioneers, we exist for our customers, we are hel ved, and we love type 2 fun. Read more about our values here.</p><p><a href='http://Sanity.io'><em>Sanity.io</em></a><em> pledges to be an organization that reflects the globally diverse audience our product serves. We believe that hiring the best talent and bringing together a diversity of perspectives, ideas, and cultures leads to better products and services. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, marital status, disability, or gender identity.</em></p>
CT, DE, FL, GA, ME, MD, MA, NH, NJ, NY, NC, RI, SC, VA, DC
DevOps/SysAdmin