Back to All Jobs

Remote DevOps/SysAdmin Jobs in the United States

Explore remote DevOps/SysAdmin jobs at US companies hiring nationwide. Apply to roles like DevOps Engineer, Site Reliability Engineer, Platform Engineer, Systems Administrator. The jobs are sourced directly from company career pages and updated every 48 hours.

Reset
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.
Lambda company logo
Senior Site Reliability Engineer - SDN
Oct 1
Lambda
Oct 1
Multiple States
Hybrid
240k-312k / Yr
<p>Lambda, The Superintelligence Cloud, is a leader in AI cloud infrastructure serving tens of thousands of customers. Our customers range from AI researchers to enterprises and hyperscalers. Lambda's mission is to make compute as ubiquitous as electricity and give everyone the power of superintelligence. One person, one GPU.</p><p>If you'd like to build the world's best AI cloud, join us.</p><p>*Note: This position requires presence in our San Francisco/San Jose/Bellevue office location 4 days per week; Lambda’s designated work from home day is currently Tuesday.</p><p>Engineering at Lambda is responsible for building and scaling our cloud offering. Our scope includes the Lambda website, cloud APIs and systems as well as internal tooling for system deployment, management and maintenance.</p><h2>What You'll Do</h2><ul><li><p>Operate and scale Lambda’s multi-tenant cloud networking platform and SDN infrastructure</p></li><li><p>Operate and improve Kubernetes-based control plane services and dataplane software running on SmartNICs</p></li><li><p>Develop tooling and automation to reduce operational toil and improve reliability</p></li><li><p>Collaborate with software, platform, and networking teams to improve service reliability and deployment workflows</p></li><li><p>Deploy and maintain network monitoring, observability, and management tools</p></li><li><p>Improve deployment safety through CI/CD pipelines, GitOps workflows, testing, and progressive rollouts</p></li><li><p>Drive operational excellence through observability, incident management, capacity planning, postmortems, and participation in the on-call rotation</p></li></ul><h2>You</h2><ul><li><p>Have 5+ years of experience in Site Reliability Engineering, Production Engineering, or a similar role</p></li><li><p>Have experience operating and supporting large-scale distributed systems in production</p></li><li><p>Have experience with Kubernetes application lifecycle management, upgrades, troubleshooting, and production operations</p></li><li><p>Have experience participating in on-call rotations and incident response</p></li><li><p>Have strong troubleshooting skills across Linux systems, Kubernetes, distributed systems, and networking</p></li><li><p>Have experience with observability platforms, monitoring, alerting, and metrics</p></li><li><p>Are comfortable working on the Linux command line and have a solid understanding of the Linux networking stack</p></li><li><p>Have experience with multi-datacenter and hybrid cloud environments</p></li><li><p>Have experience automating infrastructure and operational workflows using Python, Ansible, or similar tools</p></li><li><p>Have experience designing and operating CI/CD and GitOps deployment workflows</p></li></ul><h2>Nice To Have</h2><ul><li><p>Experience building and operating Software Defined Networks (SDN), including OpenStack Neutron, OVN, and OVS</p></li><li><p>Experience operating production-scale SDNs in a cloud environment (e.g., infrastructure powering AWS VPC-like networking services)</p></li><li><p>Software development experience in Go and/or Python (C is a plus)</p></li><li><p>Experience automating infrastructure and network configuration using Kubernetes, Helm, Terraform, and Ansible</p></li><li><p>Deep understanding of the Linux networking stack and its interaction with network virtualization technologies, SR-IOV, and DPDK</p></li><li><p>Understanding of the SDN ecosystem and modern cloud networking architectures</p></li><li><p>Experience diagnosing complex production issues across infrastructure, networking, and application layers</p></li></ul><h2>Salary Range Information</h2><p>The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.</p><h2>About Lambda</h2><ul><li><p>Founded in 2012, with 500+ employees, and growing fast</p></li><li><p>Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove</p></li><li><p>We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG</p></li><li><p>Our values are publicly available: <a href='https://lambda.ai/careers'>https://lambda.ai/careers</a></p></li><li><p>We offer generous cash & equity compensation</p></li><li><p>Health, dental, and vision coverage for you and your dependents</p></li><li><p>Wellness and commuter stipends for select roles</p></li><li><p>401k Plan with 2% company match (USA employees)</p></li><li><p>Flexible paid time off plan that we all actually use</p></li></ul><h2>Equal Opportunity Employer</h2><p>Lambda is an Equal Opportunity employer. Applicants are considered without regard to race, color, religion, creed, national origin, age, sex, gender, marital status, sexual orientation and identity, genetic information, veteran status, citizenship, or any other factors prohibited by local, state, or federal law.</p>
CA, WA
DevOps/SysAdmin
Planet company logo
Senior Site Reliability Engineer
Sep 9
Planet
Sep 9
All U.S.
Remote & Hybrid Options
142.8k-178.5k / Yr
<p><strong>Welcome to Planet. We believe in using space to help life on Earth.</strong></p><p>Planet designs, builds, and operates the largest constellation of imaging satellites in history. This constellation delivers an unprecedented dataset of empirical information via a revolutionary cloud-based platform to authoritative figures in commercial, environmental, and humanitarian sectors. We are both a space company and data company all rolled into one.</p><p>Customers and users across the globe use Planet's data to develop new technologies, drive revenue, power research, and solve our world’s toughest obstacles.</p><p>As we control every component of hardware design, manufacturing, data processing, and software engineering, our office is a truly inspiring mix of experts from a variety of domains.</p><p>We have a people-centric approach toward culture and community and we strive to iterate in a way that puts our team members first and prepares our company for growth. Join Planet and be a part of our mission to change the way people see the world.</p><p>Planet is a global company with employees working remotely world wide and joining us from offices in San Francisco, Washington DC, Germany, Austria, Slovenia, and The Netherlands.</p><p><strong>About the Role:</strong></p><p>Planet designs, builds, and operates the largest constellation of imaging satellites in history. This constellation delivers an unprecedented dataset of empirical information via a revolutionary cloud-based platform to authoritative figures in commercial, environmental, and humanitarian sectors. We are both a space company and data company rolled into one.</p><p>In this role, you will join Planet's Direct Access Service Infrastructure team, directly contributing to our next-generation Constellation as a Service platform. This platform represents a major new offering to our customers that goes beyond traditional cloud-based platforms and supports on-premises deployments. </p><p>You will be responsible for building, deploying, and operating critical compute software that supports end-to-end imaging operations within customer on-premises and/or cloud environments. You will use your understanding of internal compute requirements as well as customers' environmental-specific constraints to help design, implement, and support a robust system for reproducible deployments across operating environments, to guarantee the reliability, scalability, and availability of our services. To do this, you will partner closely with cross-functional engineering teams to enable and empower the integration of software solutions and the troubleshooting of distributed systems.</p><p>This is a full-time, remote position based in the United States or Canada. If located near an office, you are expected to work from that office 3 days per week.</p><p><strong>Impact You'll Own:</strong></p><ul><li>Build and deploy computing services and infrastructure in customer environments for a next-generation satellite operations and image processing end-to-end platform</li><li>Operate in a high-impact, tight knit team to architect novel systems for air-gapped deployments at scale</li><li>Clarify and surface requirements from ambiguous use cases defined by cross-functional stakeholders, including internal users and external customers</li><li>Responsible for operations such as deployments, service orchestration, and documentation for cross platform stakeholders</li><li>Scale architecture while ensuring availability of services</li><li>Improve reliability and scalability by resolving edge cases, studying failure modes, and writing tests</li><li>Participate in on-call rotations to ensure operational excellence </li></ul> <p><strong>What You Bring:</strong></p><ul><li>6+ years of experience building services that leverage cloud-native infrastructure and tooling</li><li>Bachelor’s degree in Computer Science or similar</li><li>Experience deploying and maintaining bare-metal and cloud kubernetes through tools such as Talos, RKE2, Proxmox, or k3s</li><li>Proficiency with Terraform, Ansible, Helm, Kustomize, and/or similar IaC / GitOps tooling</li><li>Experience with CI/CD tooling, such as Jenkins, GitLab CI/CD, Argo CD, or CircleCI</li><li>Experience successfully building, releasing, and supporting highly available, consistently performant services</li><li>Knowledge of hardware and network level implications of on-prem compute</li><li>Experience with platform optimization, particularly resource optimization, management, and cluster tuning in a constrained environment</li><li>Ability to observe and troubleshoot distributed systems with tools such as Alloy, Prometheus, Grafana, and OpenTelemetry</li><li>Advanced skills in Python, Bash, and other tooling as appropriate to build services and meet product goals</li><li>Excellent communication skills and the ability to work through collaboration with cross-functional engineering teams</li><li>Experience working with Jira for task management and progress tracking</li></ul> <p><strong>What Makes You Stand Out:</strong></p><ul><li>Experience with CUDA-based GPU programs</li><li>Security expertise in sensitive environments, including implementing zero-trust architectures, hardening Kubernetes clusters, conducting security audits, and deploying workloads in air-gapped environments</li></ul> <p><strong>Application Deadline:</strong></p><p>November 11, 2026 by 11:59p / 23:59 CET (Central European Time)</p><p><strong>EAR/ITAR Requirements:</strong></p><p><em>This position requires access to export-controlled information, and as such, employment (or hiring of a contractor) is contingent upon the candidate’s ability to access all applicable export-controlled information without additional export licensing being required by the Bureau of Industry and Security and/or the Directorate of Defense Trade Controls.</em></p><p><strong>Benefits While Working at Planet:</strong></p><p><em>These offerings are dependent on employment type and geographical location, based upon applicable law or company policy.</em></p><ul><li>Comprehensive Medical, Dental, and Vision plans</li><li>Health Savings Account (HSA) with a company contribution</li><li>Generous Paid Time Off in addition to holidays and company-wide days off </li><li>16 Weeks of Paid Parental Leave</li><li>Wellness Program and Employee Assistance Program (EAP)</li><li>Home Office Reimbursement</li><li>Monthly Phone and Internet Reimbursement</li><li>Tuition Reimbursement and access to LinkedIn Learning</li><li>Equity</li><li>Commuter Benefits (if local to an office)</li><li>Volunteering Paid Time Off</li></ul> <p><strong>Compensation:</strong></p><p>The US base salary range for this full-time position at the commencement of employment is listed below. Additionally, this role might be eligible for discretionary short-term and long-term incentives (bonus and equity). The final salary range is determined by job related experience, skills and location. The range displays our typical hiring range for new hire salaries in US locations only. Your recruiter can share more about the specific salary range for your preferred location during the hiring process.</p><h2>#Li-remote</h2><p>New York City + California Salary Range</p><p>$153,000&mdash;$191,300 USD</p><p>San Francisco Salary Range</p><p>$162,600&mdash;$203,200 USD</p><p>US National Salary Range</p><p>$142,800&mdash;$178,500 USD</p><p><strong>San Francisco Fair Chance Ordinance<br></strong>Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records.<strong><br></strong></p><p><strong>Why we care so much about Belonging. <br></strong>We’re dedicated to helping the whole Planet, and to do that we must strive to represent all of it within each of our offices and on all of our teams. That’s why Planet is guided by an ultimate north star of Belonging—dreaming big as we approach our ongoing work. If this job intrigues you, but you’re thinking you might not have all the qualifications, please... do apply! At Planet, we are looking for well-rounded people from around the world who can contribute to more ways than just what is listed in this job description. We don’t just fill positions, we aspire to fulfill people’s careers, most excited about folks who are motivated by our underlying humanitarian efforts. We are a few orbits around the sun before we get to where we want to be, so we hope you’re excited to come along for the ride. </p><p><strong>EEO statement: <br></strong>Planet is committed to building a community where everyone belongs and we invite people from all backgrounds to apply. Planet is an equal opportunity employer, and committed to providing employment opportunities regardless of race, religious creed, color, national origin, ancestry, physical disability, mental disability, medical condition, genetic information, marital status, sex, gender, gender identity, gender expression, pregnancy, childbirth and breastfeeding, age, sexual orientation, military or veteran status, or any other protected classification, in accordance with applicable federal, state, and local laws. <a href='https://www.eeoc.gov/sites/default/files/2023-06/22-088_EEOC_KnowYourRights6.12.pdf'>Know Your Rights.</a></p><p><strong>Accommodations: <br></strong>Planet is an inclusive community and we know that everyone has their own needs. If you have a disability or special need that requires accommodation during the hiring process, please reach out to [email protected] or contact your recruiter with your request. Your message will be confidential and we will be happy to assist you.</p><p><strong>Privacy Policy</strong>: By clicking 'Apply Now' at the top of this job posting, I acknowledge that I have read the <a href='https://www.planet.com/privacy/#california'>Planet Data Privacy Notice for California Staff Members and Applicants</a>, and hereby consent to the collection, processing, use, and storage of my personal information as described therein.</p><p><strong>Privacy Policy (European Applicants):</strong> By clicking 'Apply Now' at the top of this job posting, I acknowledge that I have read the <a href='https://www.planet.com/privacy/'>Candidate Privacy Notice GDPR Planet Labs Europe</a>, and hereby consent to the collection, processing, use, and storage of my personal information as described therein.</p><p><strong>AI in Our Interviewing Process</strong>: Planet is committed to providing an exceptional interview experience for all candidates. We currently use <a href='https://www.metaview.ai/'>Metaview</a> to better focus on candidates and less on trying to capture notes. As such, with the candidate's consent, select interviews may be recorded and include a 'Planet AI Notetaker' for transcription and summarization purposes. Should an interview involve use of AI interview technologies, the candidate will receive notification and have the ability to opt out both in advance and/or real-time. Opting out will not affect one's candidacy.</p><p><strong>Candidate AI Policy</strong>: Planet embraces Artificial Intelligence (AI) tools, and we encourage its responsible use. We understand that candidates may use various resources, including AI tools, to <em>prepare</em> for interviews and assessments. However, <em>during any live interview stage or when actively completing assessments for this position, the use of AI tools—e.g. Large Language Models (LLMs), deep fake technology, etc.—is strictly prohibited unless explicitly prompted by an interviewer or assessment instructions</em>. If you are unsure about acceptable use, please contact your recruiter for clarification. If an AI tool or similar technology is desired as an accommodation, please contact [email protected] with your request for assistance. Your message will be confidential, and we will be happy to assist you. Violation of this policy may result in disqualification of your application.</p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
AuthZed company logo
Sr. Site Reliability Engineer
Sep 9
AuthZed
Sep 9
All U.S.
Fully Remote
150k-195k / Yr
<h2>About AuthZed:</h2><p>We are the creators and maintainers of SpiceDB and the authorization infrastructure that companies around the world depend on to keep their engineering teams focused on what matters most - their own product.</p><p>We are a Series A company, fixing broken access control with products that eliminate complex permission management while delivering enterprise-scale performance and consistent access control.</p><p>AuthZed is a fully remote company with employees across the US, Canada, and Europe. We’re a hardworking and close-knit group with a software-driven culture (yep, even our GTM team understands and loves this technology)! We bring integrity to all our interactions, fostering confidence in decision making - trusting and respecting each voice on our team, every day.</p><h2>Company Values:</h2><ul><li><p><strong>Agency: </strong>Everyone should have the capability, freedom, and confidence to bring about changes to our business and product. Organizational processes exist to clearly define our goals, but not restrict how progress is made.</p></li><li><p><strong>Collaboration:</strong> Success is defined in various dimensions and no single person can be an expert in all of them. Without valuing the opinions of others, finding compromises, and sharing mutual trust and respect, you cannot arrive at the best possible solution.</p></li><li><p><strong>Open-mindedness: </strong>Without asking questions, testing assumptions, and questioning our pre-existing biases we risk operating within an echo-chamber. We celebrate the representation of diverse perspectives and backgrounds as a catalyst for creating an inclusive work environment that everyone can appreciate.</p></li></ul><h2>About the Role:</h2><p>As a Site Reliability Engineer, you will play a critical role in ensuring the reliability, availability, and performance of our systems. You will be responsible for designing, implementing, and maintaining scalable infrastructure solutions to support our growing customer base. This is an exciting opportunity to work in a fast-paced environment and contribute to the success of a company bringing a Google-inspired authorization system to companies around the globe.</p><h2>What you’ll own:</h2><ul><li><p>Design, implement, and maintain highly available and scalable infrastructure solutions for our projects, products, and customers.</p></li><li><p>Monitor and analyze system performance, identifying and resolving bottlenecks and issues to ensure optimal performance and reliability.</p></li><li><p>Automate infrastructure deployment and configuration management processes.</p></li><li><p>Continuously improve system reliability, security, and efficiency through proactive monitoring, capacity planning, and performance tuning.</p></li><li><p>Troubleshoot and resolve complex infrastructure and application issues in production and test environments.</p></li><li><p>Collaborate with software engineering teams to design and implement systems that are resilient, scalable, and secure.</p></li><li><p>Participate in on-call rotation and respond to production incidents in a timely manner.</p></li><li><p>Document system configurations, troubleshooting procedures, and operational guidelines.</p></li></ul><h2>What you bring:</h2><ul><li><p>Proven experience as a Site Reliability Engineer or in a similar role.</p></li><li><p>Strong understanding of networking, operating systems, and cloud infrastructure.</p></li><li><p>Experience with Site Reliability Engineering, System Design, and Distributed Computing.</p></li><li><p>Experience in various programming languages — we currently have SDKs for NodeJS, Java, Python, Ruby, and Go.</p></li><li><p>Experience with containerization technologies such as Docker and Kubernetes.</p></li><li><p>Knowledge of infrastructure-as-code tools like Terraform and Pulumi.</p></li><li><p>Familiarity with monitoring and logging tools (e.g., Prometheus, Grafana, ELK stack).</p></li><li><p>Experience with lower-level implementation details of relational databases (bonus if you have have experience with distributed SQL databased like Google Cloud Spanner or CockroachDB).</p></li><li><p>Experience working with Git and GitHub.</p></li><li><p>Experience with continuous integration and deployment systems.</p></li><li><p>Strong problem-solving and troubleshooting skills.</p></li><li><p>Excellent communication and collaboration abilities.</p></li></ul><h2>Extra shine:</h2><ul><li><p>Experience with Authorization systems.</p></li></ul><h2>Life at AuthZed:</h2><ul><li><p>Opportunity to work with cutting-edge technology in a rapidly growing sector.</p></li><li><p>A supported environment where your ideas lead to real impact.</p></li><li><p>Competitive salary based on experience.</p></li><li><p>Stock options at an early-stage startup.</p></li><li><p>Comprehensive benefits including healthcare (US-based) and other insurance.</p></li><li><p>A full remote and flexible schedule to accommodate different timezones</p></li><li><p>Twice-yearly travel for team offsites focused on team bonding, collaboration, and having fun!</p></li></ul>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Alpaca company logo
Senior DevOps Engineer
Sep 8
Alpaca
Sep 8
All U.S.
Fully Remote
<p><strong>Who We Are:</strong></p><p><strong>Alpaca is a US-headquartered, global leader in agent-first brokerage infrastructure </strong>for stocks, ETFs, options, crypto, fixed income, 24/5 trading, and more.<br><br>Amongst our subsidiaries, Alpaca is a licensed financial services company, serving hundreds of financial institutions across 40 countries with our institutional-grade APIs. This includes broker-dealers, investment advisors, wealth managers, hedge funds, and crypto exchanges, totalling over 10 million brokerage accounts.<br><br>Our global team is a diverse group of experienced engineers, traders, and brokerage professionals who are working to achieve our mission of opening financial services to everyone on the planet. We're deeply committed to open-source contributions and fostering a vibrant community, continuously enhancing our award-winning, developer-friendly API and the robust infrastructure behind it.<br><br><strong>Alpaca is proudly backed by $400 million in funding from top-tier global investors including Portage Ventures, Spark Capital, Tribe Capital, Social Leverage, Horizons Ventures, Opera Tech Ventures, SBI Group, Derayah Financial, Unbound, Peak XV, Elefund, and Y Combinator.</strong><br><br><strong>Our Team Members:</strong></p><p>We're a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world, with teammates spanning the USA, Canada, Japan, Hungary, Nigeria, Brazil, the UK, and beyond!<br><br>We're searching for passionate individuals eager to contribute to Alpaca's rapid growth. If you align with our core values—Stay Curious, Have Empathy, and Be Accountable—and are ready to make a significant impact, we encourage you to apply.</p><h2>Role</h2> <p>As a Senior DevOps Engineer you will design, build and operate the infrastructure that lets Alpaca scale globally and run trading-critical systems with confidence. You will have the autonomy to design and implement solutions against clearly defined goals - and a real voice in shaping those goals with the team.</p><p>We are not hiring a specialist in any single tool. We are looking for a well-rounded infrastructure engineer who thinks in cloud architecture and Infrastructure-as-Code, with a genuine Platform-as-a-Product mindset: someone who measures success by how quickly and safely the rest of engineering can ship, and who treats manual toil as a bug to be engineered away. You are comfortable operating our data stores (PostgreSQL, Message Brokers) at an operator level, partnering with our SRE and database specialists on the deeper work.</p><p><strong>Things You Get To Do</strong></p><ul><li><strong>Design and evolve </strong>our cloud architecture on GCP - networking, interconnects, IAM and high-availability topology - and express it entirely as code with Terraform, following GitOps as a first principle.</li><li><strong>Build and own </strong>the CI/CD pipelines that plan, review, test and safely apply IaC changes - Policy-as-Code guardrails, drift detection and progressive rollout so infrastructure changes ship as confidently as application code.</li><li><strong>Advance </strong>Platform-as-a-Product: build self-serve capabilities and paved paths so engineers can provision what they need, through a golden path rather than a hand-off.</li><li><strong>Strengthen </strong>our observability stack - metrics, logs, traces and alerting across Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - so the platform is easy to run and reason about.</li><li><strong>Operate </strong>our GKE clusters and the infrastructure services that run on them - Helm-packaged workloads, message brokers (RabbitMQ, IBM MQ) and data stores.</li><li><strong>Participate </strong>in our Follow-The-Sun on-call model: watch and triage alerts, join and declare incidents, lead structured debugging and escalation, and drive blameless post-mortems and the post-actions that actually close the loop.</li><li><strong>Embed </strong>SRE practices - SLIs/SLOs and error budgets, capacity planning - into how Core Infrastructure builds and operates, working closely with our SRE function.</li></ul> <p><strong>Who You Are (Must-Haves)</strong></p><ul><li>5+ years in a DevOps, Platform/Infrastructure, or SRE role, with a proven track record operating large-scale, high-availability, high-performance systems in production.</li><li>Deep hands-on experience designing cloud architecture on Google Cloud Platform (GCP) as the primary cloud - landing zones, networking, IAM and high-availability topology.</li><li>Strong Infrastructure-as-Code skills with Terraform, structuring large codebases across multiple environments, with GitOps as a first principle and least-privilege as a default mindset.</li><li>Proven experience building CI/CD pipelines for IaC - automated plan/apply, code review, Policy-as-Code, drift detection and safe rollout.</li><li>Significant production experience with Kubernetes (ideally GKE) and packaging/deploying workloads with Helm.</li><li>Solid cloud and L3/L4-L7 networking fundamentals (VPCs, routing, load balancing, DNS, TLS, interconnects) and comfort debugging cross-service connectivity.</li><li>Hands-on experience with a modern observability stack - Prometheus, Thanos, Grafana, Loki, Tempo and Alertmanager - across metrics, logs, traces and alerting.</li><li>Operator-level familiarity with data stores such as PostgreSQL and Message Brokers (e.g. RabbitMQ, RedPanda) - able to run and troubleshoot them in production.</li><li>A good understanding of SRE practices - SLOs/error budgets, capacity planning - and a Platform-as-a-Product mindset.</li><li>Strong grasp of incident management end to end: joining and declaring incidents, structured debugging under pressure, escalation, clear documentation, and post-mortems that drive real change.</li><li>Able and willing to take part in a Follow-The-Sun on-call rotation from APAC hours, and to work effectively in a distributed, async-first team with strong written communication.</li></ul> <p><strong>Who You Might Be (Bonus Points)</strong></p><p>You can succeed in this role without all of the below, but any of these will help you ramp faster:</p><ul><li>Policy-as-code and IaC quality tooling (OPA/Conftest, Checkov, tflint, Atlantis, or similar).</li><li>Experience managing Terraform state, module registries and versioning at scale across many teams.</li><li>Experience building self-serve developer platforms and internal golden paths (e.g. with Backstage, Tilt, or similar).</li><li>Experience with the Alloy collector and with incident tooling such as Rootly.</li><li>Working proficiency in Go for automation and tooling.</li><li>Strong Linux (Debian/Ubuntu) and container (Docker/containerd) fundamentals.</li><li>Security and compliance experience in a regulated environment (SOC 2, secrets management, audit logging).</li><li>Familiarity with trading, brokerage, or other regulated fintech domains, and with low-latency systems.</li></ul><p><h2>How We Take Care of You:</h2> <ul><li>Competitive Salary &amp; Stock Options</li><li>Health Benefits</li><li>New Hire Home-Office Setup: One-time USD $500</li><li>Monthly Stipend: USD $150 per month via a Brex Card</li></ul> <p><em>Alpaca is proud to be an equal opportunity workplace dedicated to pursuing and hiring a diverse workforce.<br></em></p><p><a href='https://files.alpaca.markets/disclosures/AlpacaRecruitmentPrivacyPolicy.pdf'><em>Recruitment Privacy Policy</em></a></p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Cloudbeds company logo
DevOps Engineer
Sep 8
Cloudbeds
Sep 8
All U.S.
Fully Remote
90k-120k / Yr
<p><strong>What Makes Cloudbeds Unique</strong><br>At Cloudbeds, we're not just building software, we’re transforming hospitality. Our intelligently designed platform powers properties across 150 countries, processing billions in bookings annually. From independent properties to hotel groups, we help hoteliers transform operations and uplevel their commercial strategy through a unified platform that integrates with hundreds of partners. And we do it with a completely remote team. Imagine working alongside global innovators to build AI-powered solutions that solve hoteliers' biggest challenges. Since our founding in 2012, we've become the World's Best Hotel PMS Solutions Provider and landed on Deloitte's Technology Fast 500 again in 2024, but we're just getting started. </p><p><strong>Location: </strong>Remote in the USA </p><h2>What You Will Do:</h2> <ul><li>Architect and improve Cloudbeds’ infrastructure in AWS.</li><li>Help build the Kubernetes platform.</li><li>Automate the platform with infrastructure-as-code (IaC) and configuration management.</li><li>Drive successful product and engineering projects to production.</li><li>Support software development processes via CI/CD pipelines.</li><li>Evolve system security, reliability, scalability, performance, and quality.</li><li>Support development teams with sharing DevOps best practices and expertise, assist in environment and application configuration from the infrastructure perspective.</li><li>On-call rotation support for the production environment outages. </li></ul> <h2>You’ll Succeed With:</h2> <ul><li>Bachelor’s degree in Computer Science or related field, or equivalent experience.</li><li>3+ years of experience as a DevOps Engineer with AWS (EC2, EKS, S3, RDS, SQS, CloudWatch, IAM).</li><li>2+ years of strong Experience in Kubernetes (preferably EKS), Helm and Docker.</li><li>Strong Experience with application containerization methodologies and delivery.</li><li>Experience with IaC methodologies such as Terraform.</li><li>Experience with designing, building, and supporting CI/CD pipelines (GitHub Actions, and ArgoCD/gitops).</li><li>Experience with web application servers (NGiNX, Ingress controllers, traffic load balancing), databases (MySQL, PostgreSQL, Aurora), cache technologies (any of Redis, Memcached), and queue technologies (SQS).</li><li>Ability to write Bash/Python scripts.</li><li>Good networking skills.</li><li>Good written and verbal communication in English.</li><li>Good team player qualities.</li><li>Ability to work remotely and manage your own time in a global team.</li></ul> <h2>Nice to Haves:</h2> <ul><li>Experience in Linux system administration.</li><li>Experience working in a PCI-compliant environment.</li><li>Experience automating service onboarding (deployment, resources provisioning).</li><li>Exposure to monitoring and alerting tools (e.g., Prometheus, Grafana, CloudWatch, DataDog, OpenTelemetry).</li></ul> <p><em>Depending on your skills and experience, you can expect your annual compensation to be between <strong>$90,000 and $120,000 USD</strong>. Please note that we maintain high flexibility, and packages above this range will be considered. (The salary is only for US-based candidates).</em></p><p>#LI-IK1 </p><p><strong>What to Expect - Your Journey with Us</strong><br>Behind Cloudbeds' revolutionary technology is a team of redefining what's possible in hospitality. We're 650+ employees across 40+ countries, bringing together elite engineers, AI architects, world-class designers, and hospitality veterans to solve challenges others haven't dared to tackle. Our diverse team speaks 30+ languages, but we all share one language: a passion for innovation and travel. From pioneering breakthroughs in machine learning to revolutionizing how hotels operate, we're not just watching the future of hospitality unfold – we're coding it, designing it, writing it, and shipping it. If you're ready to work alongside some of the brightest minds in tech who are obsessed with using AI to transform a trillion-dollar industry, this is your chance to be part of something extraordinary.</p><p>Learn more online at <a href='https://www.cloudbeds.com/'>cloudbeds.com</a> </p><p><strong>Company <a href='https://www.cloudbeds.com/press/?press_type=company-news&amp;news_category=awards'>Awards</a>:</strong></p><ul><li>Best All-In-One Hotel Management System | HotelTechAwards (2025)</li><li>Overall 10 Best Places to Work | HotelTechAwards (2025)</li><li>Most Loved Workplace® Certified (2024)</li><li>Top 10 People’s Choice (2024)</li><li>Deloitte Technology Fast 500 (2024)</li></ul> <h2>Discover our Benefits:</h2> <ul><li>Remote First, Remote Always</li><li>PTO in accordance with local labor requirements</li><li>Monthly Wellness Fridays - enjoy an extra-long weekend every month</li><li>Fully Paid Parental Leave</li><li>Home office stipend based on country of residency</li><li>Professional development courses in Cloudbeds University</li><li>Access to professional development, including manager training, upskilling and knowledge transfer </li></ul> <p><strong>Everyone is Welcome - A Culture of Inclusion </strong> </p><p>Cloudbeds is proud to be an Equal Opportunity Employer that celebrates the diversity in our global team! We do not discriminate based upon race, religion, color, national origin, gender (including pregnancy, childbirth, or related medical conditions), sexual orientation, gender identity, gender expression, age, status as a protected veteran, status as an individual with a disability, or other applicable legally protected characteristics.</p><p>Cloudbeds is committed to the full inclusion of all qualified individuals. As part of this commitment, Cloudbeds will ensure that persons with disabilities are provided reasonable accommodations in the hiring process. We encourage deaf, hard-of-hearing, deaf-blind, and deaf-disabled individuals to apply. If reasonable accommodation is needed to participate in the job application or interview process or to perform essential job functions, please contact our HR team by phone at (858) 201-7832 or via email at <a href='mailto:[email protected]'>[email protected]</a>. Cloudbeds will provide an American Sign Language (ASL) interpreter where needed as a reasonable accommodation for the hiring processes.</p><p>To all Staffing and Recruiting Agencies: Our Careers Site is only for individuals seeking a job at Cloudbeds. Staffing, recruiting agencies, and individuals being represented by an agency are not authorized to use this site or to submit applications, and any such submissions will be considered unsolicited. Cloudbeds does not accept unsolicited resumes or applications from agencies. Please do not forward resumes to our jobs alias, Cloudbeds employees, or any other company location. Cloudbeds is not responsible for any fees related to unsolicited resumes/applications.</p><p>#Li-remote</p>
AL, AK, AZ, AR, CA, CO, CT, DE, FL, GA, HI, ID, IL, IN, IA, KS, KY, LA, ME, MD, MA, MI, MN, MS, MO, MT, NE, NV, NH, NJ, NM, NY, NC, ND, OH, OK, OR, PA, RI, SC, SD, TN, TX, UT, VT, VA, WA, DC, WV, WI, WY
DevOps/SysAdmin
Sanity company logo
Senior Site Reliability Engineer
Sep 5
Sanity
Sep 5
Multiple States
Fully Remote
<p>At <a href='http://Sanity.io'><strong><u>Sanity.io</u></strong></a>, we’re building the future of AI-powered Content Operations. Our AI Content Operating System gives teams the freedom to model, create, and automate content the way their business works, accelerating digital development and supercharging content operations efficiency. <br><br>It's always peak hour somewhere: with infrastructure and customers spanning every continent, a Sanity SRE makes sure the platform we build is scalable and fast, safe to deploy, and inspiring to use. The scale is real: Content Lake alone handles around 75,000 requests a second, about 4.5m a minute, and companies like <strong>Skims</strong>, <strong>Figma</strong>, <strong>Riot Games</strong>, <strong>Anthropic</strong>, <strong>Complex</strong>, <strong>Nordstrom</strong>, <strong>Arc’teryx,</strong> and <strong>Morningbrew</strong> run their content operations on it.</p><p>Our stack is built on a mix of the tried and tested and the bleeding edge. The core technologies we currently use include Kubernetes, Prometheus, ElasticSearch, PostgreSQL, NATS, Kong, Fastly, and Google Cloud Platform.</p><p>The SRE role involves close partnership with our development teams to design and build infrastructure that supports our goal: to be the best platform for authoring, processing, and distributing content worldwide in real time. You will work close to the metal on the security, stability, and performance our customers have come to expect, and help raise the reliability bar as we grow.</p><h2>What you would do:</h2><ul><li><p>Design, build, and operate the shared platform foundations engineers ship on every day: GCP infrastructure, Kubernetes, networking, routing, CI/CD, and observability.</p></li><li><p>Diagnose and troubleshoot complex distributed systems running at high request volume.</p></li><li><p>Ensure observability and analyze the behavior of our stack.</p></li><li><p>Contribute to in-flight work like modernizing our edge, caching, and gateway layers onto Fastly and tightening observability across the platform.</p></li><li><p>Raise the reliability bar through better dashboards, alert severity, paging standards, on-call readiness, and incident response.</p></li><li><p>Make deployment boring in the best way: build golden paths, production readiness checks, safe rollouts, and useful automation so engineers have fewer places to look before they ship.</p></li><li><p>Mentor engineers and raise the technical bar through code review, design review, and pairing.</p></li><li><p>Participate in our on-call rotation and help our developer on-call rollout land well.</p></li></ul><h2>About you:</h2><ul><li><p>Based in the United States, with reasonable overlap with European engineering hours.</p></li><li><p>Experience with SRE/DevOps tools, processes, and culture.</p></li><li><p>5+ years of experience as part of an SRE on-call rotation.</p></li><li><p>Analytical approach to designing, diagnosing, and optimizing infrastructure.</p></li><li><p>Experience with managing scalable, highly available, cloud-based applications, ideally with high request volume and customer-facing uptime expectations.</p></li><li><p>Experience with Kubernetes for orchestrating, scaling, and managing containerized applications in cloud-based environments.</p></li><li><p>Experience building CI/CD pipelines.</p></li><li><p>Experience with an observability stack (Prometheus, et al.).</p></li><li><p>Comfortable working across CDNs, edge, gateways, and caching layers, or eager to go deep there.</p></li><li><p>You improve on-call and reliability by building systems, standards, and feedback loops that make production healthier over time.</p></li><li><p>You are comfortable dealing with incidents and outages and have built a practical, thoughtful communication style for handling high-pressure situations.</p></li><li><p>An open but considered approach to new technologies.</p></li></ul><p>There are many roads leading up to being an SRE. Our team is already a mix of self-taught and formally educated people. Don't self-select out!</p><h2>What we can offer:</h2><ul><li><p>A highly-skilled, inspiring, and supportive team</p></li><li><p>Real infrastructure scale and meaningful, hands-on work changing how it runs</p></li><li><p>Positive, flexible, and trust-based work environment that encourages long-term professional and personal growth</p></li><li><p>A global, multi-culturally diverse group of colleagues and customers</p></li><li><p>Comprehensive health plans and perks</p></li><li><p>A healthy work-life balance that accommodates individual and family needs</p></li><li><p>Competitive stock options program and location-based salary</p></li></ul><h2>Who we are:</h2><p><a href='http://Sanity.io'>Sanity.io</a> is a modern content operating system that replaces rigid legacy content management systems. We treat content as data, so teams can keep one governed source of truth and adapt it across websites, apps, workflows, and AI agents with less duplicated content work.</p><p>Sanity recently raised an $85m Series C led by GP Bullhound and is backed by ICONIQ Growth, Threshold Ventures, Heavybit, Shopify, and founders from Vercel, WP Engine, Twitter, Mux, Netlify, and Heroku.</p><p>Sanity is a 200+ person company with committed, ambitious people. We are pioneers, we exist for our customers, we are hel ved, and we love type 2 fun. Read more about our values here.</p><p><a href='http://Sanity.io'><em>Sanity.io</em></a><em> pledges to be an organization that reflects the globally diverse audience our product serves. We believe that hiring the best talent and bringing together a diversity of perspectives, ideas, and cultures leads to better products and services. We are committed to equal employment opportunity regardless of race, color, ancestry, religion, sex, national origin, sexual orientation, age, marital status, disability, or gender identity.</em></p>
CT, DE, FL, GA, ME, MD, MA, NH, NJ, NY, NC, RI, SC, VA, DC
DevOps/SysAdmin

No jobs match this search

Please try again using other terms or filters.

What are the most in-demand remote jobs in America?

Right now it’s engineering jobs, though that may change as AI gets better at coding.

Do I need to be based in the United States to apply?

Generally, yes. Our job board is aimed at American candidates, so those are the location details we show. We recommend not applying from outside the US unless the job description explicitly mentions they accept candidates from your country.

Can I only apply for companies based in my state?

If the location for the job says All U.S, you can be based anywhere in the United States. Though in some states the company may have to hire you as a contractor instead of an employee.

Which companies hire remotely in the US?

Our database has thousands of companies. Here is a list of Top Companies to work for: https://www.remoteworkusa.com/top-companies

What are the benefits of a remote job?

If you apply to remote jobs in other cities or states, you are getting more opportunities than you would if you only applied to job openings in your city.

Another benefit is time saved. If you spend an hour a day in traffic, it adds up to 10 full days per year(yes, seriously).