Back

Sr. Software Engineer - Cloud Infrastructure and Devops

Simility, a Paypal ServiceSan Jose, CA, USA, Austin, TX, USA +1 more
You will be redirected to the main website
Posted on: 27 Sep 2026

Job Description

Makes technical decisions affecting multiple teams, crossing organizational boundaries Establishes conventions & processes to be followed by other employees Actions determine the utilization of company resources (people, money, assets) and affect the effectiveness of the company Handles multiple, multi-team initiatives simultaneously, using judgement to prioritize among more issues than can be handled individually. Understands evolving industry capabilities & practices and can judiciously apply up--to-date information for optimal results Competent at communicating technical issues with non-technical audiences Spreads their behavior, principles, and knowledge as a means of improving technical results of other employees (via many means - modeling behavior, 1:1s, working sessions, quality documentation) Partners with product management, to ideate solutions to business problems & goals Act as a hands-on contributor, while leading by example Mentor junior engineers Enjoy a high degree of independence in decision-making while being accountable for results of yourself and your team Be part of a team dedicated to driving the scalability and reliability of Venmo's AWS cloud infrastructure Contribute to initiatives: internal teams rarely have dedicated project managers. You will define the design, identify stakeholders, navigate risks and changes, coordinate colleagues' work, implement solutions, and be accountable for timely and quality project delivery Troubleshoot incidents, identify root causes, fix and document problems, and implement preventive measures Lead by example, making meaningful contributions to the improvement of engineering teams' production operations Develop and improve tools and automation to manage infrastructure and application configuration as code Enhance the quality, reliability, and stability of our infrastructure and operations Design, implement, and operate chaos engineering experiments to proactively identify and remediate system weaknesses Build and maintain disaster recovery automation, runbooks, and infrastructure across Venmo's cloud environment Plan and execute chaos and incident game day exercises to validate system resilience and team preparedness Develop and maintain backup and restore tooling and processes for critical systems and data Define and track resilience metrics, SLOs, and recovery time objectives for owned systems 8+ years relevant experience and a Bachelor's degree OR Any equivalent combination of education and experience. Engineering at Venmo At Venmo, we are creating a product that people love. We strive to create a delightful user experience while connecting the world and empowering people through payments. We are looking for intellectually curious people who want to be inspired and inspire others to change the world. Engineering is a craft, and at Venmo we want the internals of our software to be as elegant as the end user experience we are designing. We spend our days scaling our infrastructure and building new features to meet and exceed our users' needs and wants. We work with a product team that is data-driven and human-centered in its design principles. We teach and learn from one another and push each other to be at our creative and analytical best. Bachelor's in computer science or related field of study 5+ years' experience in software development or a related field 3+ years' experience operating distributed applications 24x7x365, as part of a Cloud Engineering, DevOps, and/or SRE team Extensive hands-on experience with designing, implementing, and supporting infrastructure (AWS experience preferred) to support global-scale services Deep hands-on experience with IaaS and PaaS solutions from AWS (or similar cloud provider) Hands-on programming and scripting (Python, Java, Bash, Go) Hands-on experience with containers and container orchestration: Docker, Kubernetes Strong communication skills with the ability to understand and explain technical issues to a non-technical audience Experience with chaos engineering tools and frameworks (e.g., AWS Fault Injection Simulator, Gremlin, Chaos Monkey, Litmus) Hands-on experience designing and implementing disaster recovery solutions for distributed systems Experience developing and maintaining backup and restore tooling and strategies Experience planning and executing game day exercises or incident simulations Understanding of RTO/RPO requirements and how to design systems to meet them

Job Overview

Salary:
Not disclosed
JOB TYPE:
Not specified
Experience:
Not mentioned
Job Location:
San Jose, CA, USA, Austin, TX, USA +1 more
Job Level:
Not specified
Education:
Graduation
Sr. Software Engineer - Cloud Infrastructure and DevopsSimility, a Paypal Service