Job Description
Delivers complete solutions spanning all phases of the Software Development Lifecycle (SDLC) (design, implementation, testing, delivery and operations), based on definitions from more senior roles. Advises immediate management on project-level issues Guides junior engineers Operates with little day-to-day supervision, making technical decisions based on knowledge of internal conventions and industry best practices Applies knowledge of technical best practices in making decisions Participate in situation room activities for new product rollouts Implement automated monitoring solutions to detect single points of failure Participate in situation room activities for new product rollouts Manage multiple concurrent incidents during peak periods with efficiency and precision Provide training and guidance to engineering teams on change management best practices 3+ years relevant experience and a Bachelor's degree OR Any equivalent combination of education and experience. Your way to Impact Rather than building or maintaining infrastructure day-to-day, our team is entrusted with a broader mandate: safeguarding the reliability, resiliency, and availability of some of the world's most heavily trafficked financial platforms. We hold final decision-making authority during high-severity incidents, partner closely with executive leadership on post-incident learnings, and drive the tooling and processes that continuously strengthen our incident response capabilities. You'll also regularly interface with executive leadership during critical incidents and post-mortems, and drive implementation of tooling that advances the Command Center's capabilities. Site Resiliency & Infrastructure Management * Proactively identify and address vulnerabilities in cloud (AWS, GCP, Azure) and on-premises infrastructure Review Infrastructure as Code changes for reliability risks as part of change approval process Identify architectural anti-patterns in Kubernetes deployments and cloud migrations Review and approve changes to production systems, ensuring comprehensive risk assessment Automate change validation and rollback procedures to minimize service disruptions Streamline change management processes to reduce manual errors and bottlenecks Provide training and guidance to engineering teams on change management best practices Maintain change audit documentation and compliance requirements * Leverage deep expertise in cloud platforms (AWS, GCP, Azure) to drive incident resolution Support Braintree and Venmo cloud infrastructure operations Guide teams toward solutions by providing architectural direction during incidents Stay current with emerging cloud technologies and best practices Mentor team members on cloud technologies and incident management techniques * Implement automation, dashboards, and tooling to enhance the team's incident response capabilities Build runbooks and playbooks for cloud-native incident scenarios Develop internal tools and scripts to improve TDO operational efficiency Drive projects that advance the Command Center's operational capabilities * Significant hands-on experience with at least one major cloud provider (AWS or GCP required; multi-cloud experience preferred) Strong proficiency with Infrastructure as Code tools (Terraform, CloudFormation, Pulumi, or equivalent) including ability to read, review, and troubleshoot IaC configurations during incidents Significant hands-on experience with Kubernetes and CNCF ecosystem tools, including troubleshooting K8s deployments, manifests, and cluster issues Ability to quickly read and review code across multiple languages (Python, Go, Bash) and configuration formats (YAML, HCL, JSON)essential for effective incident troubleshooting Proven experience managing critical incidents in Infrastructure-as-Code driven environments, including troubleshooting IaC state issues, GitOps failures, and cloud-native deployment problems Professional-level certification in at least one major cloud platform (AWS Solutions Architect Professional, Google Cloud Professional Cloud Architect, or equivalent) Experience with monitoring and observability tools (Splunk, Datadog, Prometheus, Grafana) Experience with monitoring and observability tools (Splunk, Datadog, Prometheus, Grafana) * Exceptional communication skills with ability to articulate complex technical issues to both technical and non-technical stakeholders Executive presence and ability to effectively communicate with senior leadership during high-pressure incidents and post-mortems Strong analytical and problem-solving abilities with a systematic approach to troubleshooting Ability to remain calm under pressure and make critical decisions during incidents Excellent collaboration skills with experience working across global, cross-functional teams Strong documentation skills and attention to detail