Airbnb

Senior Software Engineer, Reliability Engineering Team

🌍 Remoto📍 Remoto
PUBLICIDADE

Descrição da Vaga

 By collaborating closely with other engineering teams you will help establish a culture of reliability throughout the organization by providing a comprehensive incident management platform that is being used for instrumentation, operability, and around incidents. Your ability to identify opportunities for improvement and drive their implementation will contribute significantly to our overall operational efficiency and growth, ensuring that our services remain resilient as our business continues to expand.

Additionally, as an essential part of this role, you will serve as an active member of the Production SRE team, responding to and managing high severity incidents. Your vast technical experience and leadership skills will be invaluable as you step into the role of Incident Commander during these critical events. You will guide cross-functional teams during crisis situations and ensure timely resolution, minimizing the impact on our customers and business. This aspect of your work will require not just strong technical acumen, but also excellent communication and coordination skills, resilience under pressure, and a firm commitment to our culture of blamelessness and continuous learning.

A Typical Day: 

  • Design, implement and maintain the tools and systems that support service reliability, monitoring, and alerting.
  • Collaborate with other engineering teams to ensure services are designed with reliability in mind, and provide guidance on the appropriate use of tooling and automation.
  • Identify opportunities to improve the reliability, scalability, and efficiency of our services and drive their implementation.
  • Work with infrastructure engineers to understand the challenges they face in operating our services and develop tools and systems to help them manage these challenges.
  • Participate in incident response and post-mortems to identify and address systemic issues.
  • Continuously evaluate new technologies and industry best practices to improve our SRE tooling and incident response procedures.
  • Gain and maintain an intimate understanding of how the critical parts of the site work (services, infrastructure, product, tools, and processes)
  • Lead high-urgency incidents and mentor less-experienced engineers in effectively handling incidents.

Your Expertise:

  • Bachelor's degree in Computer Science or related field.
  • 5+ years of experience in software engineering or SRE roles, with a focus on large scale distributed systems.
  • Strong coding skills in at least one programming language, such as Java, Python, or Go.
  • Experience with distributed systems and service-oriented architectures.
  • Experience with cloud computing platforms such as AWS or Google Cloud Platform.
  • Strong conviction in software development best practices, including version control, automated testing, and continuous integration and delivery.
  • Experience with containerization technologies such as Docker and Kubernetes.
  • Excellent problem-solving and analytical skills, with a strong attention to detail.
  • Ability to work effectively in a fast-paced and dynamic environment.
  • Strong communication and interpersonal skills.
  • Fluent in English (Professional Level)
PUBLICIDADE

Airbnb is a global community of Hosts and travelers—a community that was born in 2007 when two Hosts welcomed three guests to their San Francisco home, and has since grown to 4 million Hosts who have welcomed more than 1 billion guest arrivals to about 100,000 cities in almost every country and regi...[Ver Mais]

Informações Adicionais

Selecionamos as principais informações da posição. Para conferir o descritivo completo, clique em "acessar" 


PUBLICIDADE
Candidatar-se a esta vaga →

Você será redirecionado para o site da empresa

PUBLICIDADE