Senior Site Reliability Engineer, Production Engineering

Posted 6 Days Ago
Easy Apply
Be an Early Applicant
Lisbon
Senior level
Cloud • Software
We deliver visibility from switch to SaaS and everything in between—so you can deliver flawless digital experiences.
The Role
As a Senior Site Reliability Engineer, you will be responsible for enhancing the reliability and performance of the ThousandEyes platform by managing large-scale distributed systems in the cloud and collaborating with application development teams. You'll design scalable operations tooling, implement AWS cloud-native services, and improve incident response processes.
Summary Generated by Built In

Please note that we have a hybrid approach to work and would like to find someone who can come into our offices in Lagoas Park once a week.Who We Are

Cisco ThousandEyes is a leading Digital Experience Assurance platform that empowers organizations to deliver seamless digital experiences across every network—even those beyond their ownership. Leveraging AI and an unparalleled set of cloud, internet, and enterprise network telemetry data, ThousandEyes enables IT teams to proactively detect, diagnose, and resolve issues before they impact end-user experiences. ThousandEyes is deeply integrated across Cisco's extensive technology portfolio, supporting customers in scaling deployments while offering AI-powered assurance insights within Cisco’s Networking, Security, Collaboration, and Observability portfolios.

About The Role

We are seeking a skilled Senior Site Reliability Engineer (SRE) in Production Engineering with a strong background in SaaS and operations. You will design and manage large-scale, highly available distributed systems in the cloud, collaborating directly with application development teams to enhance the reliability, performance, and security of our platform.

Key Responsibilities

  • Identify and provide solutions to common obstacles hindering operational excellence across engineering teams.
  • Partner with application developers using cloud-native tools to address novel challenges around scale, performance, and reliability.
  • Generalize and standardize solutions and processes to enable repeated success across our microservice-based multi-region platform.
  • Play a key role in the ThousandEyes platform by leveraging scale testing, additional environments, and working with application teams to improve system reliability.
  • Use cloud-native observability and reliability tools such as Prometheus, Istio, and ArgoCD.
  • Manage a rapidly growing infrastructure capable of handling substantial daily data volumes, emphasizing operations/infrastructure/everything as code.

What You’ll Do

  • Collaborate with software engineers to ensure architecture and services are optimized for availability, latency, and performance.
  • Design and implement scalable operations tooling to support platform growth and scaling across multiple regions.
  • Design, deploy, and maintain AWS cloud-native services that are elastic and resilient to failure.
  • Participate in and improve our 24x7 incident response and on-call rotation.
  • Use and expand our existing CNCF solutions like Kubernetes, Service Mesh, Prometheus, OpenTelemetry, and ArgoCD to increase platform reliability.
  • Automate production operations to provide guardrails and continuous platform operation.
  • Develop automation solutions for scalable service and platform operations, including deployment, scale testing, graceful failure, and chaos testing.
  • Stay updated on industry best practices for scalability and reliability to improve the scalability of the ThousandEyes platform.

Required Qualifications

  • Expert-level knowledge of Kubernetes and its ecosystem.
  • Proficiency in software development with languages such as Python or Go.
  • In-depth knowledge of cloud providers, preferably AWS.
  • Proven ability to build and implement scalable and well-tested solutions.
  • Strong understanding of Unix/Linux systems, including kernel, system libraries, file systems, and client-server protocols.
  • Knowledge of Site Reliability principles: Incident Response, Change Management, Distributed Systems, Deployment Strategies, and SLOs.
  • Excellent communication and documentation skills.
  • Strong sense of ownership, drive, and attention to detail.

Preferred Qualifications

  • Familiarity with best practices for operating a large-scale, highly available enterprise platform.
  • 5+ years of experience in a related role.

Cisco values the perspectives and skills that emerge from employees with diverse backgrounds. That's why Cisco is expanding the boundaries of discovering top talent by not only focusing on candidates with educational degrees and experience but also placing more emphasis on unlockingpotential. We believe that everyone has something to offer and that diverse teams are better equipped to solve problems, innovate, and create a positive impact.

We encourage you to apply even if you do not believe you meet every single qualification. Not all strong candidates will meet every single qualification. Research shows that people from underrepresented groups are more prone to experiencing imposter syndrome and doubting the strength of their candidacy. We urge you not to prematurely exclude yourself and to apply if you're interested in this work.

Cisco is an Affirmative Action and Equal Opportunity Employer and all qualified applicants will receive consideration for employment without regard to race, color, religion, gender, sexual orientation, national origin, genetic information, age, disability, veteran status, or any other legally protected basis. Cisco will consider for employment, on a case by case basis, qualified applicants with arrest and conviction records.

Top Skills

Go
Python

What the Team is Saying

John
Sandy
Graham
Ehsan
Christie
Chris
Surabhi
Rekha
The Company
HQ: San Francisco, CA
1,100 Employees
Hybrid Workplace
Year Founded: 2010

What We Do

As of August 7, 2020, ThousandEyes is a part of Cisco (NASDAQ: CSCO).

ThousandEyes delivers visibility into digital experiences delivered over the Internet. The world’s largest companies rely on our platform, collective intelligence and smart monitoring agents to get a real-time map of how their customers and employees reach and experience critical apps and services across traditional, SD-WAN, Internet and cloud provider networks. ThousandEyes is used by some of the world’s largest and fastest growing brands, including more than 100 of the Fortune 500, 190 of the Global 2000, 6 of the 7 top US banks, and 9 of the top 10 global software companies.

Why Work With Us

Thousand eyes it a quickly growing company with great opportunities. We empower enterprises to see, understand, and improve digital experiences for their customers and employees. We value professional development, and work with team members to achieve their career goals.

Gallery

Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery
Gallery

Cisco ThousandEyes Offices

Hybrid Workspace

Employees engage in a combination of remote and on-site work.

Typical time on-site: 20 % of the time
Company Office Image
HQSan Francisco, CA
GR
Austin, TX
Company Office Image
Cisco Cessna Business Park Office
London, GB
Oeiras, PT
Seattle, WA
Learn more

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account