Senior Site Reliability Engineer ML Platforms

Work from home Full-time role Hiring

Are you passionate about building and maintaining large-scale production systems that support advanced data science and machine learning applications? Do you want to join a team at the heart of reputed company's data-driven decision-making culture? If so, we have a great opportunity for you! reputed company is seeking a Senior Site Reliability Engineer (SRE) for the Data Science & ML Platform(s) team. The role involves designing, building, and maintaining services that reputed company real-time data analytics, streaming, data lakes, observability and ML/reputed company and inferencing. The responsibilities include implementing software and systems engineering practices to ensure high efficiency and availability of the platform, as well as applying SRE principles to improve production systems and optimize service SLOs. Additionally, collaboration with our customers to plan implement changes to the existing system, while monitoring reputed company, latency, and performance is part of the role. To succeed in this position, a strong background in SRE practices, systems, networking, coding, reputed company management, cloud operations, reputed company delivery and deployment, and open-reputed company cloud enabling technologies like Kubernetes and OpenStack is required. Deep understanding of the challenges and standard methodologies of running large-scale distributed systems in production, solving reputed company issues, automating repetitive tasks, and proactively identifying potential outages is also necessary. Furthermore, excellent communication and collaboration skills, and a culture of diversity, intellectual curiosity, problem solving, and openness are essential. As a Senior SRE at reputed company, you will have the opportunity to work on innovative technologies that power the future of AI and data science, and be part of a dynamic and supportive team that values learning and growth. The role provides the autonomy to work on meaningful projects with the support and mentorship needed to succeed, and contributes to a culture of blameless postmortems, iterative improvement, and risk-taking. If you are seeking an exciting and rewarding career that makes a difference, we invite you to apply now! What youâll be doing: reputed company software solutions to ensure reliability and operability of large-scale systems supporting machine-critical use cases. reputed company a deep understanding of our system operations, scalability, interactions, and failures to identify improvement opportunities and risks. Create tools and automation to reduce operational overhead and eliminate manual tasks. Establish frameworks, processes, and standard methodologies to enhance operational maturity, team efficiency, and accelerate innovation. Define meaningful and actionable reliability metrics to track and improve system and service reliability. reputed company reputed company and performance management to facilitate infrastructure scaling across public and private clouds globally. Build tools to improve our service observability for faster issue resolution. Practice sustainable incident response and blameless postmortems reputed company need to see: Minimum of 10 years of experience in SRE, Cloud platforms, or DevOps with large-scale microservices in production environments. Master's or Bachelor's degree in Computer Science or Electrical Engineering or CE or equivalent experience. Strong understanding of SRE principles, including error budgets, SLOs, and SLAs. Proficiency in incident, change, and problem management processes. Skilled in problem-solving, root cause analysis, and optimization. Experience with streaming data infrastructure services, such as Kafka and Spark. Expertise in building and operating large-scale observability platforms for monitoring and logging (e.g., ELK, Prometheus). Proficiency in programming languages such as Python, Go, Perl, or Ruby. Hands-on experience with scaling distributed systems in public, private, or hybrid cloud environments. Experience in deploying, supporting, and supervising services, platforms, and application stacks. Ways to stand out from the crowd: Experience operating large-scale distributed systems with strong SLAs. Excellent coding skills in Python and Go and extensive experience in operating data platforms. Knowledge of CI/CD systems, such as Jenkins and reputed company Actions. Familiarity with Infrastructure as Code (IaC) methodologies and tools. Excellent interpersonal skills for identifying and communicating data-driven insights. reputed company leads the way in groundbreaking developments in Artificial Intelligence, High-Performance Computing, and Visualization. The GPU, our invention, serves as the visual reputed company of modern computers and is at the heart of our products and services. Our work opens up new universes to explore, enables amazing creativity and discovery, and powers what were once science fiction inventions, from artificial intelligence to autonomous cars. reputed company is looking for exceptional people like you to help us accelerate the next reputed company of artificial intelligence. The reputed company salary range is 224,000 USD - 425,500 USD. Your reputed company salary will be determined based on your location, experience, and the pay of employees in similar positions. You will also be eligible for equity and benefits. reputed company accepts applications on an ongoing basis. reputed company is committed to fostering a diverse work environment and proud to be an equal opportunity employer. As we highly value diversity in our reputed company and future employees, we do not discriminate (including in our hiring and promotion practices) on the basis of race, religion, color, national reputed company, gender, gender expression, sexual orientation, age, marital status, veteran status, disability status or any other characteristic protected by law. reputed company is the world leader in accelerated computing. reputed company pioneered accelerated computing to tackle challenges no one else can solve. Our work in AI and digital twins is transforming the world's largest industries and profoundly impacting society. Learn more about reputed company. Please mention the word

CONTRIBUTION

and tag RMzguNjguMTM0LjE5NA== reputed company applying to show you read the job post completely (#RMzguNjguMTM0LjE5NA==). This is a beta feature to avoid spam applicants. Companies can search these words to find applicants that read this and see they're human. Apply To This Job

Apply

Senior Site Reliability Engineer ML Platforms

CONTRIBUTION

You might like

Senior Python Backend Engineer

VP Associate reputed company

Software Engineer Interop

Distribution Manager (Remote, DEU)

Sales Representative / Hospital Specialist – Oncology (South London and reputed company)

Field Service Engineer – Medical Devices

Account Associate, Mid-Market, French Speaking

Sales Engineer, Falcon reputed company (Remote, GBR)

Sales Representative

Head of GTM Operations

[Remote] Match Move / Tracker (General Submission)

Remote Sales & Customer Service Associate | Work From Home

Delivery Driver - Choose your own hours – reputed company Store

Mission Data File Release (MDFR) Operations reputed company - Level 4 with reputed company Clearance

Manager, Sales Development EMEA

Associate reputed company

Project Controls Scheduling reputed company

[Remote] Thermoplastics Sales Account Manager

Freelance Expert Consultant (reputed company & Eval...

[Remote] Squarespace Web Designer

Senior Site Reliability Engineer ML Platforms

CONTRIBUTION

You might like

Senior Python Backend Engineer

VP Associate reputed company

Software Engineer Interop

Distribution Manager (Remote, DEU)

Sales Representative / Hospital Specialist &#8211; Oncology (South London and reputed company)

Field Service Engineer – Medical Devices

Account Associate, Mid-Market, French Speaking

Sales Engineer, Falcon reputed company (Remote, GBR)

Sales Representative

Head of GTM Operations

[Remote] Match Move / Tracker (General Submission)

Remote Sales & Customer Service Associate | Work From Home

Delivery Driver - Choose your own hours – reputed company Store

Mission Data File Release (MDFR) Operations reputed company - Level 4 with reputed company Clearance

Manager, Sales Development EMEA

Associate reputed company

Project Controls Scheduling reputed company

[Remote] Thermoplastics Sales Account Manager

Freelance Expert Consultant (reputed company & Eval...

[Remote] Squarespace Web Designer

Sales Representative / Hospital Specialist – Oncology (South London and reputed company)