Your browser does not support javascript! Please enable it, otherwise web will not work for you.

Principal Site Reliability Developer @ Oracle

Home > Devops

 Principal Site Reliability Developer

Job Description

Role & responsibilities

At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises as a diverse team of fellow creators and inventors. We act with the speed and attitude of a start-up, with the scale and customer-focus of the leading enterprise software company in the world. Compute is one of the core organisations within OCI. We are responsible for providing Compute power i.e. VMs and BMs. Cloud pretty much cannot exists without our org. The Compute org comprises of a family of critical foundational infrastructure services that drive OCIs hardware lifecycle activities.


Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.


Responsibilities include but not limited to


Incident Management
Support and troubleshooting of Staging/Production environments
Response and Resolve incidents as per SLA's
Organise, Anticipate, Plan and work as On-Call in shifts for multiple services (Open to work in shifts & shows flexibility)
Maintain Service High Availability
Release Management
Test and Deploy solutions and automate to replace manual processes
Build and maintain deployment tools/procedures
Zero downtime deployments and a high availability mindset
Define and build innovative solution methodologies and assets around infrastructure, cloud migration and deployment operations at scale.
Work with service teams to resolve complex issues that require troubleshooting and knowledge of code.
Keep documentation up to date and resolving similar tickets with lower turnaround time and within SLA
Ensure production security posture
Ensure monitoring is robust and effective
Change Management
Perform Root Cause Analysis


Qualifications:

  • Bachelors in computer science and Engineering or related engineering fields
  • 7+ years of experience delivering and operating large scale, highly available distributed systems.
  • 5+ years of experience with Linux System Engineering
  • 4+ years of experience with Python/Java building infrastructure Automations
  • Understanding of Networking, Cloud Computing, Load Balancers
  • Strong Infrastructure troubleshooting skills
  • Experience in CICD, Cloud Computing and networking
  • Hands on experience at Monitoring/Instrumentation tools (Prometheus/Grafana etc)

Preferred candidate profile

Job Classification

Industry: Software Product
Functional Area / Department: Engineering - Software & QA
Role Category: DevOps
Role: DevOps - Other
Employement Type: Full time

Contact Details:

Company: Oracle
Location(s): Hyderabad

+ View Contactajax loader


Keyskills:   Linux Administration Infrastructure Management Python

 Fraud Alert to job seekers!

₹ Not Disclosed

Similar positions

Principal DevOps Engineer - Autonomous Database

  • Oracle
  • 10 - 15 years
  • Pune
  • 5 days ago
₹ Not Disclosed

Site Reliability Developer 3

  • Oracle
  • 5 - 11 years
  • Bengaluru
  • 8 days ago
₹ Not Disclosed

Kafka Developer

  • SRSInfoway
  • 5 - 10 years
  • Bengaluru
  • 9 days ago
₹ -13 Lacs P.A.

Site Reliability Engineer - Java Stack

  • Genpact
  • 8 - 13 years
  • Noida, Gurugram
  • 9 days ago
₹ Not Disclosed

Oracle

About Accenture\\r\\n\\r\\n \\r\\n\\r\\nAccenture is a global professional services company with leading capabilities in digital, cloud and security. Combining unmatched experience and specialized skills across more than 40 industries, we offer Strategy and Consulting, Interactive, Technology and Op...