Principal Site Reliability Developer @ Oracle

Home > Devops

Principal Site Reliability Developer

Oracle
7 - 12 years
Hyderabad
4 months ago
Email to a friend
Report this job

Job Description

Role & responsibilities

At Oracle Cloud Infrastructure (OCI), we build the future of the cloud for Enterprises as a diverse team of fellow creators and inventors. We act with the speed and attitude of a start-up, with the scale and customer-focus of the leading enterprise software company in the world. Compute is one of the core organisations within OCI. We are responsible for providing Compute power i.e. VMs and BMs. Cloud pretty much cannot exists without our org. The Compute org comprises of a family of critical foundational infrastructure services that drive OCIs hardware lifecycle activities.

Work with Site Reliability Engineering (SRE) team on the shared full stack ownership of a collection of services and/or technology areas. Understand the end-to-end configuration, technical dependencies, and overall behavioral characteristics of production services. Responsible for the design and delivery of the mission critical stack, with focus on security, resiliency, scale, and performance. Authority for end-to-end performance and operability. Partner with development teams in defining and implementing improvements in service architecture. Articulate technical characteristics of services and technology areas and guide Development Teams to engineer and add premier capabilities to the Oracle Cloud service portfolio. Understand and communicate the scale, capacity, security, performance attributes, and requirements of the service and technology stack. Demonstrate clear understanding of automation and orchestration principles. Act as ultimate escalation point for complex or critical issues that have not yet been documented as Standard Operating Procedures (SOPs). Utilize a deep understanding of service topology and their dependencies required to troubleshoot issues and define mitigations. Understand and explain the affect of product architecture decisions on distributed systems. Professional curiosity and a desire to a develop deep understanding of services and technologies.

Responsibilities include but not limited to

Incident Management
Support and troubleshooting of Staging/Production environments
Response and Resolve incidents as per SLA's
Organise, Anticipate, Plan and work as On-Call in shifts for multiple services (Open to work in shifts & shows flexibility)
Maintain Service High Availability
Release Management
Test and Deploy solutions and automate to replace manual processes
Build and maintain deployment tools/procedures
Zero downtime deployments and a high availability mindset
Define and build innovative solution methodologies and assets around infrastructure, cloud migration and deployment operations at scale.
Work with service teams to resolve complex issues that require troubleshooting and knowledge of code.
Keep documentation up to date and resolving similar tickets with lower turnaround time and within SLA
Ensure production security posture
Ensure monitoring is robust and effective
Change Management
Perform Root Cause Analysis

Qualifications:

Bachelors in computer science and Engineering or related engineering fields
7+ years of experience delivering and operating large scale, highly available distributed systems.
5+ years of experience with Linux System Engineering
4+ years of experience with Python/Java building infrastructure Automations
Understanding of Networking, Cloud Computing, Load Balancers
Strong Infrastructure troubleshooting skills
Experience in CICD, Cloud Computing and networking
Hands on experience at Monitoring/Instrumentation tools (Prometheus/Grafana etc)

Preferred candidate profile

Job Classification

Industry: Software Product
Functional Area / Department: Engineering - Software & QA
Role Category: DevOps
Role: DevOps - Other
Employement Type: Full time

Contact Details:

Company: Oracle
Location(s): Hyderabad

+ View Contact

Login

Candidates can login here to view contacts and apply.

Sign In Sign Up

Email:

Password:

Password too short

To create your profile, apply for a job or make a registration

Your name (*)

Email (*)

Mobile (*)

Preferred City (* max. 2 w/comma)

Designation / Expected Role

Current / Recent Company (*)

Experience (*)

Expected Salary (*)

Desired Industry (*):

Functional area / Department (*):

Enter Skills (key skills, subjects, technologies & roles to use in search)

Write briefly about yourself, your experience and education (*)

Attach Resume Max 2.38 MB (RTF, PDF, DOC, DOCX formats only parsed)

Please, check the file size and type.

Add social media [ + ]

Create password

I agree with website service terms and conditions

Candidates are expected to provide most recent and accurate profile information, inappropriate content is strictly prohibited!

Keyskills: Linux Administration Infrastructure Management Python

Job seems aged, it may have been expired!
Fraud Alert to job seekers!

₹ Not Disclosed

Job application

We will notify the employer with your details. You can also attach a resume or a cover letter.

Sign In Sign Up

Email:

Password:

Password too short

To create your profile, apply for a job or make a registration

Your name (*)

Email (*)

Mobile (*)

Preferred City (* max. 2 w/comma)

Designation / Expected Role

Current / Recent Company (*)

Experience (*)

Expected Salary (*)

Desired Industry (*):

Functional area / Department (*):

Enter Skills (key skills, subjects, technologies & roles to use in search)

Write briefly about yourself, your experience and education (*)

Attach ResumeMax 2.38 MB (RTF, PDF, DOC, DOCX formats only parsed)

Please, check the file size and type.

Add social media [ + ]

Create password

I agree with website service terms and conditions

Similar positions

Application Developer-AWS Cloud FullStack

IBM

5 - 7 years

Bengaluru

5 days ago

₹ Not Disclosed

Application Developer-AWS Cloud FullStack

IBM

6 - 7 years

Pune

5 days ago

₹ Not Disclosed

Application Developer-AWS Cloud FullStack

IBM

6 - 8 years

Hyderabad

5 days ago

₹ Not Disclosed

Engineer Lead, Site Reliability

Zensar

6 - 10 years

Pune

5 days ago

₹ Not Disclosed

Oracle

3i Infotech is a global IT products & services company committed to Empowering Business Transformation. A comprehensive set of IP based software solutions (20+), coupled with a wide range of IT services, uniquely positions the company to address the dynamic requirements of a variety of industry ...

Principal Site Reliability Developer @ Oracle

Home > Devops