Job Overview
Role: Site Reliability Engineer Location: Bangalore Experience: Freshers/Experienced Qualification: B.E/B.Tech/B.Sc/BCA or equivalent Key Skills: SRE, Cloud (AWS/Azure/GCP), Kubernetes, Python/Go, Observability, Automation
Job Description
NVIDIA is hiring a Site Reliability Engineer in Bangalore to improve the reliability, scalability, and operational efficiency of enterprise systems supporting its AI-powered products and services. This role combines software and infrastructure engineering, suitable for freshers and experienced candidates interested in SRE, cloud infrastructure, Kubernetes, and AI-powered engineering. Candidates should possess foundational programming knowledge and familiarity with cloud platforms, containerization, and observability tools.
Roles and Responsibilities
- System Reliability: Apply software engineering principles to system operations, focusing on availability, scalability, performance, and resilience.
- Infrastructure Management: Contribute to distributed systems, databases, Kubernetes-based environments, and cloud infrastructure.
- Automation & Observability: Drive automation efforts and implement observability (logging, metrics, tracing) for system health.
- Incident Response: Participate in incident response and contribute to developer productivity.
- Database Operations: Automate database operations such as provisioning, scaling, backups, and failover.
Skills and Eligibility Criteria
Educational Background: BS degree in Computer Science or a related technical field (Physics, Mathematics), or equivalent practical experience (B.E/B.Tech/B.Sc/BCA)
Experience: Freshers/Experienced candidates can apply.
Mandatory Technical Skills:
- Foundational programming knowledge in Python, TypeScript, JavaScript, or Go
- Basic understanding of cloud platforms (AWS, Azure, or GCP)
- Knowledge of containerization technologies such as Docker and Kubernetes
- Exposure to infrastructure-as-code tools (Terraform, AWS CDK, or CloudFormation) or willingness to learn
- Familiarity with Linux/Unix systems, networking fundamentals, and Git
- Interest in observability concepts including logging, metrics, and tracing
- Basic knowledge of relational databases (PostgreSQL or MySQL) and SQL
Competencies:
- Personal projects, internships, or coursework involving cloud infrastructure, DevOps, or SRE
- Open-source contributions, hackathons, or technical community participation
- Exposure to AI/ML concepts or experience building/deploying machine learning models
- Experience with LLM APIs or AI-powered developer tools
- Strong problem-solving skills, curiosity, ownership, and initiative
- Good communication and teamwork skills