Job Description
Job Description
Job Description – Hardware Engineer
\nExperience: 5+ Years
\nRole Overview:
\nWe are looking for an experienced and hands-on Hardware Engineer to support the operations, maintenance, troubleshooting and lifecycle management of server hardware across our infrastructure. The role will involve working with both AI/GPU-based servers and conventional CPU-based enterprise servers in mission-critical data center environments.
\nThe ideal candidate should have strong experience in server hardware break-fix, fault diagnosis, component replacement, spare-parts management, RMA processes and coordination with OEM/ODM support teams.
\nKey Responsibilities:
\n1. Server Hardware Operations & Maintenance
\n- \n
- Perform end-to-end troubleshooting, maintenance and repair of AI/GPU and conventional enterprise servers. \n
- Ensure high availability and reliability of server hardware deployed across the infrastructure. \n
- Monitor and identify hardware faults, degradation and recurring failure patterns. \n
- Perform hardware diagnosis and identify faulty components for replacement. \n
- Ensure servers are tested and operational after repair, replacement or maintenance activities. \n
- Work within defined SLAs to ensure timely resolution of hardware-related incidents. \n
2. Server Hardware Break-Fix
\n- \n
- Handle the complete hardware break-fix lifecycle, including: \n
- Fault detection and diagnosis \n
- Faulty component identification \n
- Spare allocation \n
- Onsite component replacement \n
- Server testing and service restoration \n
- Faulty part removal and tagging \n
- RMA coordination and closure \n
- Troubleshoot and replace server Field Replaceable Units (FRUs), including: \n
- GPUs \n
- CPUs \n
- Motherboards \n
- Memory \n
- NICs \n
- Power Supply Units (PSUs) \n
- Fans \n
- Storage and other server components \n
3. AI/GPU and Enterprise Server Support
\n- \n
- Provide hardware support for NVIDIA GPU-based servers and compute nodes. \n
- Work on high-density AI and HPC server environments. \n
- Support GPU server platforms, including HGX/DGX architecture and other GPU-based server platforms. \n
- Perform troubleshooting and replacement of GPU cards and associated server components. \n
- Support conventional enterprise and CPU-based servers, including: \n
- General-purpose compute servers \n
- Application and database servers \n
- Virtualization servers \n
- High-performance compute servers \n
4. Spare Parts & Inventory Management
\n- \n
- Maintain and manage server hardware spares required for break-fix activities. \n
- Ensure proper receipt, inspection, storage and issue of server components. \n
- Maintain accurate records of spare inventory and component movement. \n
- Monitor availability of critical server FRUs and escalate requirements for replenishment. \n
- Support inventory reconciliation of replaced, repaired and available spare components. \n
- Ensure critical AI/GPU server components are available as per operational requirements. \n
5. Faulty Part & RMA Management
\n- \n
- Identify, tag and maintain proper records of faulty or replaced components. \n
- Follow the defined process for segregation and storage of faulty hardware. \n
- Coordinate with OEMs/ODMs for raising and tracking RMA cases. \n
- Prepare faulty components for shipment to designated OEM/ODM service centers. \n
- Track repair and replacement status of faulty components. \n
- Ensure repaired or replacement components are received and appropriately updated in the inventory. \n
- Maintain accurate documentation to avoid unaccounted or misplaced components. \n
6. OEM/ODM & Vendor Coordination
\n- \n
- Coordinate with OEMs, ODMs and hardware service partners for technical support and issue resolution. \n
- Follow up on delayed parts, replacement requests and unresolved hardware issues. \n
- Support warranty and service-related activities. \n
- Escalate critical or recurring hardware failures to the relevant internal and external teams. \n
- Coordinate with central engineering teams and onsite support teams for timely issue resolution. \n
7. Incident Management & Documentation
\n- \n
- Respond to server hardware incidents within defined response and resolution timelines. \n
- Maintain detailed records of hardware faults, repairs, replacements and RMA activities. \n
- Document troubleshooting steps and resolutions for recurring issues. \n
- Follow established operational processes, escalation procedures and SLAs. \n
- Participate in shift or standby support for critical 24x7 environments, as required. \n
Required Skills & Experience:
\n- \n
- 5+ years of experience in server hardware operations, data center hardware support or enterprise server infrastructure. \n
- Strong hands-on experience in server hardware troubleshooting and break-fix operations. \n
- Strong understanding of server hardware architecture and components. \n
- Experience in diagnosing and replacing server components such as GPUs, CPUs, motherboards, memory, NICs, PSUs, fans and other FRUs. \n
- Experience with spare-parts management and hardware inventory. \n
- Experience in faulty part handling and RMA management. \n
- Experience coordinating with OEMs, ODMs and hardware service partners. \n
- Understanding of hardware incident management and SLA-driven support environments. \n
- Experience working in 24x7 mission-critical data center or enterprise infrastructure environments. \n
- Ability to troubleshoot issues independently and coordinate with multiple technical teams. \n
Preferred Skills:
\nExperience with one or more of the following will be an added advantage:
\n- \n
- NVIDIA GPU-based servers and compute infrastructure \n
- NVIDIA H100, H200, B200, B300, GB200 or GB300 platforms \n
- NVIDIA HGX/DGX architecture \n
- AI/HPC server environments \n
- High-density compute infrastructure \n
- Supermicro \n
- ASUS \n
- Gigabyte \n
- Dell \n
- HPE \n
- GPUaaS, AI Cloud or large-scale AI data center environments \n
Educational Qualification:
\nBachelor's degree or diploma in Computer Science, Electronics, Electrical Engineering, Information Technology, or a related technical discipline.
\nRelevant hardware, server or OEM certifications will be an added advantage.