Site Reliability Engineer (SRE) Intern — AI Infrastructure

TencentTencentPalo Alto, California, United States
Join the waitlist to applySave this job

Invite-only right now: save jobs, track applications, build tailored resumes

Already have an account? Log in

Posted

9/22/2026

Employment

Intern

Range

$27 - $52/hr

Work style

On-site

AI documents

Powered by AI

Jigup writes these against this posting once you are in, using the profile you build once.

Career path

See where this role leads

Jigup maps the next three moves from a job like this one, with the titles and the skills each step asks for.

Join the waitlist

AI summary

Core responsibilities

The SRE intern will support the deployment, configuration, and maintenance of high-performance AI infrastructure, including servers, storage, and networking equipment. They will also provide first-line engineering support for operational issues and document incident resolutions to improve team knowledge.

Requirements overview

Candidates must be pursuing or have recently completed a degree in computer engineering, computer science, or a related technical field. Proficiency in Linux environments and exposure to scripting languages like Bash or Python are required, along with strong analytical and problem-solving skills.

Key skills

Site Reliability EngineeringAI InfrastructureLinuxBashPythonKubernetesSlurmPrometheusGrafanaELKHardware diagnosticsFirmware updatesNetworkingCI/CDConfiguration managementDistributed systems

Resume keywordsJigup Pro

This job lists resume keywords

The terms this posting uses, pulled out so you can mirror them in your resume. Jigup Pro members see the list on the job board.

Join the waitlist

Education requirements

bachelor degreepostgraduate degree

About Tencent

Industry

Software Development

Employees

89,685

Type

Public Company

Size

10,001+ employees

Tencent is a world-leading internet and technology company that develops innovative products and services to improve the quality of life of people around the world. Founded in 1998 with its headquarters in Shenzhen, China, Tencent's guiding principle is to use technology for good. Our communication and social services connect more than one billion people around the world, helping them to keep in touch with friends and family, access transportation, pay for daily necessities, and even be entertained. Tencent also publishes some of the world's most popular video games and other high-quality digital content, enriching interactive entertainment experiences for people around the globe. Tencent also offers a range of services such as cloud computing, advertising, FinTech, and other enterprise services to support our clients' digital transformation and business growth. Tencent has been listed on the Stock Exchange of Hong Kong since 2004.

View company page

Job categories

TechnologyEngineeringSoftwareData & Analytics

Description

About the Hiring Team What the Role Entails Role Summary We are seeking a motivated Site Reliability Engineer (SRE) Intern to join our AI Compute team, supporting the daily operations of AI infrastructure. In this role, you will work closely with internal business teams and external engineering partners to build and operate AI infrastructure. This is a hands-on opportunity to gain knowledge and experience in cutting-edge AI infrastructure. Key Responsibilities Support the deployment, configuration, and maintenance of high-performance AI infrastructure servers, storage servers, networking equipment, and software components in secure environments. Assist with hardware diagnostics, system functionality checks, and firmware updates as required. Collaborate with engineering teams to help deliver tailored customer environments (e.g., bare-metal systems, Kubernetes, Slurm, etc.). Provide first-line engineering support for onsite operational issues, including troubleshooting hardware, network, and software problems, and firmware compliance. Document incident details, resolutions, and lessons learned to improve future problem-solving. Maintain clear, accurate, and up-to-date documentation to support knowledge sharing across the team. Participate in team meetings and knowledge-sharing sessions to foster collaboration and continuous learning. Who We Look For Qualifications & Requirements Currently pursuing or recently completed a Bachelor's or Master's degree in computer engineering, computer science, or a related technical field. Basic understanding of server hardware, firmware lifecycle, and Linux environments, with an awareness of physical and system-level security standards. Exposure to scripting languages such as Bash or Python. Familiarity with — or strong interest in — configuration management, CI/CD tools, workload managers, and cluster software (e.g., Slurm, Kubernetes), and observability tools (e.g., Prometheus, Grafana, ELK). Strong problem-solving and analytical skills. Ability to work both independently and as part of a team. Professional fluency in English and Mandarin is highly preferred Preferred Qualifications Coursework, projects, or hands-on experience related to AI Infrastructure, distributed systems, or cloud infrastructure. Familiarity with networking fundamentals and Linux system administration. Genuine interest in AI/ML infrastructure. Location State(s) US-California-Palo Alto The expected base pay range for this position in the location(s) listed above is $27.12 to $51.93 per hour. Actual pay may vary depending on job-related knowledge, skills, and experience. This position will be eligible for 1 hour of paid sick leave for every 30 hours worked and up to 13 paid holidays throughout the calendar year. Subject to the terms and conditions of the applicable plans then in effect, full-time interns are also eligible to enroll in the Company-sponsored medical plan. Equal Employment Opportunity at Tencent As an equal opportunity employer, we firmly believe that diverse voices fuel our innovation and allow us to better serve our users and the community. We foster an environment where every employee of Tencent feels supported and inspired to achieve individual and common goals.

Requirements

  • Site Reliability Engineering
  • AI Infrastructure
  • Linux
  • Bash
  • Python
  • Kubernetes
  • Slurm
  • Prometheus
  • Grafana
  • ELK
  • Hardware diagnostics
  • Firmware updates
  • Networking
  • CI/CD
  • Configuration management
  • Distributed systems

Benefits

  • Paid sick leave
  • Paid holidays
  • Medical plan

More entry level jobs in Palo Alto, CA

All jobs in Palo Alto, CA

More jobs at Tencent

Ready to apply?

Jigup is invite-only right now. Join the waitlist to save jobs, track applications, and build tailored resumes.

Join the waitlist