Intern: AI Red Teaming (Fall 2026)

Realm LabsRealm LabsSunnyvale, California, United States
Join the waitlist to applySave this job

Invite-only right now: save jobs, track applications, build tailored resumes

Already have an account? Log in

Posted

8/23/2026

Employment

Full time

Range

Check website

Work style

On-site

AI documents

Powered by AI

Jigup writes these against this posting once you are in, using the profile you build once.

Career path

See where this role leads

Jigup maps the next three moves from a job like this one, with the titles and the skills each step asks for.

Join the waitlist

AI summary

Core responsibilities

You will stress-test and break AI systems, including LLMs and agentic models, to identify vulnerabilities and unsafe behaviors. You will document these findings through reproducible attacks and technical reports to improve model safety evaluations.

Requirements overview

Candidates should have hands-on experience in adversarial machine learning, red teaming, and deep learning frameworks like PyTorch. Proficiency in Python and familiarity with cloud environments like AWS or GCP are required.

Key skills

Adversarial MLRed TeamingLLMPythonPyTorchHuggingFaceTransformersDeep LearningPrompt InjectionJailbreakingVulnerability ResearchPenetration TestingAWSGCPGitUnix

Resume keywordsJigup Pro

This job lists resume keywords

The terms this posting uses, pulled out so you can mirror them in your resume. Jigup Pro members see the list on the job board.

Join the waitlist

About Realm Labs

Industry

Computer and Network Security

Employees

10

Type

Privately Held

Size

11-50 employees

The Runtime AI Observability and Control platform for the modern enterprise. AI doesn't think in natural language. It thinks math. Realm Labs reads the math in real time to catch hallucinations, prompt injection, and unsafe responses before they become brand-damaging incidents. Works with any chatbot, copilot, or agent.

View company page

Job categories

TechnologySoftwareSecurity & SafetyData & AnalyticsEngineering

Description

ROLE OVERVIEW * You will try to break the systems we build and the models we protect, and turn what you find into something the team can act on: a reproducible attack, an evaluation that catches it, and a written account of why it works. * Expect some mix of: eliciting unsafe behaviour from aligned LLMs and multi-modal models; prompt injection and tool-use abuse against agentic systems; automating attack generation and evaluation rather than hand-crafting one-off prompts; and measuring whether guardrails hold under pressure. Where RealmLabs' interpretability work gives you access to a model's internals, use it. * We aim for a paper or public technical report out of every internship, plus attacks that stay in our evaluation suite after you leave. EXPECTED BACKGROUND: ADVERSARIAL ML AND RED TEAMING * Hands-on experience attacking or stress-testing models, from any direction: jailbreaks, prompt injection, adversarial examples, data poisoning, model extraction, or evaluating safety and moderation systems. * Able to read a paper and implement its attack. * (nice to have) Offensive security background outside ML: CTFs, vulnerability research, penetration testing. * (nice to have) Familiarity with agentic systems and their attack surface — tool calls, retrieval, memory, multi-agent orchestration. EXPECTED BACKGROUND: ML * Machine learning tools: pytorch, huggingface, transformers, datasets. * Applied deep learning and LLM experience. * Training and evaluating deep models. * (nice to have) finetuning LLMs, multi-modal LLMs. * (nice to have) Familiarity with ML[NLP,LLM,Vision] interpretability methods, sparse autoencoders, linear probes — as a way of locating failure modes, not as an end in itself. EXPECTED BACKGROUND: SOFTWARE ENGINEERING * Development environments and tools: * unix, git, basic clouds usage on AWS and/or GCP * jupyter * Programming: * python * (nice to have) “programming languages well-roundedness” * experience in statically-typed and functional languages COMPENSATION & BENEFITS * Market aligned compensation for interns in the bay area. REQUIREMENTS * Must be authorized to work in the USA or must be able to obtain CPT (Curricular Practical Training) approval from host university.

Requirements

  • Adversarial ML
  • Red Teaming
  • LLM
  • Python
  • PyTorch
  • HuggingFace
  • Transformers
  • Deep Learning
  • Prompt Injection
  • Jailbreaking
  • Vulnerability Research
  • Penetration Testing
  • AWS
  • GCP
  • Git
  • Unix

More entry level jobs in Sunnyvale, CA

All jobs in Sunnyvale, CA

Ready to apply?

Jigup is invite-only right now. Join the waitlist to save jobs, track applications, and build tailored resumes.

Join the waitlist