Back

Intern for AI Security RL Gyms

AliceHanoi, Vietnam
You will be redirected to the main website
Posted on: 27 Sep 2026

Job Description

Alice builds adversarial evaluation infrastructure used by the world's leading AI labs to stress-test their most capable agents before deployment. We're hiring interns for our RL Gym environments: adversarial training grounds where AI agents face prompt injection, social engineering, and data exfiltration scenarios.The Role: You design prompt injection scenarios in YAML, run them against frontier models, validate success rates, and submit passing tasks.The TaskYou build adversarial prompt injection tasks for Alice's RL Gym platform. Each task is a self-contained YAML scenario simulating a realistic AI agent deployment, testing whether the agent can be manipulated into violating its safety policies.What a task includes:An agent persona with a system prompt, security policies, and a defined set of tools.A simulated environment (e-commerce site, messaging platform, enterprise app) with content the agent interacts with.An adversarial payload embedded in the environment (in messages, product listings, tool responses, documents) that attempts to trick the agent into a policy violation.A benign user request that naturally leads the agent to encounter the payload.Deterministic evaluation criteria specifying exactly what constitutes success (agent resists) vs. failure (agent is compromised).

Job Overview

Salary:
Not disclosed
JOB TYPE:
Not specified
Experience:
Not mentioned
Job Location:
Hanoi, Vietnam
Job Level:
Not specified
Education:
Graduation
Intern for AI Security RL GymsAlice