Machine Learning Engineer - RL
Machine Learning Engineer (Reinforcement Learning Systems)
Overview
We are looking for a Platform Engineer to build the infrastructure, tooling, and systems that power large-scale Reinforcement Learning (RL) workflows. This role focuses on enabling researchers to train, evaluate, and deploy RL models efficiently by providing a scalable and reliable experimentation platform.
You will work at the intersection of distributed systems engineering and ML research, building platforms that abstract away infrastructure complexity and enable “self-serve” experimentation for research teams.
About Deccan AI
Deccan AI is a fast-growing, venture-backed AI infrastructure company focused on training, evaluating, and improving next-generation AI systems. Headquartered in the Bay Area, with a growing India hub in Hyderabad, the company was founded by alumni of IIT Bombay, IIM Ahmedabad, and former Google leaders.
We work with some of the world’s leading AI frontier labs and research organizations, including Google DeepMind, Snowflake, and other cutting-edge AI teams. Backed by Prosus Ventures, Deccan AI recently raised $25M in Series A funding and is entering a significant growth phase.
With a global network of over 1 million experts, advanced automation systems, and vertically integrated platforms, we deliver the high-quality data and evaluation infrastructure that state-of-the-art AI models depend on. As the AI infrastructure market rapidly expands, Deccan AI is building the systems powering the future of AI.
About the Role
We're hiring an ML Engineer to work on Reinforcement Learning, partnering directly with a frontier AI lab. You'll design and train agents that learn, adapt, and improve working on RL and RLHF systems that sit at the core of how modern AI systems are trained and aligned.
What You'll Do
- Design, implement, and train RL agents using policy-gradient and value-based methods (PPO, DQN, SAC, etc.)
- Build and tune reward models and simulation environments for agent training
- Work on multi-agent RL systems and RLHF pipelines for model alignment
- Run large-scale training experiments and analyze agent behavior/failure modes
- Collaborate with research teams to translate RL research into production-ready systems
What We're Looking For
- 2–4 years of hands-on experience in reinforcement learning or related ML research
- Strong Python skills and experience with PyTorch, TensorFlow, or JAX
- Familiarity with Ray RLlib, Stable Baselines3, CleanRL, Gymnasium, or MuJoCo
- Solid understanding of core RL algorithms: PPO, DQN, SAC, actor-critic methods, reward shaping
- Bonus: published RL research, RLHF pipeline experience, or multi-agent systems work
Why Join
- Direct exposure to frontier-lab-scale RL and alignment problems
- Work alongside top ML research and engineering talent
- High-ownership role with real technical depth
RoundInterview FocusKey Areas to Evaluate
1st Round
Introduction & Experience
Candidate background, current work, domain expertise, key projects, technical contributions, and understanding of the company/role
2nd Round
Hands-on Coding – Agentic AI
Live/on-call coding, problem-solving ability, coding fundamentals, agentic AI use cases, system implementation, and practical engineering skills
3rd Round
RL & Model Training – Deep Dive
Reinforcement Learning concepts, model training, LLM fine-tuning, post-training techniques, RL fine-tuning, and depth of hands-on experience
Recommended Jobs
Posted just now
Posted just now
Posted just now
Posted just now
Posted just now

