As a Software Engineer on the RL Research & Environments team, you will design and operate the data, evaluation, and environment systems that improve model capabilities after pre-training.
This role focuses on post-training: identifying capability gaps, building targeted datasets, designing reward signals, and running iterative training loops that measurably improve user-facing behavior. You will own the infrastructure and experimental workflows that connect product priorities to concrete capability gains.
Magic’s long-context models introduce distinct post-training challenges: long-horizon reasoning, sustained coherence over extended trajectories, context-use quality, and tool-augmented behavior. You will build systems that expose failure modes, generate high-signal training data, and enable rapid RL iteration at scale.
This role can evolve into ownership of major capability areas, deeper RL systems work, or broader influence over post-training strategy as Magic scales long-context model performance and reliability.