Deep, hands-on Transformer expertise
This is a hard requirement. You must have personally made material architecture or training decisions in at least one substantial Transformer-based system, such as:
•
Vision-language or vision-language-action models
•
Multimodal foundation models
•
Transformer-based perception, planning, or control systems
You should be able to explain how you represented and tokenized inputs and outputs, fused modalities, structured attention and temporal context, selected losses, built data mixtures, distributed and monitored training, diagnosed failures, and changed the system to improve task performance.
Using hosted model APIs, prompting language models, or running an unchanged public training recipe does not meet this requirement.
You have worked on machine learning for a system that perceives or acts in the physical world, such as robotics, autonomous vehicles, drones, industrial automation, manipulation, mobile robots, humanoids, or wearable and egocentric systems.
At least one substantial project must have progressed beyond offline datasets or simulation into a real or operational physical system. Simulation experience qualifies only when paired with credible sim-to-real ownership and physical validation.
Hands-on technical ability
•
Strong Python engineering skills
•
Direct experience with PyTorch or an equivalent deep-learning framework
•
Ability to read and debug unfamiliar model and training code
•
Experience designing controlled experiments and analyzing rollout failures
•
Experience working with large multimodal datasets
•
Sound judgment around compute, memory, training stability, inference latency, and cost
You have set the direction for a significant research, model-development, robotics, or cross-functional technical program. You have made architecture and resource decisions, mentored or hired technical talent, stopped weak lines of work, and helped take a result into deployment.
A management title is not required, but direct technical ownership is.