Artificial Intelligence (AI) is transforming our lives and is everywhere. At INTEL, we are part of this AI revolution. Our software stack integrates seamlessly into customer frameworks, which are used by millions of end users. As our AI engineering team in Shanghai continues to grow, we are looking for a passionate engineer to help us deliver high-performance and high-quality deep learning solutions to our customers. Our team’s work includes:
• Optimizing performance for key use cases/models, debugging and resolving issues related to accuracy and memory management
• Designing and developing model deployment architectures, such as implementing new features on vLLM/SGLang to accelerate inference (e.g., P/D disaggregation)
• Developing and debugging high-performance kernels for INTEL accelerators
• Communicating with direct teammates, collaborators, and architects to discuss issues, propose solutions, provide status updates, and gather feedback
• Applying innovative ideas to enhance our products