We propose collaborative research in the following areas, with flexibility to refine topics based on mutual expertise:
•
Efficient Scheduling for Sparse & Dense LLMs:
Design token-aware, load-balanced scheduling algorithms for MoE and hybrid LLM workloads that reduce inter-GPU communication and optimize heterogeneous cluster utilization.
•
Efficient Inference for State Space Models
Develop high-throughput, low-latency inference techniques for state space models, leveraging their linear-time properties to outperform traditional attention mechanisms in long-context scenarios.
•
Memory-Aware Training & Serving
Explore advanced quantization, memory-efficient checkpointing, offloading strategies, and dynamic memory management techniques to support training and inference of ultra-large models.
•
Scalable Parallelism for LLMs
Investigate hybrid parallelism (data, model, pipeline, expert) and communication-reduction strategies tailored for scaling LLMs across thousands of GPUs.
•
Hardware-Aware Optimization
Develop compiler, kernel, and data layout optimizations that fully exploit features of modern GPU architectures, improving throughput for both dense and sparse model operations.
•
High-Throughput, Low-Latency Inference
Create optimized model serving strategies using speculative decoding, continuous batching, expert routing, and adaptive computation for production-grade LLM applications.