Responsibilities
• Research, prototype, and implement AI methods that improve model efficiency, inference performance, and deployment feasibility on constrained devices.
• Optimize modern LLMs, SLMs, VLMs, multimodal models, and agentic workloads across post-training, inference, and deployment workflows.
• Propose and evaluate novel compression methods (PTQ, QAT, pruning, low-rank approximation, etc) for on-device LLM/VLM enablement.
• Devise approaches to address challenges related to long-context inference and KV cache compression in the context of reasoning and agentic applications.
• Develop gradient-free and backpropagation-free methods for model merging, compression, and efficiency-driven optimization.
• Implement and evaluate emerging efficient architectures and modules, including MoE, SSMs, hybrid models, Looped Transformers, etc.
• Prototype inference-time optimization methods such as speculative decoding, constrained decoding, low-latency generation, and kernel-level optimization.
• Build experimental pipelines, perform evaluations on standardized language, vision, reasoning, and agentic benchmarks.
• Contribute to publications, technical reports, open-source releases, invention disclosures, and IP submissions where appropriate.