Originating from Bosch AI research in Silicon Valley, our Vision and Language AI Group advances cutting-edge research in two core AI areas: large language models (LLMs) and 3D computer vision. In the LLM space, we focus on AI agents, retrieval-augmented generation, and effective adaptation of LLMs to Bosch domain-specific applications. In 3D vision, we advance spatial intelligence with 3D world models that perceive, reconstruct, and simulate the physical world, enabling reliable embodied control and sim-to-real transfer. By integrating these two areas, we aim to build intelligent systems that can perceive, reason, and communicate seamlessly across both language and spatial environments. We also actively collaborate with leading groups in academia and industry to promote research ideas and publish research findings in internationally renowned conferences and journals such as ACL, EMNLP, CVPR, ICCV, ECCV, NeurIPS, AAAI and CoRL.