{"product_id":"vision-language-models-building-vlms-with-hugging-face-9798341624047","title":"Vision Language Models: Building Vlms with Hugging Face","description":"\u003cp\u003eVision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. \u003cem\u003eVision Language Models\u003c\/em\u003e is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.\u003c\/p\u003e \u003cp\u003eDesigned for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries.\u003c\/p\u003e \u003cp\u003e \u003c\/p\u003e\u003cul\u003e \u003cli\u003eExplore core model architectures and alignment techniques\u003c\/li\u003e \u003cli\u003eTrain and fine-tune VLMs with Hugging Face, PyTorch, and others\u003c\/li\u003e \u003cli\u003eDeploy models for applications like image search and captioning\u003c\/li\u003e \u003cli\u003eImplement advanced inference strategies, from zero-shot to agentic systems\u003c\/li\u003e \u003cli\u003eBuild scalable VLM systems ready for production use\u003c\/li\u003e \u003c\/ul\u003e\u003cbr\u003e\u003cbr\u003e\u003cb\u003eAuthor:\u003c\/b\u003e Merve Noyan,AndrÃ©s Marafioti,Miquel FarrÃ©\u003cbr\u003e\u003cb\u003ePublisher:\u003c\/b\u003e O'Reilly Media\u003cbr\u003e\u003cb\u003ePublished:\u003c\/b\u003e 07\/14\/2026\u003cbr\u003e\u003cb\u003ePages:\u003c\/b\u003e 406\u003cbr\u003e\u003cb\u003eBinding Type:\u003c\/b\u003e Paperback\u003cbr\u003e\u003cb\u003eWeight:\u003c\/b\u003e 1.43lbs\u003cbr\u003e\u003cb\u003eSize:\u003c\/b\u003e 9.19h x 7.00w x 0.84d\u003cbr\u003e\u003cb\u003eISBN:\u003c\/b\u003e 9798341624047\u003cbr\u003e\u003cbr\u003e\u003cb\u003eAbout the Author\u003c\/b\u003e\u003cbr\u003e\u003cb\u003e\u003ci\u003eFarré, Miquel:\u003c\/i\u003e\u003c\/b\u003e - Miquel Farré is a video technology expert with over 15 years of experience and more than 60 patents in machine learning and information science. His career began at the Fraunhofer Institute, where he designed advanced video codecs, and Nagravision, where he developed video streaming security modules. Transitioning to video understanding, Miquel joined Disney to architect the enterprise content metadata platform, leading machine learning initiatives across Pixar, Marvel, Lucasfilm, ABC, and ESPN. He then moved to YouTube, driving search monetization before expanding his focus to lead monetization for the platform's Home and Watch Next surfaces. Before joining Studio Jadu, he worked at Hugging Face on video multimodal large language models and founded Arbro AI to build automated farming solutions.\u003cb\u003e\u003ci\u003eNoyan, Merve:\u003c\/i\u003e\u003c\/b\u003e - Merve Noyan is a machine learning engineer working in the ML advocacy engineering team at Hugging Face. She builds tools to enable people to build with vision language models across the Hugging Face ecosystem (transformers, TRL, smolagents). Previously she worked for different companies building natural language understanding based solutions on information retrieval and conversational agents.\u003cb\u003e\u003ci\u003eZohar, Orr:\u003c\/i\u003e\u003c\/b\u003e - Orr Zohar is a PhD candidate in SVL at Stanford University, advised by Professor Serena Yeung-Levy and supported by the Knight-Hennessy Scholarship. His research centers on large multimodal models, particularly in video understanding, with a focus on self-training methodologies and agentic design. Orr has co-developed innovative approaches such as Video-STaR, a self-training method for video instruction tuning, and VideoAgent, an agent-based framework for long-form video comprehension. Notably, he led the Apollo project, a comprehensive study exploring video understanding in large multimodal models, resulting in the creation of the Apollo family of models that set new benchmarks in the field.\u003cbr\u003eEt al...\u003cbr\u003e","brand":"booksdeli.com","offers":[{"title":"Merve Noyan \/ Paperback \/ English","offer_id":48840906145949,"sku":"9798341624047","price":95.98,"currency_code":"USD","in_stock":true}],"thumbnail_url":"\/\/cdn.shopify.com\/s\/files\/1\/0619\/5648\/9373\/files\/img_e4d1221d-5c57-4cd6-b8ee-bb467ba05cb7.jpg?v=1787728181","url":"https:\/\/booksdeli.com\/products\/vision-language-models-building-vlms-with-hugging-face-9798341624047","provider":"booksdeli.com","version":"1.0","type":"link"}