{"product_id":"vision-language-models-building-vlms-with-hugging-face-9798341624047","title":"Vision Language Models: Building Vlms with Hugging Face","description":"\u003cp\u003eVision language models (VLMs) combine computer vision and natural language processing to create powerful systems that can interpret, generate, and respond in multimodal contexts. \u003cem\u003eVision Language Models\u003c\/em\u003e is a hands-on guide to building real-world VLMs using the most up-to-date stack of machine learning tools from Hugging Face, Meta (PyTorch), NVIDIA (Cuda), OpenAI (CLIP), and others, written by leading researchers and practitioners Merve Noyan, Miquel Farré, Andrés Marafioti, and Orr Zohar. From image captioning and document understanding to advanced zero-shot inference and retrieval-augmented generation (RAG), this book covers the full VLM application and development lifecycle.\u003c\/p\u003e \u003cp\u003eDesigned for ML engineers, data scientists, and developers, this guide distills cutting-edge VLM research into practical techniques. Readers will learn how to prepare datasets, select the right architectures, fine-tune and deploy models, and apply them to real-world tasks across a range of industries.\u003c\/p\u003e \u003cp\u003e \u003c\/p\u003e\u003cul\u003e \u003cli\u003eExplore core model architectures and alignment techniques\u003c\/li\u003e \u003cli\u003eTrain and fine-tune VLMs with Hugging Face, PyTorch, and others\u003c\/li\u003e \u003cli\u003eDeploy models for applications like image search and captioning\u003c\/li\u003e \u003cli\u003eImplement advanced inference strategies, from zero-shot to agentic systems\u003c\/li\u003e \u003cli\u003eBuild scalable VLM systems ready for production use\u003c\/li\u003e \u003c\/ul\u003e\u003cbr\u003e\u003cbr\u003e\u003cb\u003eBinding Type:\u003c\/b\u003e Paperback\u003cbr\u003e\u003cb\u003ePublisher:\u003c\/b\u003e O'Reilly Media\u003cbr\u003e\u003cb\u003ePublished:\u003c\/b\u003e 07\/14\/2026\u003cbr\u003e\u003cb\u003eISBN:\u003c\/b\u003e 9798341624047\u003cbr\u003e\u003cb\u003ePages:\u003c\/b\u003e 406\u003cbr\u003e\u003cb\u003eWeight:\u003c\/b\u003e 1.43lbs\u003cbr\u003e\u003cb\u003eSize:\u003c\/b\u003e 9.19h x 7.00w x 0.84d","brand":"Merve Noyan,Andr\u0026#195;\u0026#169;s Marafioti,Miquel Farr\u0026#195;\u0026#169;","offers":[{"title":"Default Title","offer_id":52732238561461,"sku":"9798341624047","price":67.99,"currency_code":"USD","in_stock":true}],"url":"https:\/\/pastforward.org\/products\/vision-language-models-building-vlms-with-hugging-face-9798341624047","provider":"Past Forward","version":"1.0","type":"link"}