AutoModelForImageTextToText
Found in 2 packages: transformers, vllm
transformers(4 releases)
v4.51.3-InternVL-previewThis preview release introduces support for the InternVL 2.5 and 3 family of multimodal models, featuring a native multimodal pre-training paradigm and state-of-the-art performance on visual-linguistic tasks.
v4.49.0-Mistral-3This release introduces Mistral 3 (Mistral Small 3.1) to the Transformers library, a 24B parameter model featuring 128k context length and advanced vision-language capabilities.
v4.49.0-AyaVisionThis release introduces Aya Vision 8B and 32B, multilingual multimodal models combining SigLIP-2 vision encoders with Cohere language models, available via a specialized transformers release tag.
v4.49.0-SmolVLM-2This release introduces SmolVLM-2, a lightweight vision-language model based on Idefics3 and SmolLM2 that supports multi-image and video processing.
vllm(1 releases)
Track Symbol Changes
Use the Change8 MCP server or GitHub Action to get notified when AutoModelForImageTextToText changes.
Learn More